Skip to main content

rl-sales-augment

A trained RL policy that picks the next sales move (rapport / pitch / objection / discount / close); your LLM writes the words. The trained model downloads once on first use (sha-verified, cached): CPU, no finetuning, works with any LLM.

pip install rl-sales-augment

Quickstart (local Gemma, no API key)

The policy was trained alongside Gemma 4; running it fully local (E2B or E4B) is the first-class path:

pip install "rl-sales-augment[gemma]"
import rl_sales_augment as rsa

gen = rsa.providers.gemma_e4b()      # google/gemma-4-E4B-it, auto-downloads, runs on CPU/MPS/CUDA
# gen = rsa.providers.gemma_e2b()    # google/gemma-4-E2B-it, the lighter 5B variant
bot = rsa.load_agent(gen, company_ctx="Acme sells AcmeBox, an $8k on-prem appliance.")

out = bot.reply("honestly it feels expensive vs AWS")
out["chosen_move"]   # 'RAPPORT'  <- the RL decision
out["belief"]        # {'interest': .5, 'trust': .5, 'budget_fit': .2, 'objection': .8, ...}
out["reply"]         # the LLM's words, executing that move

Prefer an API model? Same code, different one-liner:

gen = rsa.providers.openai_chat(model="gpt-5.5")      # or gemini_api() / anthropic_chat()

bot.reply() is stateful: keep calling it, the agent remembers. For stateless use (e.g. behind an API), pass the whole conversation in OpenAI message format:

out = bot.chat([
    {"role": "user", "content": "what does it cost?"},
    {"role": "assistant", "content": "Depends on seats. How many do you need?"},
    {"role": "user", "content": "40 seats, but budget is tight"},
])

Providers

rsa.providers.gemma_e2b()                                     # [gemma]   local Gemma 4 E2B, no API key
rsa.providers.gemma_e4b()                                     # [gemma]   local Gemma 4 E4B, no API key
rsa.providers.openai_chat(model="gpt-5.5")                    # [openai]  OPENAI_API_KEY
rsa.providers.anthropic_chat(model="claude-sonnet-5")         # [anthropic] ANTHROPIC_API_KEY
rsa.providers.gemini_api()                                    # [gemini]  GEMINI_API_KEY
rsa.providers.gemini_vertex()                                 # [gemini]  gcloud ADC + GCP_PROJECT
rsa.providers.openai_chat(base_url="http://...")              # any OpenAI-compatible server

Or bring your own: any gen(prompt) -> str works.

On the Gemma E4B injection path the prompt is NEUTRAL (no persona, no style words): the human voice comes from the RL latent itself, not prompt engineering.

Multilingual: the bot replies in the customer's language automatically (tested: Malayalam, Hindi, Tamil, Japanese, Spanish, German; a built-in script detector keeps even small local models on-language for Indic/CJK/Arabic/Cyrillic scripts). Romanized Indic is supported too: Manglish / Hinglish / Tanglish ("ntha visesham, sugano?") is detected and mirrored on frontier models.

API keys (.env)

Put credentials in a .env file next to where you run your script; providers load it automatically (real environment variables take precedence). Never commit it.

# .env
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
GEMINI_API_KEY=AIza...
GCP_PROJECT=my-gcp-project        # for gemini_vertex (gcloud ADC)

Explicit control: rsa.load_env("/path/to/.env"). Local Gemma needs no key at all.

MCP server

pip install "rl-sales-augment[mcp]"
rl-sales-augment-mcp        # stdio
{ "mcpServers": { "rl-sales-augment": { "command": "rl-sales-augment-mcp" } } }

Tools: next_move (buyer state → RL move), perception_prompt, list_moves, list_segments.

REST API

pip install "rl-sales-augment[gemini,api]"

A complete FastAPI server (POST /v1/chat, OpenAI-format messages) ships in examples/fastapi_server.py.

Why not just call GPT-5.6 / Opus 4.8 / Gemini directly?

Because what kills LLM sales conversations isn't the words, it's the timing. Frontier models are trained to be helpful and agreeable, so on a skeptical buyer they answer every objection politely, forever, and never risk asking for the deal (measured: 0/4 closes on adversarial buyers while handling every question beautifully). A bigger model writes better sentences; it doesn't fix this, because next-token training never rewards a deal that closes six turns later.

The bundled policy is different in kind, not degree:

  • Trained on outcomes, not text. PPO over millions of simulated deals with delayed, stochastic rewards. It has lost deals to premature pitching, burned reputation on spam-closing, and learned that discounting converts SMBs but insults enterprise. An API model has read about selling; the policy has sold.
  • State-dependent timing. Prompting "be assertive, always close" makes a bot uniformly pushy. The skill is when: the policy closes at high readiness and keeps building trust below it. Same LLM writing the words, right moment to ask. Result: 100% vs 19-31% close in a paired A/B, 3/4 vs 0/4 on hard buyers (simulated; harness in the repo).
  • Consistent and auditable. Sampled LLM strategy swings run-to-run; the policy is deterministic, and every turn logs chosen_move + belief, so you can see why it did what it did.
  • Complementary and tiny. A ~1MB MLP on CPU. Keep GPT-5.6 / Opus 4.8 / Gemini for language, empathy, and knowledge; add the decision layer they don't have. Retrainable on your own funnel's economics (the commercial offering).

Details, transcripts, and a 67-second demo: github.com/NandhaKishorM/rl-sales-augment

License

AGPL-3.0-or-later. Training the policy on your own market is the commercial offering: nandakishor@convaiinnovations.com (Convai Innovations Pvt. Ltd.).

Metadata

Release files for rl-sales-augment 0.9.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rl-sales-augment 0.9.7
File Size Uploaded
rl_sales_augment-0.9.7.tar.gz 68.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rl-sales-augment 0.9.7
File Interpreter ABI Platform
rl_sales_augment-0.9.7-py3-none-any.whl Python 3 none any Details

Total release size: 125.7 kB

Release files / rl_sales_augment-0.9.7.tar.gz

Download URL rl_sales_augment-0.9.7.tar.gz
Size 68.9 kB
Tags Source
SHA-256 checksum
How to use checksums
8e8ae35740f649991d5d8d267c79e60c03c36c8a1fca47ad6c75d70287e7f673
BLAKE2b-256 checksum
How to use checksums
6e16ac444c0a4a2b18b452b0e65e8f578c6319dbbc4c5618d71d1956af06ae83
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release files / rl_sales_augment-0.9.7-py3-none-any.whl

Download URL rl_sales_augment-0.9.7-py3-none-any.whl
Size 56.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b96b8a7007f434c7afa7ea3bc93370180afdf360f9112c9e4906fad4b53cadb8
BLAKE2b-256 checksum
How to use checksums
cfbc22c9a4e5de13be84066266655af7b68153aa391acf02079d4f1752fd3eb0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page