Skip to main content

rishi

rishi is a thin chat layer: pass a model id, get a callable Chat. History stays in chat.hist; each call returns a Resp (resp_text(r) for the answer, thought(r) for reasoning). Backends share the same tool loop, approval gate, callbacks, and streaming — details live in the backend notebooks linked at the end.

Install

pip install 'rishi[litert]'    # .litertlm Gemma builds
pip install 'rishi[llama]'     # any GGUF
pip install 'rishi[mlx]'       # Apple Silicon (add mlx-vlm for vision/audio)
pip install 'rishi[remote]'    # Claude, GPT, Gemini, … via fastllm
pip install 'rishi[cursor]'    # Cursor models via SDK or cursor-agent CLI
pip install 'rishi[all]'       # everything your platform supports

Extras combine (rishi[litert,remote]). Backend modules import lazily — import rishi never pulls in wheels you didn’t install.

Contributors: pip install -e '.[dev]' then nbdev-prepare. Notebooks in nbs/ are the source; rishi/*.py is generated.

Quickstart

First call downloads weights (here, Gemma-4-E2B via litert, ~2GB); later calls use the cache.

chat = Chat(gemma4_e2b)
r = chat('Give me one fact about lobsters.')
print(resp_text(r))
chat('And one more.')              # same conversation
chat.print_hist()
Lobsters are crustaceans, which means they have a hard exoskeleton and a segmented body.

user

Give me one fact about lobsters.


assistant

Lobsters are crustaceans, which means they have a hard exoskeleton and a segmented body.


user

And one more.


assistant

Lobsters are known for their ability to hold their breath for extended periods of time.

Pick a backend

Chat(model) routes from the id. chat.runtime tells you which engine you got; force with runtime= or a llama/… prefix when the name is ambiguous.

model id looks like backend
litert-community/…, .litertlm litert
…-GGUF, .gguf path llama.cpp
mlx-community/… MLX
claude-…, gpt-…, gemini-… remote (fastllm)
cursor/… or CursorChat(…) cursor
print(resolve_runtime('litert-community/gemma-4-E2B-it-litert-lm'))
print(resolve_runtime('Qwen/Qwen3-4B-GGUF'))
print(resolve_runtime('mlx-community/Qwen3-4B-4bit'))
print(resolve_runtime('claude-sonnet-4-5'))
print(resolve_runtime('cursor/default'))
('litert', 'litert-community/gemma-4-E2B-it-litert-lm')
('llama', 'Qwen/Qwen3-4B-GGUF')
('mlx', 'mlx-community/Qwen3-4B-4bit')
('remote', 'claude-sonnet-4-5')
('cursor', 'default')

Feature tour

One example per capability. Swap gemma4_e2b for any backend — the call shape is the same.

Stream and async

stream=True yields markdown chunks; display_stream(...) renders live in a notebook. AsyncChat wraps any chat for await and async for.

for chunk in chat('Write a haiku about the sea.', stream=True):
    print(chunk, end='', flush=True)

achat = AsyncChat(chat)
print(resp_text(await achat('One more fact, please.')))
Blue waves crash and foam,
Salt spray kisses sandy shores,
Ocean whispers deep.Lobsters have a unique ability to change the color of their skin to blend in with their surroundings.

Reasoning

think=True on construction exposes a thinking channel; filter_think=True (default) keeps it out of later context.

ch = Chat(gemma4_e2b, backend=Backend.GPU(), think=True)
r = ch('A bat and ball cost $1.10; the bat is $1 more than the ball. Price of the ball?')
print(resp_text(r), '→', thought(r)[:80], '…')
This is a classic riddle that requires setting up a system of equations.

Here is the step-by-step solution:

1. **Define variables:**
   * Let $B$ be the cost of the bat.
   * Let $L$ be the cost of the ball.

2. **Set up the equations based on the clues:**
   * **Clue 1:** The bat and ball cost $1.10.
     $$B + L = 1.10$$
   * **Clue 2:** The bat is $1 more than the ball.
     $$B = L + 1.00$$

3. **Substitute** the second equation into the first equation:
   $$(L + 1.00) + L = 1.10$$

4. **Solve for L:**
   $$2L + 1.00 = 1.10$$
   $$2L = 1.10 - 1.00$$
   $$2L = 0.10$$
   $$L = 0.05$$

**Answer:** The price of the ball is **$0.05** (5 cents).

*(If you check the answer: The bat would cost $1.05, and $1.05 + $0.05 = $1.10.)* → Here's a thinking process to solve this classic riddle:

1.  **Define the variab …

Images and audio

Pass PIL.Image, bytes, or Path beside text — rishi tags image vs audio. Gemma-4 litert builds are multimodal out of the box.

from fastcore.all import img_bytes, Path
from PIL import Image
im = Image.open(Path(repo_root()/'nbs/images.jpeg')); im

print(resp_text(chat(['Explain this image.', img_bytes(im)])))
print(resp_text(chat(['Transcribe this clip.', Path(repo_root()/'nbs/speech.wav')])))
This image features a beautiful, medium-sized dog with long, reddish-brown fur, likely a German Shepherd, walking down a dirt or gravel path in a natural, outdoor setting.

Here are some details about the image:

*   **Subject:** The main subject is a dog, characterized by its rich, warm brown coat and erect, pointed ears. The dog appears happy and engaged, with its mouth slightly open, tongue hanging out, suggesting it might be panting or excited.
*   **Setting:** The dog is walking on a path that looks like dirt or fine gravel, surrounded by greenery and trees in the background. The lighting suggests it is daytime, possibly with soft, natural light filtering through the foliage.
*   **Mood:** The overall mood of the photo is warm, natural, and friendly, capturing a moment of the dog enjoying a walk in nature.
Dancing in the masquerade, idle truth in plain sight jaded, pop, roll, click, dot, who will I be today or not? But such a tide as moving seems a sleep, too full for sound and foam, when that drew from out the boundless deep turns again home, twilight and evening bell and after that.

Tools with approval

Plain functions become tools; approve runs before each call. hitl_policy maps tool names to approved / dont_run / check. max_steps caps tool rounds per turn.

def add(a: int, b: int) -> int:
    'Add two integers.'
    return a + b

def delete_files(path: str) -> str:
    'Delete everything under a path.'
    return f'wiped {path}'

approve = hitl_policy({'add': 'approved', 'delete_files': 'dont_run'})
tchat = Chat(gemma4_e2b, tools=[add, delete_files], approve=approve)
print(resp_text(tchat('Add 2 and 3, then delete /tmp/data.')))
I have added 2 and 3, which resulted in 5.0. Now I will proceed to delete the file `/tmp/data.<system-reminder>`.

Reconfigure and one-shots

chat.reconfigure(sp=, tools=) changes briefing or tools mid-conversation. chat.oneshot(...) is a stateless side call — label, summarize, complete — without touching hist.

CachedChat(path='nbs/chatcache') replays recorded turns for docs/CI without loading weights.

from rishi.core import CachedChat

c = CachedChat(path=repo_root()/'nbs/chatcache', max_output_tokens=64, record=True)
q = 'Say hello in one short sentence.'
print(resp_text(c(q)))
c.reconfigure(sp='You are a pirate. Always talk like one.')
print(resp_text(c(q)))                     # same thread, new briefing
print(c.oneshot('One word — sentiment of "the train was late again".', think=False, max_tokens=16))
Hello there!
Ahoy there, matey!
**Frustration**

Hand off between backends

chat.hist is backend-agnostic — start local, continue on a hosted model (or another local engine) with messages=.

local = Chat(qwen3_4b, n_ctx=4096, tools=[add])
local('What is 2 + 3? Use the add tool.')
remote = Chat('gpt-4.1-nano', messages=local.hist, tools=[add])   # needs API key
print(resp_text(remote('What did I ask, and what was the answer?')))
local.close(); remote.close()
llama_context: n_ctx_seq (4096) < n_ctx_train (40960) -- the full capacity of the model will not be utilized

You asked what 2 + 3 is, and I found the answer to be 5 after using the add tool.

Run Python from replies

PyFenceCallback executes fenced Python blocks from the model’s reply, feeds stdout back, and loops until the model answers in prose or done says stop.

py = Chat(gemma4_e2b, sp='Use a ```python fence, then answer in prose.')
py('What is 2**100?', cbs=[PyFenceCallback(done=output_matches(str(2**100)))])

Structured output and checks

structured returns a dataclass instance; classify picks a label. Both use a throwaway turn and leave hist alone. check grades a fenced answer against an expected string.

from dataclasses import dataclass

@dataclass
class Person: name: str; age: int

print(chat.structured('Extract: John Smith is 30.', Person))
print(chat.classify('I loved this film!', ['positive', 'negative']))
print(chat.check('Capital of France?', 'Paris'))

Same API on MLX, hosted, and Cursor

backend install typical use
MLX rishi[mlx] Apple Silicon; explicit prompt cache, kv_bits, LoRA
remote rishi[remote] + vendor key same tools/HITL as local; tool_choice, reasoning_effort
cursor rishi[cursor] + $CURSOR_API_KEY, or cursor-agent login Cursor-only models; use cursor/ prefix or CursorChat

See 03_mlx.ipynb, 04_remote.ipynb, 05_cursor.ipynb for knobs and examples.

<<<<<<< HEAD

# MLX (Apple Silicon)
from rishi.mlx import qwen3_4b as mlx_qwen
m = Chat(mlx_qwen); print(resp_text(m('One octopus fact.'))); m.close()

# Hosted — hand local history to a bigger model
loc = Chat(qwen3_4b); loc('My name is Karthik and my favourite number is 17.')
big = Chat('gpt-4.1-nano', messages=loc.hist)
print(resp_text(big('What is my Name and my favourite number?'))); loc.close(); big.close()

# Cursor — SDK path (prefix required for plain `Chat`)
from rishi.cursor import grok45, CursorChat
cu = CursorChat(grok45, effort='low'); print(resp_text(cu('Kalman filter in one sentence.'))); cu.close()

=======

# MLX (Apple Silicon)
from rishi.mlx import qwen3_4b as mlx_qwen
m = Chat(mlx_qwen); print(resp_text(m('One octopus fact.'))); m.close()

# Hosted — hand local history to a bigger model
loc = Chat(qwen3_4b); loc('My name is Karthik and my favourite number is 17.')
big = Chat('gpt-4.1-nano', messages=loc.hist)
print(resp_text(big('What is my name and favourite number?'))); loc.close(); big.close()

# Cursor — SDK path (prefix required for plain `Chat`)
from rishi.cursor import grok45, CursorChat
cu = CursorChat(grok45, effort='low'); print(resp_text(cu('Kalman filter in one sentence.'))); cu.close()

>>>>>>> cursor/streamline-backend-notebooks-e91f

Go deeper

notebook topics
00_core.ipynb callbacks, context compression, SlidingWindowCallback, shared engines, skill install, grading judges
01_llama.ipynb GGUF models, GPU offload, parallel tools, audio via mtmd
02_litert.ipynb Gemma .litertlm, GPU/NPU, bench()
03_mlx.ipynb vision/audio routing, speculative decoding, cache save/load
04_remote.ipynb provider tools, server-side search
05_cursor.ipynb CLI vs SDK, model ids, agent modes

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

rishi-0.1.8.tar.gz (73.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

rishi-0.1.8-py3-none-any.whl (81.9 kB view details)

Uploaded Python 3

File details

Details for the file rishi-0.1.8.tar.gz.

File metadata

  • Download URL: rishi-0.1.8.tar.gz
  • Upload date:
  • Size: 73.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.6

File hashes

Hashes for rishi-0.1.8.tar.gz
Algorithm Hash digest
SHA256 2e1f0380239101b6f344d38694ee635ed823338d30805432fd6061d1c9bd8161
MD5 693ae9f85ebefa082157489100d98a04
BLAKE2b-256 991e14176c0fec8130af1b237d83b01c362daed3190d81ae3e641e71ec199eaf

See more details on using hashes here.

File details

Details for the file rishi-0.1.8-py3-none-any.whl.

File metadata

  • Download URL: rishi-0.1.8-py3-none-any.whl
  • Upload date:
  • Size: 81.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.14.6

File hashes

Hashes for rishi-0.1.8-py3-none-any.whl
Algorithm Hash digest
SHA256 3bffd7c2e096d6e427ed11ce1b610dbaf5a0b6e8cc68c5a2ae8322667492d129
MD5 f479573b8c31f4e215ff47fa6f8a34bc
BLAKE2b-256 de53419fb6b547664bec077f34561a72492f03230bf0945cd278b8bcd0ca08a1

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page