sallm-agent
Minimal tool-calling agent (sallm) for local small LLMs (default: Gemma 4 4B via Ollama).
Focus: keep context predictable over long sessions — durable SQLite state, LanceDB retrieval, a skill stack, and a visible ContextReceipt that explains token spend.
Not an invisible RAG black box: raw messages stay canonical; vectors are a rebuildable index; derived facts must cite source message ids.
Setup
ollama pull gemma4:e4b-it-qat
ollama pull qwen3-embedding:0.6b
uv sync --extra dev
Core deps include peewee (SQLite ORM) and lancedb (vector index). There is no DSPy/Pydantic dependency.
Examples
examples/imap_inbox/— durable IMAP inbox Q&A (CLI tools + long-session recall). See the docstring inagent.py.examples/linux_history/— durable bash/zsh history Q&A + livebash_run(usage-story inference). See the docstring inagent.py.
Turn pipeline
user → persist raw message
→ goal/skill control (small JSON call)
→ vector retrieve (Qwen embed + LanceDB)
→ budgeted prompt + ReAct ```run tools
→ persist answer
→ extract grounded facts + index chunks
Library
from sallm import Agent, RetrievalConfig, Skill, SkillRegistry
from sallm.tools import builtin_tools
agent = Agent(
tools=builtin_tools(("calc", "echo")),
state_path="/tmp/sallm/state.db",
vector_path="/tmp/sallm/vectors",
session_id="demo",
retrieval=RetrievalConfig(
memory_gate=True,
search_mode="dense", # or "hybrid"
use_instruct=True,
use_rewrite=False,
use_hyde=False,
),
)
result = agent.ask("Remember the code is PURPLE-42.")
print(result["answer"])
print(result["receipt"]) # ContextReceipt as dict
print(result["goal"], result["stack"])
Resume by reusing state_path + session_id.
Preload memory (meaning-first): agent.remember(text, source="…") runs an ingest LLM prompt, stores English facts for retrieval, and does not fill the recent-history window. Use for shell-history blocks and other raw dumps that should answer ordinary-English questions later. See docs/agent-instructions.md.
VectorStore contract
Implement upsert / search / delete_session / close (see sallm.memory.types.VectorStore). Default: LanceVectorStore. A future pgvector adapter can satisfy the same dataclasses (VectorRecord, VectorQuery, VectorHit) without changing the agent.
SQLite stores chunk text + indexed flags; LanceDB is rebuilt from those rows after a crash.
Skills
Default skill is converse. Register more with SkillRegistry (name, description, prompt fragment, optional tool subset).
Compiled profiles
Neutral JSON under sallm/profiles/ (instructions + demos + budgets). Offline:
uv run sallm optimize --dataset data/cases.jsonl --task controller --out /tmp/profile.json
sallm chat never optimizes at startup; it only loads a profile.
CLI
# Durable long session (recommended)
uv run sallm chat \
--state-path .sallm/state.db \
--vector-path .sallm/vectors \
--session long1 \
--retrieval-query instruct \
--search dense \
--memory-gate \
--extract waterfall \
--tools echo,calc
uv run sallm chat --show-prompt
uv run sallm chat --script tests/fixtures/sample_conversation.txt
Slash commands: /help, /clear, /history, /prompt, /state, /stack, /memory, /context, /quit.
Tool contract
| Rule | Detail |
|---|---|
| Identity | Tool name = first argv token |
| Help | Every tool supports --help |
| Args | CLI flags only (no JSON blobs) |
| Intermediate | stdout may start with [intermediate] |
```run
calc --expression "2**10"
```
Shipped tools: echo, calc, dig.
Legacy context optimizers
Still available without durable state: --context max-messages|summarize. Prefer --state-path + retrieval for hour-scale sessions.
Tests
uv run pytest tests/ -v
# E2E needs Ollama + gemma4:e4b-it-qat (+ qwen3-embedding:0.6b for stack memory)
Limits
Retrieval improves grounding; it does not guarantee the model never invents facts. Source-tagged memory and ContextReceipt make misses inspectable.
How it works
Walkthrough with an example, stage-by-stage flow, token-budget simulation, and hypothesis checks: docs/how-the-agent-works.md.
When a script turn (briefing / transcript) is larger than the history budget: docs/oversized-briefings.md.
Offline prompt/parameter tuning: docs/optimize-prompts.md.
Skills (selection, stack, tools): docs/skills.md.
License
Metadata
Release files for sallm-agent 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| sallm_agent-0.2.0.tar.gz | 56.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| sallm_agent-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 130.4 kB
Release files / sallm_agent-0.2.0.tar.gz
| Download URL | sallm_agent-0.2.0.tar.gz |
|---|---|
| Size | 56.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
ddc7b6155deb59caa27f55df786b917c7b6ecc91f94e85324ec5d9e1572eff31
|
|
BLAKE2b-256 checksum How to use checksums |
5570e2a637779387b57efc54c31825d4cac64b95db1f1836f88d31abacb721fd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 10, 2026.
Transparency logRelease files / sallm_agent-0.2.0-py3-none-any.whl
| Download URL | sallm_agent-0.2.0-py3-none-any.whl |
|---|---|
| Size | 74.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
66ca32e4dcb64171e4a9660d6801a88297fe1f7365e27a4d563d83be120324ce
|
|
BLAKE2b-256 checksum How to use checksums |
340132fe8f11337ec4e234e60dc062674d5b717f63b9f5fdea52011c78f1853f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 10, 2026.
Transparency log