Skip to main content

SALLM logo

sallm-agent

Minimal tool-calling agent (sallm) for local small LLMs (default: Gemma 4 4B via Ollama).

Focus: keep context predictable over long sessions — durable SQLite state, LanceDB retrieval, a skill stack, and a visible ContextReceipt that explains token spend.

Not an invisible RAG black box: raw messages stay canonical; vectors are a rebuildable index; derived facts must cite source message ids.

Setup

ollama pull gemma4:e4b-it-qat
ollama pull qwen3-embedding:0.6b

uv sync --extra dev

Core deps include peewee (SQLite ORM) and lancedb (vector index). There is no DSPy/Pydantic dependency.

Examples

  • examples/imap_inbox/ — durable IMAP inbox Q&A (CLI tools + long-session recall). See the docstring in agent.py.
  • examples/linux_history/ — durable bash/zsh history Q&A + live bash_run (usage-story inference). See the docstring in agent.py.

Turn pipeline

user → persist raw message
  → goal/skill control (small JSON call)
  → vector retrieve (Qwen embed + LanceDB)
  → budgeted prompt + ReAct ```run tools
  → persist answer
  → extract grounded facts + index chunks

Library

from sallm import Agent, RetrievalConfig, Skill, SkillRegistry
from sallm.tools import builtin_tools

agent = Agent(
    tools=builtin_tools(("calc", "echo")),
    state_path="/tmp/sallm/state.db",
    vector_path="/tmp/sallm/vectors",
    session_id="demo",
    retrieval=RetrievalConfig(
        memory_gate=True,
        search_mode="dense",  # or "hybrid"
        use_instruct=True,
        use_rewrite=False,
        use_hyde=False,
    ),
)
result = agent.ask("Remember the code is PURPLE-42.")
print(result["answer"])
print(result["receipt"])  # ContextReceipt as dict
print(result["goal"], result["stack"])

Resume by reusing state_path + session_id.

Preload memory (meaning-first): agent.remember(text, source="…") runs an ingest LLM prompt, stores English facts for retrieval, and does not fill the recent-history window. Use for shell-history blocks and other raw dumps that should answer ordinary-English questions later. See docs/agent-instructions.md.

VectorStore contract

Implement upsert / search / delete_session / close (see sallm.memory.types.VectorStore). Default: LanceVectorStore. A future pgvector adapter can satisfy the same dataclasses (VectorRecord, VectorQuery, VectorHit) without changing the agent.

SQLite stores chunk text + indexed flags; LanceDB is rebuilt from those rows after a crash.

Skills

Default skill is converse. Register more with SkillRegistry (name, description, prompt fragment, optional tool subset).

Compiled profiles

Neutral JSON under sallm/profiles/ (instructions + demos + budgets). Offline:

uv run sallm optimize --dataset data/cases.jsonl --task controller --out /tmp/profile.json

sallm chat never optimizes at startup; it only loads a profile.

CLI

# Durable long session (recommended)
uv run sallm chat \
  --state-path .sallm/state.db \
  --vector-path .sallm/vectors \
  --session long1 \
  --retrieval-query instruct \
  --search dense \
  --memory-gate \
  --extract waterfall \
  --tools echo,calc

uv run sallm chat --show-prompt
uv run sallm chat --script tests/fixtures/sample_conversation.txt

Slash commands: /help, /clear, /history, /prompt, /state, /stack, /memory, /context, /quit.

Tool contract

Rule Detail
Identity Tool name = first argv token
Help Every tool supports --help
Args CLI flags only (no JSON blobs)
Intermediate stdout may start with [intermediate]
```run
calc --expression "2**10"
```

Shipped tools: echo, calc, dig.

Legacy context optimizers

Still available without durable state: --context max-messages|summarize. Prefer --state-path + retrieval for hour-scale sessions.

Tests

uv run pytest tests/ -v
# E2E needs Ollama + gemma4:e4b-it-qat (+ qwen3-embedding:0.6b for stack memory)

Limits

Retrieval improves grounding; it does not guarantee the model never invents facts. Source-tagged memory and ContextReceipt make misses inspectable.

How it works

Walkthrough with an example, stage-by-stage flow, token-budget simulation, and hypothesis checks: docs/how-the-agent-works.md.

When a script turn (briefing / transcript) is larger than the history budget: docs/oversized-briefings.md.

Offline prompt/parameter tuning: docs/optimize-prompts.md.

Skills (selection, stack, tools): docs/skills.md.

License

MIT

Metadata

Release files for sallm-agent 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sallm-agent 0.2.0
File Size Uploaded
sallm_agent-0.2.0.tar.gz 56.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sallm-agent 0.2.0
File Interpreter ABI Platform
sallm_agent-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 130.4 kB

Release files / sallm_agent-0.2.0.tar.gz

Download URL sallm_agent-0.2.0.tar.gz
Size 56.3 kB
Tags Source
SHA-256 checksum
How to use checksums
ddc7b6155deb59caa27f55df786b917c7b6ecc91f94e85324ec5d9e1572eff31
BLAKE2b-256 checksum
How to use checksums
5570e2a637779387b57efc54c31825d4cac64b95db1f1836f88d31abacb721fd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 10, 2026.

Transparency log

Release files / sallm_agent-0.2.0-py3-none-any.whl

Download URL sallm_agent-0.2.0-py3-none-any.whl
Size 74.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
66ca32e4dcb64171e4a9660d6801a88297fe1f7365e27a4d563d83be120324ce
BLAKE2b-256 checksum
How to use checksums
340132fe8f11337ec4e234e60dc062674d5b717f63b9f5fdea52011c78f1853f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 10, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page