Scribe
🌐 English · 简体中文
Local-first chat + research + coding agent. Connects to any OpenAI-compatible server (llama.cpp, Ollama, LM Studio, or a cloud API key) and uses web search, RAG and semantic memory to research, write, and remember — across sessions. Runs comfortably on a 12 GB VRAM machine with Gemma 4 12B at 128k context.
Why Scribe
Three properties that hold by construction, not by prompt-tuning:
- The tool call cannot break. On a llama.cpp server Scribe enforces tool calls with a GBNF grammar generated from the tool schemas — a malformed call is grammatically impossible, and a model that fumbles is re-asked under the grammar. (scribe/grammar.py)
- Answers cite their sources, or say they can't. Grounded Q&A maps every
claim to a numbered source
[n], tags[CONTRADICTION]when sources disagree, and refuses to answer outside the sources. (scribe/prompts.py) - Grounding is measured, not asserted.
scribe-llm benchreports a deterministic Source-Presence Index (SPI) over a checksum-locked held-out suite, andbench --modelsranks whole model fleets on it — see the grounding leaderboard. (scribe/evolve/spi.py). See Gauntlet Soak Test Results for performance under load.
Honest scope: GBNF and constrained decoding aren't new — llama.cpp added grammars in 2023, and the same idea ships elsewhere as "structured outputs". Scribe's contribution is the integration: auto-generating the grammar from your tool schemas and wiring it as an automatic tool-call safety net for small local models. Guarantees 1–3 hold on llama.cpp; other backends degrade to a best-effort text parser. The grammar guarantees a call's form, not the model's judgment (it can still pick the wrong tool — it just can't emit a malformed one).
📖 Full overview on the project site.
Grounding leaderboard
SPI over the 24-task checksum-locked grounded suite — deterministic, no LLM
judge. Every model runs under the same grounded prompt; the ranking is pure
citation discipline. Reproduce with scribe-llm bench --models and a
[scribe.bench] model list; full report in
docs/leaderboard.md.
| Rank | Model | SPI |
|---|---|---|
| 1 | gemma-4-12B (local, llama.cpp) | 0.865 |
| 2 | llama-3.3-70b (Groq) | 0.826 |
| 3 | gpt-oss-120b (Groq) | 0.823 |
| 4 | qwen3-32b (Groq) | 0.807 |
| 5 | gemma-4-E2B (local, llama.cpp) | 0.664 |
| 6 | llama-3.1-8b (Groq) | 0.341 |
Under the same harness, a local 12B out-cites 70B and 120B cloud models — the discipline comes from the harness + model pairing, not model size alone.
Features
- Universal LLM Adapter — llama.cpp, Ollama, LM Studio, or any OpenAI-compatible cloud API (OpenRouter, Groq, ...) — see docs/providers.md
- GBNF Tool Enforcement — grammar-constrained tool calls on llama.cpp; auto-repair when a model emits a malformed call
- Grounded Q&A — hybrid retrieval (FTS5 + vectors, RRF) with mandatory citations and contradiction tagging
- Quality Gate —
scribe-llm benchruns a judge-scored fitness suite and the deterministic SPI grounding metric - Safe Code Mode —
/codewith a destructive-command gate, Python AST gate, bubblewrap sandbox, and git checkpoint/rollback - Internet Research —
web_search+web_fetchtools; DuckDuckGo out of the box, Brave with a key - Writing Agent — deep-research and writer skills for books, papers and reports, with sandboxed workspace file tools
- Book Studio (web) — dark, VSCode-style web editor: three resizable panes, an integrated terminal, model-drafted table of contents, chapter-by-chapter writing, and Markdown / EPUB / PDF export
- Open Knowledge Format —
scribe-llm wiki distillcurates sessions into portable OKF markdown (YAML frontmatter +index.md/log.md+ links); SME/RAG are derived indexes over the files - Persistent Self — a WorldModel persona injected into every prompt, plus pulse heartbeat and nightly diary
- Observability — ORORO session traces and a machine-readable
scribe-llm status --jsoncontract - Project Vaults —
scribe-llm initgives a directory its own isolated RAG/SME stores - Model Discovery & Blind Compare — auto-find local servers; A/B two models without bias
- Cross-Session Memory — SME (Semantic Memory Engine) for seamless session continuity
- Email Bridge — get results by email and send commands from one approved address (stdlib only)
- Language games — explicit command vocabularies (Wittgenstein-inspired) for stable LLM behavior
Installation
From PyPI
pip install scribe-llm # provides the `scribe-llm` command (alias: `scb`)
The PyPI distribution is named
scribe-llm; the CLI command isscribe-llm(short aliasscb). The import package staysscribe.
From source (🐧 Linux / 🍎 macOS)
# Clone the repository
git clone https://github.com/pedjaurosevic/scribe-llm.git
cd scribe-llm
# Run the install script: installs the package (editable), creates the
# config, and scaffolds ~/scribe-workspace.
./scripts/install.sh
🪟 Windows (WSL)
scribe-llm chat and the scribe-llm web Book Studio run on native Windows too; only
the web UI's integrated terminal needs a POSIX PTY (it degrades gracefully
without one). For the full experience, run Scribe via WSL (Windows Subsystem
for Linux).
- Open your WSL terminal (e.g., Ubuntu).
- Run the Linux installation commands shown above.
- To configure Scribe to work with Ollama (either running natively on Windows or inside WSL), check out the WSL & Ollama Integration Guide.
Or install by hand (skip the script):
pip install -e .
Requires Python 3.10+. The install is editable, so
git pullupdates Scribe in place.
Quick Start
- Start your llama-server:
./scripts/start-server.sh
- Start Scribe:
scribe-llm chat
- Scribe will automatically recall your last session and ask if you want to continue.
Configuration
Edit ~/.config/scribe/config.toml:
[scribe]
base_url = "http://127.0.0.1:18083/v1" # llama.cpp / Ollama / LM Studio / cloud
model = "default" # auto-detects the loaded model
api_key = "not-needed" # set a real key for cloud providers
Or use environment variables:
export SCRIBE_BASE_URL=http://localhost:11434/v1 # e.g. Ollama
export SCRIBE_MODEL=gemma4:12b
export SCRIBE_API_KEY=sk-... # cloud providers only
Full provider guide — llama.cpp, Ollama, LM Studio, OpenRouter/Groq, plus a recipe for Gemma 4 12B with 128k context on a 12 GB GPU: docs/providers.md.
Web search needs no setup (DuckDuckGo). For Brave Search, set
BRAVE_API_KEY or brave_api_key under [scribe].
CLI Commands
scribe-llm chat # Interactive TUI chat (streaming by default)
scribe-llm chat --textual # Full-screen Textual UI (experimental)
scribe-llm chat --resume TAG # Resume a past session (no TAG = last one)
scribe-llm web # Book Studio web UI at http://localhost:8765 (localhost-only by default)
scribe-llm web --host 0.0.0.0 # Expose on the network (prints a warning — the UI has a shell terminal)
scribe-llm memory recall "query" # Recall from semantic memory
scribe-llm rag search "query" # Hybrid search (FTS5 + vectors); --semantic-only to opt out
scribe-llm rag ask "question" # Grounded Q&A — answers cite sources or refuse
scribe-llm rag reindex # Rebuild the lexical (FTS5) index
scribe-llm session last # Show last session
scribe-llm session list # List all sessions
scribe-llm session search "query" # Full-text search across all session transcripts
scribe-llm init [DIR] # Create a project-local vault (config + ./.scribe)
scribe-llm discover [--tailscale] # Find OpenAI-compatible model servers
scribe-llm compare "q" --a M1 --b M2 # Blind A/B two models on one prompt
scribe-llm bench [--fitness|--spi] # Quality gate: judge fitness + SPI grounding
scribe-llm trace [ID] [--json] # Show a session's ORORO trace
scribe-llm pulse # Record one heartbeat (wire to a systemd timer)
scribe-llm diary # Reflect on today's sessions
scribe-llm remember "fact" # Add a durable fact to the WorldModel
scribe-llm config show # Show current config
scribe-llm status [--json] # System status (--json = machine-readable contract)
scribe-llm evolve eval # Run the held-out fitness suite (Phase 0)
scribe-llm mail send "Subj" "Body" # Email yourself a notification
scribe-llm mail watch # Accept commands by email (see below)
Email bridge
Scribe can email you results and accept commands by email — using only the Python standard library (no extra dependencies).
[scribe.email]
enabled = true
address = "you@gmail.com"
approved_sender = "you@gmail.com" # the ONLY address allowed to command Scribe
secret = "pick-a-token" # must appear in command subjects
export SCRIBE_EMAIL_PASSWORD="your-gmail-app-password" # never in the config file
scribe-llm mail send "Done" "The report is ready." # send a notification
scribe-llm mail watch # poll inbox, run commands, reply
To run a command, email yourself with the secret in the subject:
Subject:
[scribe:pick-a-token] summarize the notes in research/
Scribe runs it with the sandboxed workspace tools (read/write/list files, no shell) and replies with the answer.
Security: commands are accepted only when both the sender matches
approved_sender and the subject carries [scribe:secret]. Since From:
headers can be spoofed, the secret is the real gate — leave it empty to keep
command intake off (sending still works). Gmail requires an
App Password (with 2-Step
Verification enabled).
Architecture
┌─────────────────────────────────────────┐
│ SCRIBE TUI │
│ (Rich-based interface) │
├─────────────────────────────────────────┤
│ CORE KERNEL │
│ Session Manager │ Skills │ Config │
├─────────────────────────────────────────┤
│ LLM ADAPTER LAYER │
│ OpenAI-compatible (llama.cpp) │
├─────────────────────────────────────────┤
│ MEMORY LAYER │
│ SME (cross-session) │ RAG (documents) │
├─────────────────────────────────────────┤
│ TOOLS LAYER │
│ web_search │ web_fetch │ bash │
└─────────────────────────────────────────┘
Philosophy
Language games (Wittgenstein-inspired). Each command word has a fixed,
explicit meaning, so the model knows exactly what each action means. Reasoning,
when enabled, stays in a <think> block and never leaks into the answer; by
default Scribe answers directly.
License
MIT
Release files for scribe-llm 3.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| scribe_llm-3.0.0.tar.gz | 287.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| scribe_llm-3.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 506.2 kB
Release files / scribe_llm-3.0.0.tar.gz
| Download URL | scribe_llm-3.0.0.tar.gz |
|---|---|
| Size | 287.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8df0adb4c1d236df1a43b3d721eb8c49df97d9604e07ba6b7ef1724680e2bbf1
|
|
BLAKE2b-256 checksum How to use checksums |
c29c4c7282067250cd7fe33842c3794f948e666ad493ea565de359f97442458d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 17, 2026.
Transparency logRelease files / scribe_llm-3.0.0-py3-none-any.whl
| Download URL | scribe_llm-3.0.0-py3-none-any.whl |
|---|---|
| Size | 218.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a1ce000144a81ce52aa03b8190970d81e904de95bbc20d9c7e751fb40d6598c3
|
|
BLAKE2b-256 checksum How to use checksums |
b466d1b8b771fdfa1799b29dfc4ec751a21543bf5f262dcc6a9a59c271b36bdb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.12
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Jul 17, 2026.
Transparency log