freeaiagent — free local AI agent, zero setup
Run powerful AI models fully on-device. No cloud, no cost, no compromise on privacy.
Run a local AI agent with one pip install. No OpenAI key, no cloud required — it downloads and runs a local LLM for you, or connects to free cloud tiers (Groq, Gemini, OpenRouter). Any app calls it over HTTP with no LLM code on its side.
Runs as a persistent HTTP server on localhost:7731. Stores conversation history in SQLite. Any app — script, CLI tool, personal project — can delegate tasks to it with a single HTTP call, no LLM code required on the caller's side.
Built on free LLM backends:
- Local, zero-install — downloads and runs a local GGUF model for you via llamafile (no Ollama, no keys). Supports 1B–14B models.
- SDX — text + vision in one bundle — attach images in any chat; a local vision model reads them and the conversation continues seamlessly. Five tiers from ~3 GB to ~14 GB. See SDX.
- In-process (optional) — load a GGUF straight into Python via
llama-cpp-python, no subprocess or local server. - Ollama — local, no key, if you already use it.
- Free cloud tiers — Groq, Google Gemini, OpenRouter, Together, Cerebras (bring a free API key).
Plus streaming, tool/function calling, per-app sessions, summarizing & per-backend context windows, hybrid RAG retrieval over chat history (SDX), and ensemble inference out of the box.
Call it from Python with the built-in Client SDK (one import, live download progress, no HTTP boilerplate), or point any OpenAI SDK / LangChain / LlamaIndex client at its drop-in /v1 endpoint. Install it as an auto-start service so it's always available.
Architecture
Your apps talk to one local HTTP service — over plain HTTP, the Python Client SDK,
or the OpenAI-compatible /v1 API. freeaiagent owns model routing, persistent context,
the automatic fallback chain, and tool calls — so no app needs any LLM code, keys, or
model management of its own.
Same diagram as text (Mermaid)
flowchart TD
subgraph apps["YOUR APPS"]
A1["Web app<br/>Flask / Django / FastAPI"]
A2["Python app<br/>freeaiagent.Client SDK"]
A3["OpenAI SDK / LangChain<br/>→ /v1 endpoint"]
A4["Browser<br/>built-in Chat UI /ui"]
end
apps -->|"HTTP · SDK · OpenAI /v1"| API
API["HTTP API · localhost:7731<br/>/chat · /task · /chat/stream · /tools · /sessions<br/>/pull/stream · /models · /config · /v1/chat/completions"]
API --> CORE
subgraph CORE["freeaiagent server · FastAPI"]
R["Router<br/>backend + model per call"]
S["Sessions + Context<br/>SQLite, per-app (X-Caller-ID)"]
F["Fallback chain<br/>local → ollama → cloud"]
T["Tool loop<br/>calls your HTTP tools"]
end
CORE -->|"whichever backend is available"| B1 & B2 & B3
B1["Local model (default)<br/>llamafile + GGUF · 1B–14B<br/>zero install · offline · no key"]
B2["Ollama (optional)<br/>if you already run it"]
B3["Free cloud tiers<br/>Groq · Gemini · OpenRouter<br/>Together · Cerebras"]
DB[("Local storage ~/.freeaiagent/<br/>config.json · context.db · models/ · engine/")]
CORE -.-> DB
Why
Embedding LLM logic directly into every app that needs it means duplicating prompt management, context handling, model configuration, and install detection across projects. When something changes — a new model, a different provider, a context bug — you fix it in every app separately.
freeaiagent is the single place that owns all of that. Your apps just call an endpoint.
Install
pip install freeaiagent
Requires Python 3.10+. That's everything — local models and all cloud presets work with no extra packages. (For the optional in-process backend: pip install "freeaiagent[llama-cpp]".) Then either pull a local model or add a free key (below).
freeaiagent pull # one-time local model download (~2.3 GB), then fully offline
freeaiagent start --open # starts the server and opens the Chat UI in your browser
Backends & free keys
No budget needed. Every option below is free.
Local model — zero install, no key, fully offline (default)
A self-contained model the agent downloads and runs for you. No Ollama, no service.
freeaiagent pull # one-time ~2.3 GB download (Llama-3.2-3B), with a progress bar
freeaiagent start --open # starts the server and opens the Chat UI in your browser
Pick a different size or browse the catalog:
freeaiagent models --available # see all local models (1B → 14B)
freeaiagent pull qwen2.5-7b # a stronger model (also fetches a one-time engine)
freeaiagent config set default_model qwen2.5-7b
Ollama — local, no key, no internet
Runs entirely on your machine. Best if you already use Ollama.
- Download from ollama.com
- Pull a model:
ollama pull llama3.2:3b freeaiagent config set default_backend ollama && freeaiagent start
Groq — fastest free cloud inference
No credit card. ~1000 requests/day free.
- Sign up at console.groq.com → API Keys → Create
- Wire it in:
freeaiagent config set backends.groq.api_key gsk_...
freeaiagent config set default_backend groq
freeaiagent config set default_model openai/gpt-oss-20b
Cloud presets — Gemini, OpenRouter, Together, Cerebras
These are built in as presets. Just add a key, pick the backend, pick a model — no base_url/type wiring needed.
Google Gemini — 1500 free requests/day. Key: aistudio.google.com/apikey
freeaiagent config set backends.gemini.api_key AIza...
freeaiagent config set default_backend gemini
freeaiagent config set default_model gemini-2.0-flash
Free models: gemini-2.0-flash, gemini-2.0-flash-lite, gemini-1.5-flash, gemini-1.5-flash-8b
OpenRouter — many :free models. Key: openrouter.ai
freeaiagent config set backends.openrouter.api_key sk-or-...
freeaiagent config set default_backend openrouter
freeaiagent config set default_model meta-llama/llama-3.1-8b-instruct:free
Together AI — free tier. Key: api.together.xyz
freeaiagent config set backends.together.api_key ...
freeaiagent config set default_backend together
freeaiagent config set default_model meta-llama/Llama-3.3-70B-Instruct-Turbo-Free
Cerebras — free tier, very fast. Key: cloud.cerebras.ai
freeaiagent config set backends.cerebras.api_key csk-...
freeaiagent config set default_backend cerebras
freeaiagent config set default_model llama-3.3-70b
LM Studio / Jan / LocalAI — local GUI apps, no key
These run models locally and expose an OpenAI-compatible server.
| App | Download | Default port |
|---|---|---|
| LM Studio | lmstudio.ai | 1234 |
| Jan | jan.ai | 1337 |
| LocalAI | localai.io | 8080 |
# Example: LM Studio running on default port
freeaiagent config set backends.lmstudio.type openai_compat
freeaiagent config set backends.lmstudio.base_url http://localhost:1234
freeaiagent config set default_backend lmstudio
No keys? Run
freeaiagent keysanytime to see this setup guide.
Local models
The local backend is self-contained — no Ollama, no Python ML deps. It downloads a model and runs it as a local OpenAI-compatible server.
freeaiagent models --available # browse the catalog
* llama-3.2-1b 1.3 GB RAM>= 2GB [low ] fused Fastest. Classify / extract / tag.
gemma-2-2b 2.0 GB RAM>= 4GB [mid ] fused Concise; strong at short summaries.
llama-3.2-3b 2.3 GB RAM>= 4GB [mid ] fused Balanced default. Light reasoning / Q&A.
phi-3-mini 2.4 GB RAM>= 4GB [mid ] fused Strong reasoning per byte.
qwen2.5-7b 4.7 GB RAM>= 8GB [high] engine Strong reasoning, Q&A, summaries.
llama-3.1-8b 4.9 GB RAM>= 8GB [high] engine Broad capability, 128k context.
qwen2.5-14b 9.0 GB RAM>=16GB [max ] engine Strongest local option.
- fused models are a single self-contained file (kept under 4 GB so they run on Windows).
- engine models are external GGUF weights run via a one-time ~305 MB llamafile engine — this is how 7B–14B models run anywhere, including Windows.
freeaiagent pull llama-3.2-3b # a catalog model by name
freeaiagent pull qwen2.5-7b # also fetches the shared engine, once
freeaiagent config set default_model qwen2.5-7b
freeaiagent rm qwen2.5-7b # delete a downloaded model to free disk
Interrupted downloads resume automatically — re-running pull continues a
partial .part file via an HTTP range request instead of starting over. Catalog
entries may pin a SHA256 checksum, which is verified after download.
Any model from HuggingFace
Search the whole Hub for GGUF models and pull any of them — no key for public repos.
freeaiagent search qwen2.5 # find GGUF repos (most-downloaded first)
freeaiagent search bartowski/Qwen2.5-7B-Instruct-GGUF # list that repo's GGUF files + sizes
freeaiagent pull hf:bartowski/Qwen2.5-7B-Instruct-GGUF/Qwen2.5-7B-Instruct-Q4_K_M.gguf
Models download to ~/.freeaiagent/models/, the engine to ~/.freeaiagent/engine/.
SDX — chat with images, fully local
SDX (Smart Decision eXecution) is a compound engine that bundles a text model + vision model behind one model name. Attach an image in the Chat UI (or pass image on /chat) and SDX reads it locally; text-only turns skip the vision model entirely. One pull fetches the whole bundle — nothing ever leaves your machine.
freeaiagent pull sdx-standard # one bundle: text + vision + embedder
freeaiagent config set default_backend sdx
freeaiagent start --open # attach images right in the Chat UI
| Tier | Text model | Bundle | Min RAM |
|---|---|---|---|
sdx-nano |
Qwen2.5-0.5B | ~3.3 GB | 6 GB |
sdx-mini |
Llama-3.2-1B | ~3.7 GB | 6 GB |
sdx-standard |
Llama-3.2-3B | ~4.9 GB | 8 GB |
sdx-plus |
Qwen2.5-7B | ~9.7 GB | 16 GB |
sdx-max |
Qwen2.5-14B | ~13.7 GB | 24 GB |
Pick a tier with freeaiagent config set backends.sdx.model sdx-plus. Each tier uses its model's native chat template and real tokenizer-based context budgeting. Switching tiers automatically frees the previous one's RAM; the vision model unloads after 5 idle minutes.
How images work: the vision model extracts a description that is stored in the session as text ([SDX-Image]: ...). That means image content stays part of the conversation permanently — later turns (and retrieval, below) can refer back to a chart or screenshot long after the image itself is gone.
Hybrid RAG over your chat history (SDX-only)
Long sessions forget their beginnings under any sliding window. With context_strategy: rag, SDX embeds every exchange locally (shared ~84 MB nomic embedder, cached in SQLite) and blends the most relevant earlier exchanges with the recent window — clearly fenced as memory, so the model never sees a scrambled dialogue. Because images become text, pictures are retrievable too ("what was that error screenshot from earlier?").
freeaiagent config set backends.sdx.context_strategy rag
Every /chat response reports what actually ran — "sdx": {"strategy_used": "rag"} — and RAG silently falls back to the recency window if anything is missing, so it can never break a chat. Tune with backends.sdx.rag.top_k, .min_similarity, and .recent_pairs.
Quick start
1. Start the server
freeaiagent start
# Running at http://localhost:7731
# API docs at http://localhost:7731/docs
2. Chat with it
freeaiagent chat
# You: what is the capital of France?
# Agent [llama-3.2-3b]: Paris.
3. Call it from any app
import urllib.request, json
req = urllib.request.Request(
"http://localhost:7731/chat",
data=json.dumps({"message": "summarize this for me: ..."}).encode(),
headers={"Content-Type": "application/json"},
)
response = json.loads(urllib.request.urlopen(req).read())["response"]
No pip dependencies needed in the calling app. Pure stdlib.
Chat UI
A browser-based chat interface — similar to ChatGPT, no signup, runs entirely on your machine.
Open it:
freeaiagent start # then visit http://localhost:7731/ui
┌──────────────┬──────────────────────────────────────────────┐
│ + New Chat │ Model: [openai/gpt-oss-20b ▾] │
├──────────────┤──────────────────────────────────────────────┤
│ Chat 1 │ │
│ Chat 2 │ [assistant response] │
│ Chat 3 ··· │ [your message] │
│ │ [assistant response] │
│ ├──────────────────────────────────────────────┤
│ │ [ Type a message… ] [↑] │
└──────────────┴──────────────────────────────────────────────┘
- Sessions sidebar — all your chats, ordered by most recent. Click to switch.
- New Chat — starts a fresh session instantly, no naming required.
- Model picker — dropdown populated live from your configured backends. Change mid-chat.
- Rename — double-click any session title to edit it inline.
- Delete — hover a session and click
×. - Keyboard — Enter to send, Shift+Enter for a newline.
- Images — on an SDX model, attach an image to any message; you'll see an "Analyzing image…" step, then the reply (see SDX).
Sessions and history are stored in ~/.freeaiagent/context.db and persist across restarts.
CLI
freeaiagent start # start server (default port 7731)
freeaiagent start --port 8080 # custom port
freeaiagent start --reload # dev mode: auto-reload on code change
freeaiagent chat # interactive chat (default session)
freeaiagent chat "quick one-liner" # single message, then exit
freeaiagent chat --session work # chat in a named session
freeaiagent chat --session work "hi" # single message to named session
freeaiagent sessions # list all sessions
freeaiagent task "explain this" \
--input "$(cat file.py)" # one-shot task, no context read/written
freeaiagent task "translate to French" \
--input "Hello world" \
--model mistral:7b # override model for this task
freeaiagent status # health check + active backend/model
freeaiagent models # list models on the active backend
freeaiagent models --available # browse the local model catalog
freeaiagent keys # show where to get free API keys
freeaiagent pull # download the default local model
freeaiagent pull qwen2.5-7b # download a catalog model by name
freeaiagent pull sdx-standard # download an SDX text+vision bundle
freeaiagent pull hf:owner/repo/file.gguf # download any GGUF from HuggingFace
freeaiagent rm qwen2.5-7b # delete a downloaded model (frees disk)
freeaiagent search qwen2.5 # find GGUF models on HuggingFace
freeaiagent search owner/repo # list a repo's GGUF files
freeaiagent context show # print conversation history (default session)
freeaiagent context show --session work # print history for a named session
freeaiagent context clear # wipe default session
freeaiagent context clear --session work # wipe a named session
freeaiagent config show # print current config
freeaiagent config set default_model mistral:7b
freeaiagent config set default_backend groq
freeaiagent config set backends.groq.api_key gsk_...
freeaiagent install # auto-start at login (no admin)
freeaiagent service status # is the auto-start service installed?
freeaiagent uninstall # remove the auto-start service
HTTP API
All endpoints accept and return JSON.
POST /chat
Send a message. Conversation history is read and updated automatically per session.
curl -X POST http://localhost:7731/chat \
-H "Content-Type: application/json" \
-d '{"message": "what did I just ask you?", "session_id": "work"}'
{
"response": "You asked me what you just asked me.",
"model": "llama3.2:3b",
"session_id": "work",
"context_length": 4
}
| Field | Type | Description |
|---|---|---|
message |
string | required |
session_id |
string | session to read/write context for (default: "default") |
system |
string | optional system prompt override for this message |
model |
string | optional model override for this message |
backend |
string | optional backend override for this message |
tools |
bool | optional — let the model call registered tools (default false) |
max_messages |
int | optional — context-window override for this message (see Context strategies) |
ensemble |
bool | string[] | optional — fan out to multiple models and judge the best (see Ensemble) |
image |
string | optional — base64 or data: URI image for SDX vision backends (see SDX) |
On an SDX backend the response also includes "sdx": {"strategy_used": "rag" | "recency"} —
which context strategy actually built this turn.
Auto sessions: if you don't pass session_id, the session is taken from the
X-Caller-ID request header — so an app can set one header and get its own
context thread automatically. Resolution order: body session_id → X-Caller-ID → "default".
POST /chat/stream
Same as /chat, but streams the reply token-by-token as Server-Sent Events. The
full response is persisted to the session when the stream finishes.
curl -N -X POST http://localhost:7731/chat/stream \
-H "Content-Type: application/json" \
-d '{"message": "write a haiku about the sea"}'
data: {"token": "Waves"}
data: {"token": " roll"}
data: {"token": " in"}
data: [DONE]
POST /task
One-shot task. No context is read or written — clean slate every time.
curl -X POST http://localhost:7731/task \
-H "Content-Type: application/json" \
-d '{"task": "list all TODO comments", "input": "..."}'
{
"result": "Line 42: TODO fix this\nLine 87: TODO add tests",
"model": "llama3.2:3b"
}
| Field | Type | Description |
|---|---|---|
task |
string | required — the instruction |
input |
string | optional — content to work on |
model |
string | optional — override model for this call |
system |
string | optional — override system prompt |
GET /context?session=default
Returns the conversation history for a session.
{
"messages": [
{"role": "user", "content": "hello", "timestamp": "2026-06-21T10:00:00+00:00"},
{"role": "assistant", "content": "hi there", "timestamp": "2026-06-21T10:00:01+00:00"}
],
"total": 2,
"session_id": "default"
}
DELETE /context?session=default
Clears conversation history for a session.
{"cleared": 4, "message": "Cleared 4 messages.", "session_id": "default"}
GET /sessions
Lists all sessions ordered by most recent activity.
{
"sessions": [
{
"id": "work",
"title": "Explain the difference between...",
"model": "openai/gpt-oss-20b",
"message_count": 14,
"last_updated": "2026-06-21T10:45:00Z"
}
]
}
POST /sessions
Create a session explicitly (also auto-created on first /chat message).
{"id": "my-project", "title": "My Project"}
PATCH /sessions/{id}
Rename a session: {"title": "New name"}
DELETE /sessions/{id}
Delete a session and all its messages.
GET /health
{"status": "ok", "active_backend": "ollama", "default_model": "llama3.2:3b"}
Returns "status": "degraded" with an "error" field if no backend is reachable.
GET /models
{"models": ["llama3.2:3b", "mistral:7b", "phi3:mini"]}
Tools / function calling
Register an HTTP tool once; the model may call it during a /chat with tools: true.
When the model calls a tool, freeaiagent POSTs the arguments to your tool's
endpoint and feeds the result back, looping until the model produces an answer.
# Register a tool
curl -X POST http://localhost:7731/tools/register \
-H "Content-Type: application/json" \
-d '{
"name": "get_weather",
"description": "Get the current weather for a city",
"endpoint": "http://localhost:9000/weather",
"parameters": {"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]}
}'
# Use it
curl -X POST http://localhost:7731/chat \
-H "Content-Type: application/json" \
-d '{"message": "what is the weather in Paris?", "tools": true}'
| Endpoint | Description |
|---|---|
POST /tools/register |
Register a tool (name, description, endpoint, optional parameters) |
GET /tools |
List registered tools |
DELETE /tools/{name} |
Unregister a tool |
Tool calling needs a backend/model that supports the OpenAI tool protocol (Groq, most hosted providers, some local models). Backends without support answer normally instead.
Models & config management
| Endpoint | Description |
|---|---|
GET /models/catalog |
Curated catalog with an installed flag per entry |
GET /models/installed |
Local model files on disk (name, path, size, kind) |
DELETE /models/installed/{name} |
Delete a downloaded model (catalog name or filename); frees disk |
GET /config |
Effective configuration as JSON |
POST /config/set |
Set a dotted key — {"key": "default_backend", "value": "groq"} |
POST /pull/stream
Server-side model download, streamed as SSE so a UI can show a real progress bar.
GGUF models emit an engine phase (the one-time shared runtime) before model.
One download at a time (409 otherwise); an unknown model is a 400.
data: {"type": "start", "phase": "model", "label": "qwen2.5-7b", "total_mb": 4700}
data: {"type": "progress", "phase": "model", "pct": 12, "downloaded_mb": 564, "speed_mbps": 8.1}
data: {"type": "done", "path": "~/.freeaiagent/models/Qwen2.5-7B-Instruct-Q4_K_M.gguf"}
data: [DONE]
OpenAI-compatible (/v1)
| Endpoint | Description |
|---|---|
POST /v1/chat/completions |
OpenAI chat completions (supports stream: true) |
GET /v1/models |
OpenAI-shaped model list |
Point any OpenAI SDK / LangChain / LlamaIndex client at http://localhost:7731/v1.
Configuration
Config lives at ~/.freeaiagent/config.json and is created on first run.
{
"default_backend": "llamafile",
"default_model": "llama-3.2-3b",
"port": 7731,
"max_messages": 0,
"context_strategy": "sliding",
"summarize_threshold": 40,
"summarize_batch": 30,
"summarize_model": null,
"backends": {
"llamafile": {"type": "llamafile", "port": 8080, "auto_download": false},
"llama_cpp": {"type": "llama_cpp", "n_ctx": 4096, "n_gpu_layers": 0},
"sdx": {"type": "sdx", "model": "sdx-standard", "n_gpu_layers": 0,
"context_strategy": "recency",
"rag": {"top_k": 6, "min_similarity": 0.30, "recent_pairs": 4},
"keep_loaded": false, "vision_idle_unload_s": 300},
"ollama": {"base_url": "http://localhost:11434"},
"groq": {"api_key": ""},
"together": {"type": "openai_compat", "base_url": "https://api.together.xyz", "api_key": ""},
"openrouter": {"type": "openai_compat", "base_url": "https://openrouter.ai/api", "api_key": ""},
"cerebras": {"type": "openai_compat", "base_url": "https://api.cerebras.ai", "api_key": ""},
"gemini": {"type": "openai_compat",
"base_url": "https://generativelanguage.googleapis.com/v1beta/openai",
"api_prefix": "", "api_key": ""}
},
"ensemble": {"enabled": false, "models": [], "judge": null, "strategy": "llm_judge"},
"fallback_order": ["llamafile", "ollama", "groq"]
}
Each backend may also set its own max_messages (overriding the global one),
e.g. a short window for an 8k model and a long one for a 128k model. See
Context strategies & ensemble.
Cloud presets are inert until you set their api_key. default_model for the
local backend is a catalog name (freeaiagent models --available); legacy
values are migrated automatically.
Edit directly or use freeaiagent config set <key> <value> with dotted keys:
freeaiagent config set port 8080
freeaiagent config set backends.ollama.base_url http://192.168.1.10:11434
freeaiagent config set backends.groq.api_key gsk_abc123
freeaiagent config set default_backend groq
freeaiagent config set default_model llama-3.1-8b-instant
Context strategies & ensemble
Context window
By default the agent uses a sliding window: only the last max_messages
messages are sent to the model (0 = unlimited). The effective window resolves
in order: per-call max_messages (on /chat) → backends.<name>.max_messages →
global max_messages.
freeaiagent config set max_messages 20 # global window
freeaiagent config set backends.groq.max_messages 100 # per-backend override
# or per request:
curl -X POST http://localhost:7731/chat -H "Content-Type: application/json" \
-d '{"message": "hi", "max_messages": 10}'
Summarization
Instead of dropping old messages, summarize them: once a session grows past
summarize_threshold, the oldest summarize_batch messages are folded into a
single system "memory" block — so long sessions keep their early context
(decisions, constraints, names) at the cost of one extra LLM call per fold.
freeaiagent config set context_strategy summarize
freeaiagent config set summarize_threshold 40 # start summarizing past 40 messages
freeaiagent config set summarize_batch 30 # fold the oldest 30 into one summary
freeaiagent config set summarize_model llama-3.2-3b # optional; defaults to the chat model
Ensemble inference
Send the same prompt to several models in parallel and return the best answer — catches per-model blind spots for high-stakes one-shot questions (costs N× tokens, ~max-of-N latency). All ensemble models run on the active backend, so they must be models it can serve.
# Per request — pass an explicit model list:
curl -X POST http://localhost:7731/chat -H "Content-Type: application/json" \
-d '{"message": "explain X", "ensemble": ["llama-3.1-8b-instant", "llama-3.3-70b-versatile"]}'
# Or configure it once, then send "ensemble": true (or set enabled):
freeaiagent config set ensemble.models '["llama-3.1-8b-instant", "llama-3.3-70b-versatile"]'
freeaiagent config set ensemble.strategy llm_judge # llm_judge | longest | majority
freeaiagent config set ensemble.enabled true
The response includes an ensemble_votes array (each model's answer, or its
error) alongside the winning response. Judge strategies: llm_judge (a
small model picks the best; falls back to longest on failure), longest
(longest answer, penalised for repetition), majority (most common answer).
Failed models are dropped; only an all-failed ensemble errors.
Backend reference
Local (llamafile) — default
Self-contained. No Ollama, no key, no data leaves your machine. Downloads a model once and runs it locally. See Local models for the catalog, GGUF/engine models, and HuggingFace search.
freeaiagent pull # download the default local model
freeaiagent start
In-process (llama-cpp-python) — optional
Loads a GGUF model directly into the Python process — no llamafile
subprocess, no local HTTP hop. Useful if you want a pure-Python path or finer
control over n_ctx / GPU offload. Reuses the same GGUF weights freeaiagent
pull downloads.
pip install "freeaiagent[llama-cpp]" # optional extra (compiles llama-cpp-python)
freeaiagent pull qwen2.5-7b # any GGUF (engine-mode) catalog model
freeaiagent config set default_backend llama_cpp
freeaiagent config set backends.llama_cpp.n_gpu_layers 35 # offload to GPU (0 = CPU only)
Inert until both the package and a GGUF model are present, so it never interferes with the default setup.
Ollama
Runs locally. No API key. No data leaves your machine.
# Install Ollama from https://ollama.com
ollama pull llama3.2:3b
freeaiagent start
Any model available on your Ollama instance works. Switch with:
freeaiagent config set default_model mistral:7b
Groq (free-tier cloud)
Fast inference. Free API key at console.groq.com.
freeaiagent config set backends.groq.api_key gsk_...
freeaiagent config set default_backend groq
freeaiagent config set default_model llama-3.1-8b-instant
Models are fetched live from the Groq API when your key is set — the list below is the fallback.
| Model | Context | Notes |
|---|---|---|
openai/gpt-oss-20b |
131k | Current production — replaces llama-3.1-8b |
openai/gpt-oss-120b |
131k | Current production — most capable |
groq/compound |
131k | Agentic system with web search + code execution |
groq/compound-mini |
131k | Lighter compound system |
llama-3.1-8b-instant |
131k | Deprecated Aug 16 2026 |
llama-3.3-70b-versatile |
131k | Deprecated Aug 16 2026 |
qwen/qwen3.6-27b |
131k | Preview |
qwen/qwen3-32b |
131k | Preview — deprecated Jul 17 2026 |
meta-llama/llama-4-scout-17b-16e-instruct |
131k | Preview — deprecated Jul 17 2026 |
Cloud presets
together, openrouter, cerebras, and gemini are built-in OpenAI-compatible
presets — set an api_key and select the backend (see Cloud presets). Add any other OpenAI-compatible
server (LM Studio, Jan, LocalAI) with type: openai_compat and a base_url.
Automatic fallback
If the default backend is unreachable, freeaiagent tries the next one in
fallback_order automatically (default: llamafile → ollama → groq). So if a
cloud key is exhausted or you're offline, the local model still answers.
Using from another app
The agent runs as a separate process. Your app calls it over HTTP — no LLM dependencies, no model management, no context handling in your code.
Python SDK (recommended)
pip install freeaiagent ships a synchronous Client with full server parity and zero HTTP boilerplate. Streaming and downloads are plain iterators — no async.
from freeaiagent import Client
# name → X-Caller-ID, so this app gets its own context thread.
# auto_start=True launches the server if it isn't already running.
agent = Client(name="my-app", auto_start=True)
# Chat (context preserved per session)
print(agent.chat("hello"))
print(agent.chat("write a tagline", session="marketing", model="qwen2.5-7b"))
# Stream tokens
for token in agent.stream("write a haiku about caching"):
print(token, end="", flush=True)
# One-shot task (no shared context)
print(agent.task("summarize concisely", input=long_text))
# Download a model with live progress (for a real progress bar in your UI)
for p in agent.pull("qwen2.5-7b"):
if p.type == "progress":
print(f"[{p.phase}] {p.pct:.0f}% {p.downloaded_mb:.0f}/{p.total_mb:.0f} MB")
elif p.type == "done":
print("Saved to", p.path)
# Namespaced management — mirrors the CLI
agent.models.catalog() # downloadable models, each flagged installed
agent.models.installed() # local files on disk
agent.sessions.list()
agent.context.get(session="marketing")
agent.config.set("default_backend", "groq")
agent.tools.register("get_weather", description="Weather for a city",
endpoint="http://localhost:9000/weather",
parameters={"type": "object", "properties": {"city": {"type": "string"}}})
agent.is_running() # fast health check
The client auto-discovers the running port via ~/.freeaiagent/server.json, so apps survive port changes with no config. Errors are typed: ServerNotRunning, BackendUnavailable, DownloadInProgress.
Install once so it's always up and use Client(auto_start=False):
freeaiagent install # auto-start at login (no admin); Linux/macOS/Windows
freeaiagent service status
freeaiagent uninstall
Drop-in OpenAI API
freeaiagent speaks the OpenAI wire protocol, so anything built on the OpenAI SDK, LangChain, or LlamaIndex works unchanged — just point base_url at it:
from openai import OpenAI
llm = OpenAI(base_url="http://localhost:7731/v1", api_key="none")
llm.chat.completions.create(model="qwen2.5-7b", messages=[{"role": "user", "content": "hi"}])
Raw HTTP
If you'd rather not add a dependency, call the endpoints directly.
Python (stdlib only):
import urllib.request, json
def ask(message):
body = json.dumps({"message": message}).encode()
req = urllib.request.Request(
"http://localhost:7731/chat",
data=body,
headers={
"Content-Type": "application/json",
"X-Caller-ID": "my-app", # this app gets its own context thread
},
)
return json.loads(urllib.request.urlopen(req).read())["response"]
def run_task(task, input_text=None):
body = json.dumps({"task": task, "input": input_text}).encode()
req = urllib.request.Request(
"http://localhost:7731/task",
data=body,
headers={"Content-Type": "application/json"},
)
return json.loads(urllib.request.urlopen(req).read())["result"]
Shell:
curl -s -X POST http://localhost:7731/task \
-H "Content-Type: application/json" \
-d "{\"task\": \"summarize\", \"input\": \"$(cat notes.txt)\"}" \
| python -c "import sys,json; print(json.load(sys.stdin)['result'])"
JavaScript / Node:
const res = await fetch("http://localhost:7731/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ message: "hello" }),
});
const { response } = await res.json();
Context storage
Conversation history is stored in ~/.freeaiagent/context.db (SQLite). It persists across server restarts.
Each session has its own history. Clear a session any time:
freeaiagent context clear # clear the default session
freeaiagent context clear --session work # clear a named session
# or via HTTP:
curl -X DELETE "http://localhost:7731/context?session=work"
curl -X DELETE "http://localhost:7731/sessions/work" # also deletes session record
Roadmap
Done
- Zero-install local backend (llamafile) with a model catalog, engine mode for 7B–14B GGUF models, and HuggingFace search/pull
- Ollama, Groq, and OpenAI-compatible cloud presets (Gemini, OpenRouter, Together, Cerebras; plus LM Studio, Jan, LocalAI)
- Per-call model and backend overrides
- Sliding window context (
max_messages) + per-backend & per-call context windows - Summarization context strategy (fold old messages into a memory block)
- Auto caller detection (
X-Caller-IDheader → per-app session) - Streaming responses (
/chat/streamSSE) - Tool use / function calling (
/tools,tools=true) - Chat web UI at
localhost:7731/ui - Python SDK (
freeaiagent.Client) with livepullprogress and port auto-discovery - Server-side streaming downloads (
/pull/stream), model/config management endpoints - OpenAI-compatible
/v1/chat/completions+/v1/modelsproxy - Auto-start service install (
freeaiagent install) on Linux/macOS/Windows - Ensemble inference (fan out the same query to multiple models, pick the best output)
- Optional in-process engine (
llama-cpp-python) - Resumable downloads + SHA256 checksum verification,
freeaiagent rm - SDX compound engine — text + vision in one local bundle, 5 tiers (Nano → Max), image chat in the UI and API
- SDX v2 — native chat templates per model, real tokenizer-based context budgets, engine manager with tier eviction and vision idle-unload
- Hybrid RAG over chat history (SDX) — local embeddings, semantic retrieval blended with a recency window, retrievable image content,
strategy_usedresponse metadata - Concurrency hardening — serialized generation per model, safe engine eviction under load, per-call stop
Planned
- Hardware auto-detect (
freeaiagent doctor) — set GPU offload automatically, recommend a model tier - Structured output — JSON-schema-constrained generation (GBNF grammars)
/v1/embeddingsendpoint (OpenAI-compatible, backed by the local embedder)- Document RAG — attach a folder, get indexed + cited answers
- Quality evals per catalog model (text + vision), published honestly
- Custom catalog entries
Development
git clone <repo>
cd freeaiagent
pip install -e ".[dev,groq]"
pytest # unit + integration (no LLM needed)
pytest tests/smoke/ -m smoke -v # smoke tests (requires Ollama running)
License
MIT
Release files for freeaiagent 1.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| freeaiagent-1.3.0.tar.gz | 120.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| freeaiagent-1.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 220.9 kB
Release files / freeaiagent-1.3.0.tar.gz
| Download URL | freeaiagent-1.3.0.tar.gz |
|---|---|
| Size | 120.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
30ec3d92fbad77313818a82411b1bc41c63a9552151300e384e175f23870ccd1
|
|
BLAKE2b-256 checksum How to use checksums |
432b0c26d604ce71b18c6720c6a83e9c1d3498686cb01dc02f0504ad973a9211
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.5
|
Release files / freeaiagent-1.3.0-py3-none-any.whl
| Download URL | freeaiagent-1.3.0-py3-none-any.whl |
|---|---|
| Size | 100.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
78e747decc094b4122793a36b94fec0163c0b64b09d57be8b93a1f71846e1724
|
|
BLAKE2b-256 checksum How to use checksums |
37a93d583341192991be745afc718aff083f590b19c29446cf5756666cb16862
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.5
|