Local AI agent service — HTTP endpoints, persistent context, multi-model LLM backend
Project description
freeaiagent
A local AI agent service you pip install once and call from anywhere.
Runs as a persistent HTTP server on localhost:7731. Stores conversation history in SQLite. Any app — script, CLI tool, personal project — can delegate tasks to it with a single HTTP call, no LLM code required on the caller's side.
Built on free LLM backends: Ollama (local, no API key) and Groq (free-tier cloud).
Why
Embedding LLM logic directly into every app that needs it means duplicating prompt management, context handling, model configuration, and install detection across projects. When something changes — a new model, a different provider, a context bug — you fix it in every app separately.
freeaiagent is the single place that owns all of that. Your apps just call an endpoint.
Install
pip install freeaiagent # Ollama only
pip install "freeaiagent[groq]" # + Groq support
Requires Python 3.10+.
Quick start
1. Start the server
freeaiagent start
# Running at http://localhost:7731
# API docs at http://localhost:7731/docs
2. Chat with it
freeaiagent chat
# You: what is the capital of France?
# Agent [llama3.2:3b]: Paris.
3. Call it from any app
import urllib.request, json
req = urllib.request.Request(
"http://localhost:7731/chat",
data=json.dumps({"message": "summarize this for me: ..."}).encode(),
headers={"Content-Type": "application/json"},
)
response = json.loads(urllib.request.urlopen(req).read())["response"]
No pip dependencies needed in the calling app. Pure stdlib.
CLI
freeaiagent start # start server (default port 7731)
freeaiagent start --port 8080 # custom port
freeaiagent start --reload # dev mode: auto-reload on code change
freeaiagent chat # interactive chat (context preserved)
freeaiagent chat "quick one-liner" # single message, then exit
freeaiagent task "explain this" \
--input "$(cat file.py)" # one-shot task, no context read/written
freeaiagent task "translate to French" \
--input "Hello world" \
--model mistral:7b # override model for this task
freeaiagent status # health check + active backend/model
freeaiagent models # list models on active backend
freeaiagent context show # print conversation history
freeaiagent context clear # wipe conversation history
freeaiagent config show # print current config
freeaiagent config set default_model mistral:7b
freeaiagent config set default_backend groq
freeaiagent config set backends.groq.api_key gsk_...
HTTP API
All endpoints accept and return JSON.
POST /chat
Send a message. Conversation history is read and updated automatically.
curl -X POST http://localhost:7731/chat \
-H "Content-Type: application/json" \
-d '{"message": "what did I just ask you?"}'
{
"response": "You asked me what you just asked me.",
"model": "llama3.2:3b",
"context_length": 4
}
| Field | Type | Description |
|---|---|---|
message |
string | required |
system |
string | optional system prompt override |
POST /task
One-shot task. No context is read or written — clean slate every time.
curl -X POST http://localhost:7731/task \
-H "Content-Type: application/json" \
-d '{"task": "list all TODO comments", "input": "..."}'
{
"result": "Line 42: TODO fix this\nLine 87: TODO add tests",
"model": "llama3.2:3b"
}
| Field | Type | Description |
|---|---|---|
task |
string | required — the instruction |
input |
string | optional — content to work on |
model |
string | optional — override model for this call |
system |
string | optional — override system prompt |
GET /context
Returns the full conversation history.
{
"messages": [
{"role": "user", "content": "hello", "timestamp": "2026-06-21T10:00:00+00:00"},
{"role": "assistant", "content": "hi there", "timestamp": "2026-06-21T10:00:01+00:00"}
],
"total": 2
}
DELETE /context
Clears all conversation history.
{"cleared": 4, "message": "Cleared 4 messages."}
GET /health
{"status": "ok", "active_backend": "ollama", "default_model": "llama3.2:3b"}
Returns "status": "degraded" with an "error" field if no backend is reachable.
GET /models
{"models": ["llama3.2:3b", "mistral:7b", "phi3:mini"]}
Configuration
Config lives at ~/.freeaiagent/config.json and is created on first run.
{
"default_backend": "ollama",
"default_model": "llama3.2:3b",
"port": 7731,
"backends": {
"ollama": {
"base_url": "http://localhost:11434"
},
"groq": {
"api_key": ""
}
},
"fallback_order": ["ollama", "groq"]
}
Edit directly or use freeaiagent config set <key> <value> with dotted keys:
freeaiagent config set port 8080
freeaiagent config set backends.ollama.base_url http://192.168.1.10:11434
freeaiagent config set backends.groq.api_key gsk_abc123
freeaiagent config set default_backend groq
freeaiagent config set default_model llama-3.1-8b-instant
Backends
Ollama (default)
Runs locally. No API key. No data leaves your machine.
# Install Ollama from https://ollama.com
ollama pull llama3.2:3b
freeaiagent start
Any model available on your Ollama instance works. Switch with:
freeaiagent config set default_model mistral:7b
Groq (free-tier cloud)
Fast inference. Free API key at console.groq.com.
freeaiagent config set backends.groq.api_key gsk_...
freeaiagent config set default_backend groq
freeaiagent config set default_model llama-3.1-8b-instant
Free-tier models available on Groq:
| Model | Context |
|---|---|
llama-3.1-8b-instant |
128k |
llama-3.3-70b-versatile |
128k |
mixtral-8x7b-32768 |
32k |
gemma2-9b-it |
8k |
Automatic fallback
If the default backend is unreachable, freeaiagent tries the next one in fallback_order automatically. No configuration needed for this to work — just have both set up.
Using from another app
The agent runs as a separate process. Your app calls it over HTTP — no LLM dependencies, no model management, no context handling in your code.
Python (stdlib only):
import urllib.request, json
def ask(message):
body = json.dumps({"message": message}).encode()
req = urllib.request.Request(
"http://localhost:7731/chat",
data=body,
headers={"Content-Type": "application/json"},
)
return json.loads(urllib.request.urlopen(req).read())["response"]
def run_task(task, input_text=None):
body = json.dumps({"task": task, "input": input_text}).encode()
req = urllib.request.Request(
"http://localhost:7731/task",
data=body,
headers={"Content-Type": "application/json"},
)
return json.loads(urllib.request.urlopen(req).read())["result"]
Shell:
curl -s -X POST http://localhost:7731/task \
-H "Content-Type: application/json" \
-d "{\"task\": \"summarize\", \"input\": \"$(cat notes.txt)\"}" \
| python -c "import sys,json; print(json.load(sys.stdin)['result'])"
JavaScript / Node:
const res = await fetch("http://localhost:7731/chat", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ message: "hello" }),
});
const { response } = await res.json();
Context storage
Conversation history is stored in ~/.freeaiagent/context.db (SQLite). It persists across server restarts. Clear it any time:
freeaiagent context clear
# or
curl -X DELETE http://localhost:7731/context
Roadmap
Phase 2 — Named sessions
Multiple apps keep separate conversation histories via a session_id field.
{"message": "hello", "session_id": "my-project"}
Phase 3 — Auto caller detection
Sessions created automatically per caller using X-Caller-ID header or caller port. Zero config for multi-app setups.
Phase 4 — More backends and features
- llamafile (single-file model, no Ollama install needed)
- Together AI, OpenRouter
- Streaming responses (
/chat/streamSSE) - Tool use / function calling
- Minimal web UI at
localhost:7731
Development
git clone <repo>
cd freeaiagent
pip install -e ".[dev,groq]"
pytest # unit + integration (no LLM needed)
pytest tests/smoke/ -m smoke -v # smoke tests (requires Ollama running)
License
MIT
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file freeaiagent-0.1.0.tar.gz.
File metadata
- Download URL: freeaiagent-0.1.0.tar.gz
- Upload date:
- Size: 14.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
ed6b26cab479a759751ea8238d88e1082349b5714aec42ebf0b64ea659f43773
|
|
| MD5 |
8fb538e071bac9cc7dd7157d00f8cd18
|
|
| BLAKE2b-256 |
45e65244d3b191dab2dff6f44dbbc9d9b9b6218d97593599ea56ebd572bd1645
|
File details
Details for the file freeaiagent-0.1.0-py3-none-any.whl.
File metadata
- Download URL: freeaiagent-0.1.0-py3-none-any.whl
- Upload date:
- Size: 14.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.14.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1a0f742a85d196f38858d6ac8506ea80c51e4c7e36e033b98aee5fe5720eca48
|
|
| MD5 |
90aef61c5f022ff7f0fd510e5182d44c
|
|
| BLAKE2b-256 |
bad502b345e1a3bfaa45f7d1bdb75e6ce1e7db45be8b0b9cc7a571e984e29b6a
|