This release is a pre-release and may not be stable for production use.
llm-scope-observer
Lightweight observability and diagnostics for local and self-hosted LLMs.
I know I messed it up for both of us and I am sorry, she will understand.. you can go ahead and use the library.
llm-scope-observer is a small Python package that wraps your local LLM calls (Ollama, FastAPI backends, OpenWebUI integrations, custom Python code) and records:
- Latency per call
- Token usage (input, output, total, tokens/sec)
- CPU / RAM / (optional) GPU utilization
- Simple hallucination-risk heuristics
- Error information
All metrics are stored locally (SQLite by default) and visualized in a small FastAPI-based dashboard.
Features
-
Request interceptor: Decorator to wrap any Python function that calls an LLM.
from llm_scope import monitor import ollama @monitor(model="llama3") def generate(prompt: str) -> str: result = ollama.generate(model="llama3", prompt=prompt) return result["response"]
-
Token estimation:
- Approximate input and output tokens
- Track total tokens and tokens/sec per call
-
System metrics snapshot (per request):
- CPU %
- RAM %
- GPU % (optional, via
pynvmlif installed)
-
Hallucination risk heuristic (simple, signal-based):
- Very long answer vs. short prompt
- Strong claims without references
- Repetition patterns
- Basic self-contradiction patterns
-
Local dashboard:
- FastAPI backend + simple HTML UI
- SQLite storage by default
- Shows latency, token trends, errors, and resource correlation per model
Installation
pip install llm-scope-observer
Optional GPU metrics:
pip install "llm-scope-observer[gpu]"
Requires Python 3.9+.
Quickstart
1. Instrument your LLM call
from llm_scope import monitor
import time
@monitor(model="test-model")
def generate(prompt: str) -> str:
time.sleep(0.1)
return "hello from llm-scope-observer"
Every time generate(...) runs, a record is written to a local SQLite database (llm_scope.db by default).
2. Run the dashboard
After some traffic:
llm-scope ui --host 127.0.0.1 --port 8000
# or
python -m llm_scope.cli ui --host 127.0.0.1 --port 8000
Open:
and you’ll see:
- Average latency per model
- Slowest calls (tail latency)
- Token usage and tokens/sec
- Error counts
- CPU / RAM / GPU vs. latency
- Hallucination score per call
How it works (high level)
-
Middleware / decorator:
@monitor(model="llama3")wraps any function.- Captures
start/endtimes, prompt, response, and errors. - Sends a metrics record to the storage backend.
-
Metrics:
- Token estimation from prompt and response text.
- System stats from
psutil(and optionallypynvml). - Simple heuristics for hallucination risk.
-
Storage:
- SQLite via
sqlite3by default. - One table:
llm_callswith timestamps, model, metrics, error, tags.
- SQLite via
-
Dashboard:
- FastAPI app.
- Reads from the same SQLite file.
- Renders an HTML summary page (no external JS required).
Roadmap
This is an early MVP. Planned next steps include:
- Prompt clustering and slow-prompt detection
- Model A vs. Model B comparison
- Basic alerting hooks and export to tools like Grafana
- Optional HTTP ingestion mode (sidecar / agent model)
License
MIT
Metadata
Release files for llm-scope-observer 0.1.0a4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| llm_scope_observer-0.1.0a4.tar.gz | 10.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| llm_scope_observer-0.1.0a4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 20.9 kB
Release files / llm_scope_observer-0.1.0a4.tar.gz
| Download URL | llm_scope_observer-0.1.0a4.tar.gz |
|---|---|
| Size | 10.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a02dbb5a9652b82a14d06478f2aef8c94889ac6296c66a8782ba50c01a024aaf
|
|
BLAKE2b-256 checksum How to use checksums |
358bb75330fd11848d0f1cb4a7704789f8e16b39627b190a17af6a94bd94fc93
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.2
|
Release files / llm_scope_observer-0.1.0a4-py3-none-any.whl
| Download URL | llm_scope_observer-0.1.0a4-py3-none-any.whl |
|---|---|
| Size | 10.6 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2810e38a533c0ddee8212004dee15564eb4b922c321aff6911a9a577b6aa5996
|
|
BLAKE2b-256 checksum How to use checksums |
4b9ad8a30077aaa50868c8fb34ee7c9e63bfc6ccb44f2c114e03edcd1b7d5a9d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.2
|