Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

llm-scope-observer

Lightweight observability and diagnostics for local and self-hosted LLMs.

I know I messed it up for both of us and I am sorry, she will understand.. you can go ahead and use the library.

llm-scope-observer is a small Python package that wraps your local LLM calls (Ollama, FastAPI backends, OpenWebUI integrations, custom Python code) and records:

  • Latency per call
  • Token usage (input, output, total, tokens/sec)
  • CPU / RAM / (optional) GPU utilization
  • Simple hallucination-risk heuristics
  • Error information

All metrics are stored locally (SQLite by default) and visualized in a small FastAPI-based dashboard.


Features

  • Request interceptor: Decorator to wrap any Python function that calls an LLM.

    from llm_scope import monitor
    import ollama
    
    @monitor(model="llama3")
    def generate(prompt: str) -> str:
        result = ollama.generate(model="llama3", prompt=prompt)
        return result["response"]
    
  • Token estimation:

    • Approximate input and output tokens
    • Track total tokens and tokens/sec per call
  • System metrics snapshot (per request):

    • CPU %
    • RAM %
    • GPU % (optional, via pynvml if installed)
  • Hallucination risk heuristic (simple, signal-based):

    • Very long answer vs. short prompt
    • Strong claims without references
    • Repetition patterns
    • Basic self-contradiction patterns
  • Local dashboard:

    • FastAPI backend + simple HTML UI
    • SQLite storage by default
    • Shows latency, token trends, errors, and resource correlation per model

Installation

pip install llm-scope-observer

Optional GPU metrics:

pip install "llm-scope-observer[gpu]"

Requires Python 3.9+.


Quickstart

1. Instrument your LLM call

from llm_scope import monitor
import time

@monitor(model="test-model")
def generate(prompt: str) -> str:
    time.sleep(0.1)
    return "hello from llm-scope-observer"

Every time generate(...) runs, a record is written to a local SQLite database (llm_scope.db by default).

2. Run the dashboard

After some traffic:

llm-scope ui --host 127.0.0.1 --port 8000
# or
python -m llm_scope.cli ui --host 127.0.0.1 --port 8000

Open:

and you’ll see:

  • Average latency per model
  • Slowest calls (tail latency)
  • Token usage and tokens/sec
  • Error counts
  • CPU / RAM / GPU vs. latency
  • Hallucination score per call

How it works (high level)

  • Middleware / decorator:

    • @monitor(model="llama3") wraps any function.
    • Captures start/end times, prompt, response, and errors.
    • Sends a metrics record to the storage backend.
  • Metrics:

    • Token estimation from prompt and response text.
    • System stats from psutil (and optionally pynvml).
    • Simple heuristics for hallucination risk.
  • Storage:

    • SQLite via sqlite3 by default.
    • One table: llm_calls with timestamps, model, metrics, error, tags.
  • Dashboard:

    • FastAPI app.
    • Reads from the same SQLite file.
    • Renders an HTML summary page (no external JS required).

Roadmap

This is an early MVP. Planned next steps include:

  • Prompt clustering and slow-prompt detection
  • Model A vs. Model B comparison
  • Basic alerting hooks and export to tools like Grafana
  • Optional HTTP ingestion mode (sidecar / agent model)

License

MIT

Metadata

Release files for llm-scope-observer 0.1.0a4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-scope-observer 0.1.0a4
File Size Uploaded
llm_scope_observer-0.1.0a4.tar.gz 10.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-scope-observer 0.1.0a4
File Interpreter ABI Platform
llm_scope_observer-0.1.0a4-py3-none-any.whl Python 3 none any Details

Total release size: 20.9 kB

Release files / llm_scope_observer-0.1.0a4.tar.gz

Download URL llm_scope_observer-0.1.0a4.tar.gz
Size 10.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a02dbb5a9652b82a14d06478f2aef8c94889ac6296c66a8782ba50c01a024aaf
BLAKE2b-256 checksum
How to use checksums
358bb75330fd11848d0f1cb4a7704789f8e16b39627b190a17af6a94bd94fc93
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.2

Release files / llm_scope_observer-0.1.0a4-py3-none-any.whl

Download URL llm_scope_observer-0.1.0a4-py3-none-any.whl
Size 10.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2810e38a533c0ddee8212004dee15564eb4b922c321aff6911a9a577b6aa5996
BLAKE2b-256 checksum
How to use checksums
4b9ad8a30077aaa50868c8fb34ee7c9e63bfc6ccb44f2c114e03edcd1b7d5a9d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.2
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page