Skip to main content

llm-speed-web

Source for llm-speed.com — the canonical, crowdsourced source of truth for how fast LLMs actually run, across hosted APIs, consumer GPUs, and prosumer rigs.

Install

One-liners

  • Bare machine (no Python yet): curl -fsSL https://llm-speed.com/install.sh | sh — provisions Python via uv (with consent), installs the CLI, then runs llm-speed doctor to set up a backend. The pipx/uv lines below assume Python ≥3.10 is already present.
  • pipx (recommended if you already have Python): pipx install llm-speed
  • uv: uv tool install llm-speed
  • Homebrew: brew install llm-speed/tap/llm-speed (coming soon)
  • Docker: docker run --rm -it llmspeed/llm-speed bench (coming soon)
  • npm: npm install -g llm-speed (coming soon)
  • Standalone binary: download from Releases (coming soon)

After any install, run llm-speed doctor — it checks every dependency and, on a TTY, walks you through installing whatever's missing (a backend, a model, the daemon). On a non-TTY it prints the exact fix commands and exits non-zero.

Optional backends

  • MLX (Apple Silicon): pip install 'llm-speed[mlx]'
  • vLLM (NVIDIA): pip install 'llm-speed[vllm]'
  • ExLlamaV2 (NVIDIA): pip install 'llm-speed[exllamav2]'
  • llama.cpp + ollama are detected as binaries on PATH; no extras needed.

See docs/RELEASE.md for the publishing runbook (release steps + rollback).

Status

Phase 0 (seeding). The CLI and website come next; see docs/ for the brief and plan.

License

Code (CLI, ingest, web) is Apache-2.0 — see LICENSE. The crowdsourced benchmark data is CC BY 4.0 — see LICENSE-DATA; reuse it freely, including commercially, with attribution (a link back to llm-speed.com or the specific run). Attribution = the link, which is also how the project earns inbound citations.

Layout

docs/                Strategic & design documents
  BRIEF.md           Project brief — why, who, monetization, phases
  CLI.md             llm-speed CLI requirements + portability strategy
  DATA_SOURCES.md    Folklore source inventory + scrape policy
  MARKETING.md       CLI launch & flywheel marketing strategy

db/
  schema.sql         SQLite seed schema (mirrors the eventual Postgres prod schema)
  seed.sqlite        Created on first run

seed/                The seeding pipeline (Python)
  models.py            Plain dataclasses (RawDocument, Claim)
  db.py                Connection + insert helpers
  extractors.py        Regex-based (model × hardware × backend × tok/s) extractor
  reddit/client.py     Reddit (PRAW) — r/LocalLLaMA + neighbors, full sweep
  scrapers/hn.py             Hacker News (Algolia API)
  scrapers/openrouter.py     OpenRouter API
  scrapers/localscore.py     LocalScore HTML + Next.js data
  scrapers/artificial_analysis.py   AA public pages (cross-reference)
  scrapers/github.py         GitHub issues/PRs in inference backends
  scrapers/blogs.py          Curated blog list
  run.py               Top-level orchestrator

Quick start

# 1. Install
python -m venv .venv && source .venv/bin/activate
pip install -e .                       # uses pyproject.toml

# 2. Set credentials (only what you have available; missing ones get skipped)
export REDDIT_CLIENT_ID=...
export REDDIT_CLIENT_SECRET=...
export REDDIT_USER_AGENT="llm-speed-seeder/0.1 (+https://llm-speed.com)"
# Reddit's API rules want a description + contact info. The project URL is
# the right contact for a project-operated bot — DO NOT put `by u/<handle>`
# here, since Reddit logs the user-agent on every request and that ties
# every scrape back to a personal handle.
export GITHUB_TOKEN=...                # optional but strongly recommended

# 3. Smoke test
python -m seed.run --quick --only hn artificial_analysis blogs

# 4. Full sweep
python -m seed.run

# 5. Inspect what landed
python -m seed.run --stats
sqlite3 db/seed.sqlite \
  'SELECT model_family, hardware_name, backend, AVG(decode_tps), COUNT(*)
   FROM claims
   WHERE confidence > 0.5
   GROUP BY 1,2,3 ORDER BY 5 DESC LIMIT 30;'

Per-seeder usage

Each module is also runnable on its own:

python -m seed.reddit.client --quick                      # auth smoke test
python -m seed.scrapers.hn --query "Qwen3 Coder tok/s"
python -m seed.scrapers.openrouter --no-endpoints
python -m seed.scrapers.localscore --max-tests 50
python -m seed.scrapers.github --repo ggerganov/llama.cpp
python -m seed.scrapers.blogs --urls-file my_urls.txt
python -m seed.scrapers.artificial_analysis

Local CI (no GitHub Actions)

This project does its CI on the developer's machine — the lint / test / smoke / API-roundtrip / build pipeline that used to run on every push to GitHub Actions now lives at scripts/check.sh.

# Full check (~30s — lint, test, smoke, API roundtrip, build)
./scripts/check.sh

# Inner-loop fast pass (~10s — lint + test only)
./scripts/check.sh --quick

# Skip individual phases:
SKIP_BUILD=1 ./scripts/check.sh
SKIP_LINT=1 SKIP_BUILD=1 ./scripts/check.sh

To run it automatically on every git push (skippable per-push with --no-verify), opt in to the bundled hook once per clone:

git config core.hooksPath .githooks

The pre-push hook runs --quick by default; CHECK_FULL=1 git push runs the full check.

The previous .github/workflows/ci.yml was deleted; only manual / release-event workflows remain (release.yml, daily-metrics.yml, scheduled-seed.yml, reddit-poster.yml). None of them fire on push.

Design notes

  • Idempotent. Re-running a seeder is safe; documents are uniqued by (source, source_id).
  • Provenance preserved. Every claim links back to its source URL + author + scrape time. Folklore stays distinguishable from CLI-verified canonical results.
  • Confidence-scored. Regex extraction caps at ~0.85; structured-API extraction reaches ~0.9; nothing in the seed phase counts as canonical (that's the CLI's job).
  • Polite. All scrapers honor source-appropriate rate limits and identify themselves in User-Agent.
  • No LLM calls in this phase. Heuristic extraction only. An LLM second-pass for ambiguous cases is a Phase 1 add-on (see docs/CLI.md).

Next milestone

docs/CLI.md — design for the llm-speed benchmark CLI. The seeded folklore is inventory; the CLI is the actual data-quality moat. Ship CLI in 2–4 weeks; if adoption fails the kill criterion (see docs/MARKETING.md), stop before building the website.

Metadata

Release files for llm-speed 0.0.4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for llm-speed 0.0.4
File Size Uploaded
llm_speed-0.0.4.tar.gz 122.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for llm-speed 0.0.4
File Interpreter ABI Platform
llm_speed-0.0.4-py3-none-any.whl Python 3 none any Details

Total release size: 268.3 kB

Release files / llm_speed-0.0.4.tar.gz

Download URL llm_speed-0.0.4.tar.gz
Size 122.9 kB
Tags Source
SHA-256 checksum
How to use checksums
ecb366021e57ba0d5338c6b8156b23cf1d142577a550a4c7e76ad7fa0f36d4c7
BLAKE2b-256 checksum
How to use checksums
ddfe79addbb76d18b7533d3aea0396f8a844e0544b86810e1ce3be9e4446e624
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 26, 2026.

Transparency log

Release files / llm_speed-0.0.4-py3-none-any.whl

Download URL llm_speed-0.0.4-py3-none-any.whl
Size 145.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6a9e9f89ecdc92fbae7d02a7e0066fd05b5f1fdd295e144d34ff4b76d9fbffda
BLAKE2b-256 checksum
How to use checksums
e35e0f3838ed30de0f14e0e9db1f28be3b49057a7019354511f5979f5456b116
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Jul 26, 2026.

Transparency log

Release history Release notifications | RSS feed

0.0.7

2 release files

0.0.6

2 release files

0.0.5

2 release files

This release

0.0.4 This release

2 release files

0.0.3

2 release files

0.0.2

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page