Skip to main content

llm-code-benchmark

AI Cost Tracking

PyPI Version Python License AI Cost Human Time Model

  • 🤖 LLM usage: $1.0836 (42 commits)
  • 👤 Human dev: ~$1004 (10.0h @ $100/h, 30min dedup)

Generated on 2026-07-19 using openrouter/qwen/qwen3-coder-next


Private benchmark repository for selecting OpenRouter models for repair-agent and validator-agent.

The benchmark never stores API keys. Paid runs require OPENROUTER_API_KEY from the environment or GitHub Actions Secrets. Standard CI and catalog-smoke do not make paid model requests.

python -m pip install -e ".[test]"
python -m pytest tests -q
llm-code-benchmark catalog-smoke --max-models 2

Paid smoke benchmark, after explicit approval and budget setup:

llm-code-benchmark live-smoke --max-models 2 --max-tasks 2 --repetitions 1 --budget-usd 0.25

Recommended next smoke uses explicit models instead of router auto-selection:

llm-code-benchmark live-smoke --model meta-llama/llama-4-scout,deepseek/deepseek-v4-flash --max-models 2 --max-tasks 2 --repetitions 1 --budget-usd 0.25 --dry-run

Remove --dry-run only after explicit paid-run approval.

Reports are written to reports/latest/ and immutable snapshots under reports/history/. Repair schema validation is intentionally separate from semantic patch validation, so path mismatches and invalid diffs are reported as repair task statuses rather than JSON schema failures.

Repair Output Modes

Repair benchmark requests now prefer file_edits: the model returns repository-relative paths and complete final file contents, and the benchmark generates a deterministic unified diff locally before validation and git apply --check. Native patch output remains supported and is scored separately through native_patch_success_rate; structured edits are tracked through structured_edit_success_rate. search_replace, create, and delete are not enabled in this version.

License

Licensed under Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

llm_code_benchmark-0.2.1.tar.gz (54.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

llm_code_benchmark-0.2.1-py3-none-any.whl (49.0 kB view details)

Uploaded Python 3

File details

Details for the file llm_code_benchmark-0.2.1.tar.gz.

File metadata

  • Download URL: llm_code_benchmark-0.2.1.tar.gz
  • Upload date:
  • Size: 54.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.7

File hashes

Hashes for llm_code_benchmark-0.2.1.tar.gz
Algorithm Hash digest
SHA256 3c7fa5ea5d72d7f8df7368dc5e8670ec5af661aab3ef550affb1c0554cb3d418
MD5 52fff571020f36968b92f8828e15a7e4
BLAKE2b-256 da2647f5fd743ee6292e1a48e2ef5e6ca3fc4c93a8ffae3fa6871b126e8db952

See more details on using hashes here.

File details

Details for the file llm_code_benchmark-0.2.1-py3-none-any.whl.

File metadata

File hashes

Hashes for llm_code_benchmark-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 a73efaa4cf6b6dfbaa936db576bb66a9fcd8f44bdc35d14b7d2614190fbb180d
MD5 ce17b0a86b2b0e2c34eae3558923e580
BLAKE2b-256 9b454ce2cfe7ca89aed8fdbf6b6dc03d2209946c445366f064edb6818c1673f2

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page