InferenceFit
InferenceFit is a local-first toolkit for evaluating and selecting LLM configurations for a specific workload. Instead of asking which model is universally "best," it runs your versioned test cases against candidate providers and models, measures quality, reliability, latency, token use, and cost, applies your hard constraints, and produces a deterministic recommendation and an open routing policy.
InferenceFit 0.1.0 is a pre-1.0 release. Public APIs and serialized schemas may change before 1.0.
Installation
InferenceFit requires Python 3.11 or newer.
pip install inferencefit
Offline quick start
The repository includes a deterministic fixture workload that needs no API key and makes no network requests. From a source checkout, run:
inferencefit validate examples/basic/eval.yaml
inferencefit benchmark examples/basic/eval.yaml
The example compares two fixtures and evaluates a schema-gated fallback cascade. The benchmark
prints the run ID, recommendation, and artifact directory under .inferencefit/runs/.
The same benchmark entry point is available from Python:
import asyncio
from inferencefit import benchmark
result = asyncio.run(benchmark("examples/basic/eval.yaml"))
print(result.recommendation)
Opt-in provider examples
Live examples are intentionally separate from the offline quick start. They make network requests and can spend provider credits. Set the named environment variable in your shell, then explicitly run the corresponding smoke spec:
# Requires FIREWORKS_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.fireworks.smoke.yaml
# Requires DEEPSEEK_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.deepseek.smoke.yaml
These specs use synthetic lead-extraction cases. See the example guide and the E0.5 validation report for their contract, pricing snapshots, and observed results.
How evaluation works
An EvaluationSpec YAML file connects a JSONL dataset to candidates, validators, hard constraints,
an optimization objective, and bounded execution settings. Candidate pricing is an explicit
snapshot in USD per million input and output tokens; InferenceFit does not fetch prices.
Built-in validators cover exact values, regular expressions, enums, JSON Schema, numeric values, containment, and local Python callables. Exact, numeric, and Python validators can use hidden evaluation answers and are evaluation-only. Runtime routing gates must use runtime-capable checks, such as JSON Schema. Local Python validators execute with the current process permissions and are not sandboxed, so do not run untrusted validator code.
Each candidate is summarized with end-to-end success, provider reliability, p50/p95 latency,
input/output tokens, and configured cost. Hard constraints can exclude configurations by success,
error rate, latency, or cost. Remaining candidates appear on a Pareto frontier and are ranked by
min_cost, min_latency, max_quality, max_reliability, or the versioned balanced objective.
E0 also simulates one two-stage cascade. Fallback occurs only after a provider error or failure of a runtime-capable routing gate; hidden expected answers are never production routing signals. If no configuration satisfies the constraints, the result is explicitly non-routable.
Results, artifacts, and resume
The Python API returns a ResultBundle. Every completed CLI or Python run also writes a versioned
run-artifact directory at .inferencefit/runs/<run-id>/ with:
manifest.json— provenance, settings, and lifecycle state;spec.yamlanddataset.jsonl— exact input snapshots;observations.jsonl— one flushed terminal record per planned evaluation;result.json— metrics, constraints, Pareto frontier, ranking, and recommendation;summary.md— a human-readable result summary;routing-policy.yaml— a single, fallback, or explicitly non-routable policy.
Interrupted compatible runs can reuse completed observations:
inferencefit benchmark examples/basic/eval.yaml --resume <run-id>
Resume checks the spec and dataset hashes before skipping completed
(case, candidate, repetition) identities.
Providers and credentials
The fixture provider is deterministic and offline. The generic, non-streaming OpenAI-compatible
adapter accepts an explicit base URL and includes endpoint presets for Fireworks, DeepSeek,
OpenRouter, Ollama, and vLLM. It uses the providers' chat-completions-shaped HTTP interface without
vendor SDKs.
Credentials are resolved from opaque references. For example, credential_ref: fireworks-main
checks INFERENCEFIT_CREDENTIAL_FIREWORKS_MAIN and then FIREWORKS_API_KEY; DeepSeek similarly
falls back to DEEPSEEK_API_KEY, and OpenRouter to OPENROUTER_API_KEY. Secret values are not
written to run artifacts.
Local daemon
inferencefit serve
The daemon defaults to 127.0.0.1:8787 and exposes health, run creation, status, result, and
cancellation endpoints. It has no authentication and is intended only for trusted localhost use;
do not expose it to a network.
Current limitations
Version 0.1.0 uses a local process job manager and filesystem artifact store. It supports non-streaming chat completions, a single two-stage fallback, and Python only. Pricing is static configuration rather than provider billing data. Exact validators intentionally do not provide semantic-equivalence scoring. There is no hosted Cloud/SaaS service, account system, traffic proxy, browser UI, distributed worker system, learned routing, or model training.
Roadmap (not available in 0.1.0)
Potential post-0.1 work includes richer request modalities, more provider-specific metadata, scalable artifact-store adapters, and additional language SDKs. These are directions, not current features or commitments.
Development and documentation
For an editable development install:
python -m venv .venv
pip install -e ".[dev]"
pytest
ruff check .
ruff format --check .
Normal tests use fixtures and do not spend provider credits. Live provider testing is opt-in; run it only explicitly, with the required credential configured and awareness of provider charges.
Detailed references:
License
InferenceFit is licensed under the Apache License 2.0 (Apache-2.0).
Release files for inferencefit 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| inferencefit-0.1.0.tar.gz | 67.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| inferencefit-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 105.4 kB
Release files / inferencefit-0.1.0.tar.gz
| Download URL | inferencefit-0.1.0.tar.gz |
|---|---|
| Size | 67.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
47cf71bcd93835cd445df7ecbc82b0a252ad4a4eb3ab0fbc91f2b6bfd366103c
|
|
BLAKE2b-256 checksum How to use checksums |
1f7163392cfb3e4950c5b02d048551b0d387921e0670d8ccd631b58cedd029d5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency logRelease files / inferencefit-0.1.0-py3-none-any.whl
| Download URL | inferencefit-0.1.0-py3-none-any.whl |
|---|---|
| Size | 38.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c2512c938ec10c876a826966d10596904e1dbe20a6128c2ac8bd8fd55a3546ca
|
|
BLAKE2b-256 checksum How to use checksums |
c303b3ae8f777f5597d8b438423c143d4bb5b787a06590bf43ac01b4a0d049fd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 27, 2026.
Transparency log