Skip to main content

InferenceFit

InferenceFit is a local-first toolkit for evaluating and selecting LLM configurations for a specific workload. Instead of asking which model is universally "best," it runs your versioned test cases against candidate providers and models, measures quality, reliability, latency, token use, and cost, applies your hard constraints, and produces a deterministic recommendation and an open routing policy.

InferenceFit 0.2.0 is a pre-1.0 release. Public APIs and serialized schemas may change before 1.0.

Installation

InferenceFit requires Python 3.11 or newer.

pip install inferencefit

Offline quick start

The repository includes a deterministic fixture workload that needs no API key and makes no network requests. From a source checkout, run:

inferencefit validate examples/basic/eval.yaml
inferencefit benchmark examples/basic/eval.yaml

The example compares two fixtures and evaluates a schema-gated fallback cascade. The benchmark prints the run ID, recommendation, and artifact directory under .inferencefit/runs/.

The same benchmark entry point is available from Python:

import asyncio

from inferencefit import benchmark

result = asyncio.run(benchmark("examples/basic/eval.yaml"))
print(result.recommendation)

Opt-in provider examples

Live examples are intentionally separate from the offline quick start. They make network requests and can spend provider credits. Set the named environment variable in your shell, then explicitly run the corresponding smoke spec:

# Requires FIREWORKS_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.fireworks.smoke.yaml

# Requires DEEPSEEK_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.deepseek.smoke.yaml

# Requires OPEN_ROUTER_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.openrouter.smoke.yaml

# Requires GEMINI_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.gemini.smoke.yaml

# Requires OPENAI_API_KEY
inferencefit benchmark examples/lead_semantic_units/eval.openai.smoke.yaml

These specs use synthetic lead-extraction cases. See the example guide and the E0.5 validation report for their contract, pricing snapshots, and observed results.

How evaluation works

An EvaluationSpec YAML file connects a JSONL dataset to candidates, validators, hard constraints, an optimization objective, and bounded execution settings. Candidate pricing is an explicit snapshot in USD per million input and output tokens; InferenceFit does not fetch price catalogs. An authoritative per-request cost reported by OpenRouter takes precedence over configured pricing. When neither reported cost nor sufficient token counts and pricing are available, cost is unknown.

Built-in validators cover exact values, regular expressions, enums, JSON Schema, numeric values, containment, and local Python callables. Exact, numeric, and Python validators can use hidden evaluation answers and are evaluation-only. Runtime routing gates must use runtime-capable checks, such as JSON Schema. Local Python validators execute with the current process permissions and are not sandboxed, so do not run untrusted validator code.

Each candidate is summarized with end-to-end success, provider reliability, p50/p95 latency, input/output tokens, and configured cost. Hard constraints can exclude configurations by success, error rate, latency, or cost. Remaining candidates appear on a Pareto frontier and are ranked by min_cost, min_latency, max_quality, max_reliability, or the versioned balanced objective.

E0 also simulates one two-stage cascade. Fallback occurs only after a provider error or failure of a runtime-capable routing gate; hidden expected answers are never production routing signals. If no configuration satisfies the constraints, the result is explicitly non-routable.

Results, artifacts, and resume

The Python API returns a ResultBundle. Every completed CLI or Python run also writes a versioned run-artifact directory at .inferencefit/runs/<run-id>/ with:

  • manifest.json — provenance, settings, and lifecycle state;
  • spec.yaml and dataset.jsonl — exact input snapshots;
  • observations.jsonl — one flushed terminal record per planned evaluation;
  • result.json — metrics, constraints, Pareto frontier, ranking, and recommendation;
  • summary.md — a human-readable result summary;
  • routing-policy.yaml — a single, fallback, or explicitly non-routable policy.

Interrupted compatible runs can reuse completed observations:

inferencefit benchmark examples/basic/eval.yaml --resume <run-id>

Resume checks the spec and dataset hashes before skipping completed (case, candidate, repetition) identities.

Providers and credentials

The supported paths use non-streaming HTTP requests without vendor SDKs:

Provider API path Default credential environment variable
Fireworks (fireworks) OpenAI-compatible chat completions FIREWORKS_API_KEY
DeepSeek (deepseek) OpenAI-compatible chat completions DEEPSEEK_API_KEY
OpenRouter (openrouter) Chat completions with reported cost and backend metadata OPEN_ROUTER_API_KEY
Gemini (gemini) Google's OpenAI-compatible chat completions endpoint GEMINI_API_KEY
OpenAI (openai) Native Responses API with store: false OPENAI_API_KEY
Fixture (fixture) Deterministic offline testing None
Ollama (ollama) and vLLM (vllm) Local OpenAI-compatible chat completions Optional explicit reference
Custom compatible endpoint Chat completions at an explicit base_url Optional explicit reference

Credentials are resolved from opaque references. For example, credential_ref: fireworks-main checks INFERENCEFIT_CREDENTIAL_FIREWORKS_MAIN first, then FIREWORKS_API_KEY. The other hosted providers fall back to their variables in the table. OpenRouter's exact name is OPEN_ROUTER_API_KEY; the older OPENROUTER_API_KEY spelling is not a resolver fallback. Credential values are never written to run artifacts.

Minimal candidate configurations for the new providers can be placed under candidates in an evaluation spec with schema_version: "0.1" and a dataset of chat messages:

candidates:
  - id: router
    provider: openrouter
    model: openrouter/free
    credential_ref: openrouter-main
    parameters: {max_tokens: 32}
  - id: gemini
    provider: gemini
    model: gemini-3.5-flash-lite
    credential_ref: gemini-main
    parameters: {max_tokens: 32}
  - id: openai
    provider: openai
    model: gpt-6-luna
    credential_ref: openai-main
    parameters: {max_output_tokens: 64}

These are the bundled smoke defaults, not a guarantee of model availability. Model IDs and supported parameters vary by model and account; consult the provider's current documentation. For example, some OpenAI models reject temperature. OpenAI normalizes max_tokens or max_completion_tokens to max_output_tokens, rejects conflicting limits, and reserves model, input, and store. Each evaluation is stateless. An explicit base_url selects generic chat completions, including when overriding a hosted provider's native path.

For a local endpoint, use provider: ollama or provider: vllm with a model served there. For another custom compatible server, use provider: openai-compatible, an explicit base_url (such as http://localhost:8000/v1), and its model ID. Add an opaque credential_ref if authentication is required. The examples guide indexes the complete runnable specs.

Local daemon

inferencefit serve

The daemon defaults to 127.0.0.1:8787 and exposes health, run creation, status, result, and cancellation endpoints. It has no authentication and is intended only for trusted localhost use; do not expose it to a network.

Current limitations

Version 0.2.0 uses a local process job manager and filesystem artifact store. It supports non-streaming chat completions and OpenAI Responses text output, a single two-stage fallback, and Python only. Cost uses reported request cost where available, then static configured prices; it is not an invoice reconciliation system. Exact validators intentionally do not provide semantic-equivalence scoring. There is no hosted Cloud/SaaS service, account system, traffic proxy, browser UI, distributed worker system, learned routing, or model training.

Roadmap (not available in 0.2.0)

Potential future work includes richer request modalities, more provider-specific metadata, scalable artifact-store adapters, and additional language SDKs. These are directions, not current features or commitments.

Development and documentation

For an editable development install:

python -m venv .venv
pip install -e ".[dev]"
pytest
ruff check .
ruff format --check .

Normal tests use fixtures and do not spend provider credits. Live provider testing is opt-in; run it only explicitly, with the required credential configured and awareness of provider charges.

python -m pytest tests/test_provider_live.py -m provider_live -q

Each live check skips when its own key is absent. To override the bundled model defaults, set OPEN_ROUTER_TEST_MODEL, GEMINI_TEST_MODEL, or OPENAI_TEST_MODEL for the matching provider. Ordinary CI supplies no credentials and runs the offline suite.

Detailed references:

License

InferenceFit is licensed under the Apache License 2.0 (Apache-2.0).

Release files for inferencefit 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for inferencefit 0.2.0
File Size Uploaded
inferencefit-0.2.0.tar.gz 82.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for inferencefit 0.2.0
File Interpreter ABI Platform
inferencefit-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 126.3 kB

Release files / inferencefit-0.2.0.tar.gz

Download URL inferencefit-0.2.0.tar.gz
Size 82.6 kB
Tags Source
SHA-256 checksum
How to use checksums
06c504aa1ca624c4412f1cbb9f8643e5bff5979e6f785105636492bec1f0e56c
BLAKE2b-256 checksum
How to use checksums
465eb1e006175adbdb477f7a324509285809930d447a10f864eae5293a95ad21
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release files / inferencefit-0.2.0-py3-none-any.whl

Download URL inferencefit-0.2.0-py3-none-any.whl
Size 43.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9a3000a7707c84f77de96f44bbe522b4a82089cf7030ca7d771ad62ce6ff9828
BLAKE2b-256 checksum
How to use checksums
afbbba18ea53d51364ef50d38c5225d7e206405f4b70e3dba28e653be115b18a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 28, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page