Skip to main content

smollest logo
Quickly find the smollest viable language model for your task, for faster and cheaper intelligence

The basic idea is to run your OpenAI/Anthropic API queries to other, smaller models on Hugging Face API (or local), allowing you to quickly find the smollest/cheapest/fastest model that would work for your use case.

smollest dashboard screenshot

Install

pip install smollest[openai]       # for OpenAI
pip install smollest[anthropic]    # for Anthropic
pip install smollest[all]          # both

Usage

Install openai from smollest and then write your code as normal!

from smollest import openai

client = openai.OpenAI(
    api_key="sk-...",
    project="my-classifier",  # organizes results by project
)

# By default, replays to 3 models of different sizes on HF Inference API
result = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Classify as positive/negative: I love this!"}],
)

Override candidates per-client or per-call:

# Per-client
client = openai.OpenAI(
    candidates=["mistralai/Mistral-7B-Instruct-v0.3", "http://localhost:1234/v1"],
)

# Per-call
result = client.chat.completions.create(
    model="gpt-4o",
    messages=[...],
    candidates=["microsoft/Phi-3.5-mini-instruct"],
)

Works the same way with Anthropic:

from smollest import anthropic

client = anthropic.Anthropic(project="my-classifier")
result = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Classify: I love this!"}],
)

How it works

  1. Your API call goes to the baseline model as normal
  2. The same prompt is replayed to each candidate (HuggingFace serverless or local OpenAI-compatible server)
  3. Structured outputs (JSON) are compared field-by-field via exact match
  4. Results are printed to console and logged to ~/.smollest/

Remote candidates run in parallel; local candidates run sequentially.

Dashboard

smollest show

Opens a web dashboard with projects in the sidebar, a results table with truncation for long outputs, latency and cost per model, and aggregate match rates. The image above shows the UI, which you can reproduce by cloning this repo and running: python examples/demo_dashboard.py

Roadmap

  • Allow adding additional models directly through the UI
  • Add LLM as judge to score outputs that are not structured
  • Let developers eaisly fine tune models on outputs

License

MIT

Metadata

Release files for smollest 0.2.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for smollest 0.2.1
File Size Uploaded
smollest-0.2.1.tar.gz 784.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for smollest 0.2.1
File Interpreter ABI Platform
smollest-0.2.1-py3-none-any.whl Python 3 none any Details

Total release size: 801.4 kB

Release files / smollest-0.2.1.tar.gz

Download URL smollest-0.2.1.tar.gz
Size 784.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f91cd8339d59c2246fad1662fe79484dc1801460b07c13fa91da17d409358539
BLAKE2b-256 checksum
How to use checksums
395b33cd8eb98d731a6e2e2e556279a1d86f342f61053f061c3331d6333ff586
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Apr 16, 2026.

Transparency log

Release files / smollest-0.2.1-py3-none-any.whl

Download URL smollest-0.2.1-py3-none-any.whl
Size 16.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b066a7ce5ebab0f4c1fcfad8bd60f46f3725656f504bfd88940042239dbfc4b8
BLAKE2b-256 checksum
How to use checksums
0694bcfdded00a3bef63a2880338743b650dc7925e28e153514d03fc6a7cc28a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Apr 16, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.1 This release

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page