Skip to main content

RedCrown

Localized model benchmarking with receipts. Run head-to-head evals across every model, provider, and config on your own data, locally, then turn the results into ranked, receipted proof you can hand to a client, a CFO, or a regulator.

Agents and harnesses run benchmarks for free. RedCrown is the neutral layer that turns them into a decision you can defend, and keeps re-proving it as prices and models change. Run anywhere, prove here.

Install

pip install redcrown

Python 3.11+. The CLI has a small dependency footprint; server extras (pip install "redcrown[server]") are only needed if you run the API yourself.

Quickstart (no keys needed)

Try the bundled offline demo in about five seconds, no provider keys required:

redcrown eval --sample extract

This runs a small bundled extraction sample using stub outputs. The output says demo mode because no real models are called. Connect your provider keys in ~/.redcrown/credentials.json (or via redcrown login) for real quality scores.

Bring your own data (text tasks)

If you have a CSV or JSONL file of inputs, import it directly:

redcrown build-dataset --from-csv mydata.csv --task classify --out exp.json
redcrown eval exp.json --report-json out.json

CSV/JSONL columns:

column required notes
input yes the text sent to the model
reference no expected output (ground truth); omit to use the no-labels path below
id no stable row ID; auto-generated if absent

--task values: extract, classify, summarize, qa

The importer fills in the provider scaffold for you. Edit exp.json to swap in the candidates you want to compare before running.

The exp.json schema

Every eval is a plain JSON file. Here is a minimal 2-candidate text example you can copy, edit, and run:

{
  "name": "classify-support-tickets",
  "quality_metric": "similarity",
  "quality_bar": 0.75,
  "reference_source": "labels",
  "pipeline": [
    {
      "id": "step1",
      "order": 1,
      "name": "classify",
      "input_type": "text",
      "output_type": "text",
      "input_binding": "dataset",
      "incumbent_candidate_id": "gpt4o-mini",
      "candidates": [
        {
          "id": "gpt4o-mini",
          "provider": "openai",
          "model": "gpt-4o-mini",
          "label": "GPT-4o Mini (incumbent)"
        },
        {
          "id": "llama-8b",
          "provider": "groq",
          "model": "llama-3.1-8b-instant",
          "label": "Llama 3.1 8B via Groq"
        }
      ]
    }
  ],
  "dataset": [
    {
      "id": "row1",
      "payload": { "kind": "text", "text": "My invoice is wrong." },
      "reference": "billing"
    },
    {
      "id": "row2",
      "payload": { "kind": "text", "text": "App crashes on login." },
      "reference": "bug"
    }
  ]
}

Required fields: name, quality_metric, quality_bar, reference_source, pipeline[].incumbent_candidate_id, pipeline[].candidates[].id, pipeline[].candidates[].provider, pipeline[].candidates[].model, dataset[].id, dataset[].payload.text.

No labels? Use the incumbent as reference

If you do not have labeled ground truth, leave reference off your dataset rows and set reference_source: "incumbent":

{
  "reference_source": "incumbent",
  ...
  "dataset": [
    { "id": "row1", "payload": { "kind": "text", "text": "My invoice is wrong." } }
  ]
}

RedCrown scores every candidate against your current model's own output, so you only need to beat or match what you run today. This is the fastest path to a cost proof when you have real traffic but no labeled examples.

Running a full eval

# run the fan-out locally, on your machine, with your keys
redcrown eval exp.json --report-json out.json

# (optional) sign in once per machine via device-code OAuth
redcrown login

# (optional) push the results to a shareable, no-login proof page
redcrown push out.json --proof-link

redcrown eval ranks every config on cost, quality, and latency against your own ground truth and names the cheapest one that clears your quality bar. Example:

RANKED  transcription · cheapest config at or above your 0.85 quality bar
  deepgram · nova-3-medical    quality 0.883    $294/mo   winner, 40% cheaper
  aws · transcribe-standard    quality 0.879    $487/mo   incumbent
  openai · whisper-1           quality 0.820    $122/mo   below your bar

That run is published as a live, no-login proof page: https://app.redcrown.ai/proof/O9iYVdWuaYjaL6mnImeIsD6TB1W_S6h4Frbx04YqAYQ

Note on build-dataset primock57: this command downloads the PriMock57 clinical-transcription corpus, which requires git-lfs. It is an audio eval designed for transcription benchmarking, not a general first step. Start with redcrown eval --sample extract or build-dataset --from-csv instead.

Free by construction

Evals run on your machine with your own provider keys, so RedCrown never sees your raw data and the run costs you nothing beyond your own inference. Only the results you choose to push become a cloud proof. --no-receipts keeps raw outputs local and uploads aggregates only.

Already ran an eval elsewhere?

You do not have to run anything through RedCrown to get a proof. Take the results from an eval you already ran, as a JSON in the RedCrown results format, and push them:

redcrown push results.json --proof-link

You get the same ranked, receipted, shareable report. The fastest path, with no install, is the web app at https://app.redcrown.ai/upload.

For coding agents (MCP)

Coding agents (Claude, Cursor, Codex) drive the whole loop over the hosted MCP server at mcp.redcrown.ai: scaffold an experiment, run it, review outputs, and mint a proof. The server is open source: https://github.com/RedCrown-ai/redcrown-mcp

Links

License

Proprietary. (c) Method Data Science LLC.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

redcrown-0.1.5.tar.gz (185.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

redcrown-0.1.5-py3-none-any.whl (126.9 kB view details)

Uploaded Python 3

File details

Details for the file redcrown-0.1.5.tar.gz.

File metadata

  • Download URL: redcrown-0.1.5.tar.gz
  • Upload date:
  • Size: 185.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for redcrown-0.1.5.tar.gz
Algorithm Hash digest
SHA256 f9c17a8eb742cef70e4a96aaa3728e48184c3c3aba263df339c3caa2d8011fe1
MD5 4e5a6ca673697c2b019fa6a41ac6389d
BLAKE2b-256 08880cca762ef3c2d60b1a7e71aa286624c373ab3355762f955606fd09dd7cf5

See more details on using hashes here.

File details

Details for the file redcrown-0.1.5-py3-none-any.whl.

File metadata

  • Download URL: redcrown-0.1.5-py3-none-any.whl
  • Upload date:
  • Size: 126.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.13

File hashes

Hashes for redcrown-0.1.5-py3-none-any.whl
Algorithm Hash digest
SHA256 9c321309a332337c84c2a0ea93a2cb747201aecf2d6391d77b12c70e28b81c25
MD5 a689b8147225ec464699bc379de2ee9e
BLAKE2b-256 c0a356ceb98cd6ecb50a6f8b5856e662647b3162c317ec9efa63cb16d5969b82

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.18

2 files

0.1.17

2 files

0.1.16

2 files

0.1.15

2 files

0.1.14

2 files

0.1.13

2 files

0.1.12

2 files

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

This release

0.1.5 This release

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page