Skip to main content

sgl-eval

License Python

One-click accuracy evaluation harness for SGLang.

Point at any OpenAI-compatible endpoint. Scoring logic (graders, evaluators, prompts, dataset configs) is vendored from NeMo-Skills; sgl-eval contributes the transport, runner, and benchmark wiring.


Quick start

pip install sgl-eval

sgl-eval ping --base-url http://localhost:30000/v1
sgl-eval run gsm8k --base-url http://localhost:30000/v1 --num-examples 50

Four subcommands: run, list, ping, preset. sgl-eval run --help is the full flag reference -- endpoint, sampling overrides (--temperature, --seed, --thinking, ...), and any flags the benchmark itself adds.


Reading a run

Each run prints the headline metric first -- single-shot accuracy, averaged across the k repeats when k > 1 -- and writes the same payload plus provenance (model, endpoint, sampling config, vendored NS commit) as metrics.json under --out-dir.

== aime25 ==
30 examples x 16 repeats  |  823.7s  |  4293 tok/s  |  3.5M tokens

* pass@1[avg-of-16]  =  78.96% +/- 1.21% (SEM 0.30%)
  pass@16            =  93.33%
  majority@16        =  93.33%
  no_answer          =  20.00%  [warn: consider --max-tokens]

While the run is going, the progress bar carries a live accuracy. For a sanity check that is usually the whole point: watch it, decide, stop.

gsm8k:  34%|###4      | 452/1319 [02:11<04:12, 3.4it/s, acc=81.42%]

Every scored sample is streamed to <out-dir>/sgl_eval_<name>_<stamp>/output-rs*.jsonl as it lands (disable with --no-dump-predictions), so the per-sample record survives however the run ends.


Running less than the whole thing

  • --num-examples N -- only the first N examples.
  • Ctrl-C -- kills in-flight requests, keeps everything already scored, and writes metrics.json flagged partial: true with how much ran, so a half-run can't later be mistaken for a full one. Exits 130; a second Ctrl-C hard-exits if cleanup hangs. The preset expected_vs_actual comparison is skipped -- a half-run isn't comparable to a baseline.
  • --from-dataset <path> -- swap in your own NS-shape JSONL ({id?, problem, expected_answer}) for one run. Only the questions change; scoring still goes through the vendored grader.

Benchmarks

sgl-eval list for the registered set, sgl-eval list -v for each one's defaults. See benchmarks.md for the ones that need more than an endpoint (today: ruler2), and for how to match a NeMo-Skills run.

Presets

Save a (benchmark, endpoint, sampling, n_repeats, expected) bundle to ~/.sgl_eval/presets/<name>.yaml and replay with sgl-eval run --preset <name>. See preset.md for schema, example, usage, and override priority.

For repository-maintained model defaults, select an exact supported model ID:

sgl-eval run BENCHMARK \
  --base-url BASE_URL \
  --load-preset-from-model-id MODEL_ID

This sets the served model and its recommended generation parameters, but not the deployment-specific --base-url. See preset.md for the supported model list, resolved values, and override priority.


Architecture

Anything that decides a score is vendored verbatim from NeMo-Skills. sgl-eval contributes only transport: an OpenAI client, a threadpool runner, a CLI, and the thin glue that wires upstream pieces into one command.

+----------------------------------------------------+
|  sgl-eval                                          |
|    cli, sampler, runner, registry, metrics         |
|    evals/                                          |
+----------------------------------------------------+
|  vendored from NeMo-Skills                         |
|    math_grader, evaluator/, metrics/,              |
|    dataset/<bench>/, prompts/*.yaml                |
+----------------------------------------------------+

The slice is pinned at a specific commit in sgl_eval/_vendored/nemo_skills/SOURCES.yaml. To upgrade, bump synced_from_sha there and run:

python scripts/sync_vendored.py    # re-fetch all vendored files
pytest                             # upstream's own tests run against the
                                   # new slice -- catches behavior drift

Adding a benchmark inside an existing category (math, multichoice) is one row in _registry.py:_TABLE. A new category needs a runner alongside it -- graders are usually already in NeMo-Skills.


Scope

The goal is to be the single accuracy-eval client SGLang's CI calls, in place of sglang.test.run_eval and the assorted per-test harnesses.

Not in scope: performance benchmarking (latency / throughput / scheduling -- that is SGLang's bench_serving.py; sgl-eval records them only as side metrics, never as the headline), training or fine-tuning, multi-server orchestration (one endpoint per invocation), and OS-level agent loops.


License

Apache-2.0. See LICENSE. Vendored NeMo-Skills sources are also Apache-2.0; see NOTICE for attribution and the list of vendored files.

Metadata

Release files for sgl-eval 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sgl-eval 0.1.0
File Size Uploaded
sgl_eval-0.1.0.tar.gz 168.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sgl-eval 0.1.0
File Interpreter ABI Platform
sgl_eval-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 353.6 kB

Release files / sgl_eval-0.1.0.tar.gz

Download URL sgl_eval-0.1.0.tar.gz
Size 168.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a1f4bccfeebf3e5a5cb176885f1195e9999124d07eb3cd8b8930cfc1c2a5b0b6
BLAKE2b-256 checksum
How to use checksums
3d5e9528ce6c0182854e98d45c170e9eaf99ffe16c6113efaa36bc30b772875d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.

Transparency log

Release files / sgl_eval-0.1.0-py3-none-any.whl

Download URL sgl_eval-0.1.0-py3-none-any.whl
Size 185.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2d80e7d7d8ba13ced70cf09757bb0f0805cff881f7da7503db7397000628a25f
BLAKE2b-256 checksum
How to use checksums
4eaba62c0afcff2e7948ed6b619e499979a0b5da9a356abf054b1fb0c1e13f90
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 2, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page