Skip to main content

sgl-eval

License Python

One-click accuracy evaluation harness for SGLang.

Point at any OpenAI-compatible endpoint. Scoring logic (graders, evaluators, prompts, dataset configs) is vendored from NeMo-Skills; sgl-eval contributes the transport, runner, and benchmark wiring.


Quick start

pip install sgl-eval

sgl-eval ping --base-url http://localhost:30000/v1
sgl-eval run gsm8k --base-url http://localhost:30000/v1 --num-examples 50

Four subcommands: run, list, ping, preset. sgl-eval run --help is the full flag reference -- endpoint, sampling overrides (--temperature, --seed, --thinking, ...), and any flags the benchmark itself adds.


Reading a run

Each run prints the headline metric first -- single-shot accuracy, averaged across the k repeats when k > 1 -- and writes the same payload plus provenance (model, endpoint, sampling config, vendored NS commit) as metrics.json under --out-dir.

== aime25 ==
30 examples x 16 repeats  |  823.7s  |  4293 tok/s  |  3.5M tokens

* pass@1[avg-of-16]  =  78.96% +/- 1.21% (SEM 0.30%)
  pass@16            =  93.33%
  majority@16        =  93.33%
  no_answer          =  20.00%  [warn: consider --max-tokens]

While the run is going, the progress bar carries a live accuracy. For a sanity check that is usually the whole point: watch it, decide, stop.

gsm8k:  34%|###4      | 452/1319 [02:11<04:12, 3.4it/s, acc=81.42%]

Every scored sample is streamed to <out-dir>/sgl_eval_<name>_<stamp>/output-rs*.jsonl as it lands (disable with --no-dump-predictions), so the per-sample record survives however the run ends. Each record stores the assistant's final response in generation and separately parsed reasoning in reasoning_content when the endpoint provides it.


Running less than the whole thing

  • --num-examples N -- only the first N examples.
  • Ctrl-C -- kills in-flight requests, keeps everything already scored, and writes metrics.json flagged partial: true with how much ran, so a half-run can't later be mistaken for a full one. Exits 130; a second Ctrl-C hard-exits if cleanup hangs. The preset expected_vs_actual comparison is skipped -- a half-run isn't comparable to a baseline.
  • --from-dataset <path> -- swap in your own NS-shape JSONL ({id?, problem, expected_answer}) for one run. Only the questions change; scoring still goes through the vendored grader.

Benchmarks

sgl-eval list for the registered set, sgl-eval list -v for each one's defaults. See benchmarks.md for the ones that need more than an endpoint (today: ruler2), and for how to match a NeMo-Skills run.

Presets

Save a (benchmark, endpoint, sampling, n_repeats, expected) bundle to ~/.sgl_eval/presets/<name>.yaml and replay with sgl-eval run --preset <name>. See preset.md for schema, example, usage, and override priority.

For repository-maintained model defaults, select an exact supported model ID:

sgl-eval run BENCHMARK \
  --base-url BASE_URL \
  --load-preset-from-model-id MODEL_ID

This sets the served model and its recommended generation parameters, but not the deployment-specific --base-url. See preset.md for the supported model list, resolved values, and override priority.


Architecture

Anything that decides a score is vendored verbatim from NeMo-Skills. sgl-eval contributes only transport: an OpenAI client, a threadpool runner, a CLI, and the thin glue that wires upstream pieces into one command.

+----------------------------------------------------+
|  sgl-eval                                          |
|    cli, sampler, runner, registry, metrics         |
|    evals/                                          |
+----------------------------------------------------+
|  vendored from NeMo-Skills                         |
|    math_grader, evaluator/, metrics/,              |
|    dataset/<bench>/, prompts/*.yaml                |
+----------------------------------------------------+

The slice is pinned at a specific commit in sgl_eval/_vendored/nemo_skills/SOURCES.yaml. To upgrade, bump synced_from_sha there and run:

python scripts/sync_vendored.py    # re-fetch all vendored files
pytest                             # upstream's own tests run against the
                                   # new slice -- catches behavior drift

Adding a benchmark inside an existing category (math, multichoice) is one row in _registry.py:_TABLE. A new category needs a runner alongside it -- graders are usually already in NeMo-Skills.


Scope

The goal is to be the single accuracy-eval client SGLang's CI calls, in place of sglang.test.run_eval and the assorted per-test harnesses.

Not in scope: performance benchmarking (latency / throughput / scheduling -- that is SGLang's bench_serving.py; sgl-eval records them only as side metrics, never as the headline), training or fine-tuning, multi-server orchestration (one endpoint per invocation), and OS-level agent loops.


License

Apache-2.0. See LICENSE. Vendored NeMo-Skills sources are also Apache-2.0; see NOTICE for attribution and the list of vendored files.

Metadata

Release files for sgl-eval 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sgl-eval 0.1.1
File Size Uploaded
sgl_eval-0.1.1.tar.gz 162.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sgl-eval 0.1.1
File Interpreter ABI Platform
sgl_eval-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 341.9 kB

Release files / sgl_eval-0.1.1.tar.gz

Download URL sgl_eval-0.1.1.tar.gz
Size 162.6 kB
Tags Source
SHA-256 checksum
How to use checksums
0fe03c2f3bc55c8bff71867f92be9d6a2827a5c52c442ac7272d206e061e1d90
BLAKE2b-256 checksum
How to use checksums
3a70543e48b06cc5f21a8bb067c3b86716b3dd6d60ed20a2c16ab6ca549f4e2e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release files / sgl_eval-0.1.1-py3-none-any.whl

Download URL sgl_eval-0.1.1-py3-none-any.whl
Size 179.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cd82286dd83b13049269932703576f1447b8a6c1c4c406a81ee5013302c364c4
BLAKE2b-256 checksum
How to use checksums
a5f8a584e6adcc660bfbd693b5a88060a129578f2cea7d432dde61140178fae2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 12, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.2

2 release files

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page