The MNIST of Gas Sensing
AI Agent-Ready simulation engine + ML benchmark for optical gas sensing
HITRAN-grade physics. 9 tasks. 25 baselines. 5 modalities. Full pipeline via natural language.
Documentation · Interactive Demo · Leaderboard · Datasets · Baselines
You don't need to be a spectroscopist. If you work on regression, denoising, domain generalization, or anomaly detection, SPEKTRAN gives you 9 ready-to-use benchmark tasks backed by real physics — with the same convenience as MNIST, but grounded in a domain where ML has direct industrial impact.
You don't even need to type code. SPEKTRAN is fully AI Agent-ready — tell Claude Code, Cursor, or any coding agent what you want, and it operates the entire pipeline through natural language. See AGENTS.md.
Highlights
Quick Start
Zero-install — load data from Hugging Face in one line:
from datasets import load_dataset
ds = load_dataset("spektran/spektran-ch4-v0") # CH4 benchmark
ds = load_dataset("spektran/spektran-co2-v0", "da") # CO2 benchmark
ds = load_dataset("spektran/spektran-industrial-v0", "so2") # Industrial SO2
ds = load_dataset("spektran/spektran-multigas-v0", "ch4_co2_h2o") # Multi-gas
Full engine — simulate spectra locally:
pip install spektran
from spektran.physics import simulate_absorbance
nu, absorbance = simulate_absorbance(
molecule="CH4", concentration_ppm=100.0,
temperature_K=296.0, pressure_atm=1.0,
path_length_m=10.0,
wavenumber_start_cm1=6046.0, wavenumber_end_cm1=6048.0,
)
More examples — WMS, multi-species, CLI
WMS 2f signal:
from spektran.physics.wms import WMSConfig, simulate_wms
See examples/wms_ch4.py for a complete WMS example.
Multi-species (CH4 + H2O interferent):
from spektran.physics import absorption_coefficient
from spektran.physics.hitran import demo_ch4_2nu3, demo_h2o
alpha_ch4 = absorption_coefficient(nu, demo_ch4_2nu3(), 100e-6, 296.0, 1.0)
alpha_h2o = absorption_coefficient(nu, demo_h2o(), 0.01, 296.0, 1.0)
See examples/multispecies_ch4_h2o.py.
CLI (all commands support --json for AI agents):
spektran info --json # Project discovery (agent bootstrap)
spektran list tasks --json # Available benchmark tasks
spektran train --baseline ridge --json # Train with auto data generation
spektran generate configs/datasets/ch4-t1-train-v0.yaml --out data --json
spektran benchmark --task T1-concentration --truth data/test.h5 --predictions preds.csv
AI Agent Ready
SPEKTRAN is designed from the ground up for the AI agent era. Every CLI command outputs structured JSON, every resource is discoverable, and the full ML pipeline — simulate → generate → train → evaluate — runs with zero manual steps.
Works with: Claude Code, Cursor, GitHub Copilot, Windsurf, Cline, and any agent that can run shell commands.
Agent interface: AGENTS.md — the agent reads this file and immediately
understands how to operate the entire project.
You: "Train the ridge baseline on T1 and show me the scores"
Agent: spektran train --baseline ridge --task T1 --json
→ {"baseline": "ridge", "scores": {"T1": {"mae_ppm": 2.84, "mape_pct": 29.87}}}
Agent workflow examples
Discovery — agent bootstraps itself:
spektran info --json # What is this project? What's available?
spektran list tasks --json # 9 tasks with metrics and available baselines
spektran list baselines --json # 25 baselines with pre-computed scores
spektran status --json # What data exists? What's been trained?
One-command training — agent trains any baseline:
spektran train --baseline ridge --json # Auto-generates data if missing
spektran train --baseline transformer --json # Works for any registered baseline
spektran train --baseline cnn1d --task T1 --json # Target a specific task
Compare all baselines — agent scripts a leaderboard run:
for baseline in ridge cnn1d transformer; do
spektran train --baseline $baseline --task T1 --json
done
Custom model — agent writes code, evaluates with CLI:
# Agent generates training code using spektran.io.read_records
# Agent writes predictions CSV
spektran benchmark --task T1-concentration \
--truth data/ch4-t1-test-v0.h5 \
--predictions my_model_predictions.csv
Benchmark Tasks
| Task | Domain | Primary Metric |
|---|---|---|
| T1 Concentration regression | DA spectrum → ppm | MAE |
| T2 Spectral denoising | Noisy → clean spectrum | RMSE |
| T3 Cross-instrument generalization | Held-out instruments | Degradation vs T1 |
| T4 WMS concentration | 2f signal → ppm | MAE |
| T5 Drift compensation | Time-series scans | Allan variance |
| T6 OOD instrument detection | In-dist vs OOD | AUROC |
| T7 Cross-modality transfer | TDLAS → NDIR | Degradation vs T1 |
| T8 Multi-species regression | CH4 + H2O → both ppm | Aggregate MAE |
| T9 Temperature regression | Spectrum → gas temp (K) | MAE |
Leaderboard
T1 Concentration + T3 Cross-Instrument (v0 splits, CH4 DA):
| Model | T1 MAE ↓ | T1 MAPE ↓ | T3 MAE ↓ | T3 Degradation |
|---|---|---|---|---|
| SpektralNet | 2.27 | 22.5% | 3.51 | 1.54x |
| Ridge regression | 2.84 | 29.9% | 3.72 | 1.31x |
| Random Forest | 5.27 | 24.1% | 10.89 | 2.07x |
| PINN | 7.29 | 49.1% | 15.62 | 2.14x |
| Patchified Transformer | 7.39 | 22.7% | 10.81 | 1.46x |
| MLP (BPNN) | 8.08 | 44.5% | 9.85 | 1.22x |
| 1D CNN | 15.58 | 42.2% | 28.30 | 1.82x |
| BiLSTM | 29.47 | 61.7% | 51.04 | 1.73x |
| CNN-LSTM-Attention | 38.39 | 69.4% | 71.03 | 1.85x |
Key insight: linear models dominate on this benchmark because Beer-Lambert absorbance is linear in concentration. SpektralNet achieves SOTA by augmenting Ridge with physics-informed features, not by adding depth.
Results for other tasks
| Task | Best Model | Score |
|---|---|---|
| T2 Denoising | 1D U-Net | RMSE 3.62e-3 |
| T4 WMS | Ridge | MAE 15.15 ppm |
| T5 Drift | Moving average | MAE 0.270 ppm |
| T6 OOD | PCA + Mahalanobis | AUROC 0.672 |
| T7 Cross-modality | Ridge (TDLAS→NDIR) | MAE 130.68 ppm (46x) |
| T8 Multi-species | Ridge (dual) | CH4 0.89 / H2O 3937 ppm |
| T9 Temperature | Ridge | MAE 9.4 K |
| T1-CRDS Concentration | Ridge (tau) | MAE 36.5 ppm |
| T1-FTIR Concentration | Ridge (spectrum) | MAE 83.7 ppm |
| T1-DOAS Concentration | Ridge (diff OD) | MAE 1.40 ppm |
Full leaderboard → · Submit results →
Real-World Impact
SPEKTRAN's benchmarks mirror problems that matter in industry and environmental science:
- Methane leak detection — oil & gas facilities, landfills, livestock operations (T1, T3)
- CO2 monitoring — greenhouse gas quantification, indoor air quality, process control (CO2 dataset)
- Industrial emissions monitoring — SO2, NO, CO stack gas analysis (industrial dataset)
- Multi-gas mixtures — disentangling overlapping species in combustion exhaust (multi-gas dataset)
- Medical breath analysis — trace-gas biomarkers at ppb levels (T1, T9)
- Instrument-agnostic deployment — models that transfer across hardware without recalibration (T3, T7)
- Drift-resilient field sensors — long-term autonomous monitoring in harsh environments (T5)
How It Works
HITRAN line data Instrument configs Benchmark
─────────────── ────────────────── ─────────
Line positions ──► Virtual instruments ──► Official splits
Line strengths (noise, fringes, (train/val/test/
Broadening params drift, chirp) held-out)
│ │ │
▼ ▼ ▼
Forward physics ──► Noisy spectra ──► Evaluate & rank
(Voigt / HTP) with provenance on leaderboard
Physics You Can Trust
- Dual-implementation validation — independent reference implementations cross-checked against HITRAN/hapi
- Literature-anchored noise — instrument noise parameters surveyed from 18 published systems
- Sim-to-real gap report — known gap sources documented in G5 report
- Automated quality gates — G1–G5 checkpoints enforced in CI
Citing
@software{spektran,
title = {SPEKTRAN: Simulation Engine and ML Benchmark for Optical Gas Sensing},
url = {https://github.com/spektran/spektran},
doi = {10.5281/zenodo.21790394},
version = {0.6.0},
license = {Apache-2.0}
}
See CITATION.cff or use the DOI: 10.5281/zenodo.21790394.
Star History
Contributing
See CONTRIBUTING.md. We welcome new baselines, modalities, and line lists.
Release files for spektran 0.6.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| spektran-0.6.0.tar.gz | 15.3 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| spektran-0.6.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 15.5 MB
Release files / spektran-0.6.0.tar.gz
| Download URL | spektran-0.6.0.tar.gz |
|---|---|
| Size | 15.3 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
837713cdcfdf605bc67bbd5f322d7defaea69beb768b508d8efa6be3e85d479e
|
|
BLAKE2b-256 checksum How to use checksums |
bd7fffb177200a4560b2007d4638aed9ce4944af4258a91f9b2cea300e5cf143
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 12, 2026.
Transparency logRelease files / spektran-0.6.0-py3-none-any.whl
| Download URL | spektran-0.6.0-py3-none-any.whl |
|---|---|
| Size | 114.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
4b0b8b388cb7425f0d049a1f5193854a9a7af6006bcf24233b1849deb06b3eda
|
|
BLAKE2b-256 checksum How to use checksums |
45cc7168a45be3f8067a002b46bf4ffd09fab7b99cf82a5715b29325d9a7dbad
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 12, 2026.
Transparency log