NiriZan
Continuous Evaluation Infrastructure for Production AI
"Inspection through Measurement" "Engineering Trust Through Continuous Evaluation"
Why NiriZan?
Modern AI systems are probabilistic rather than deterministic. Traditional software testing alone cannot determine whether a retrieval pipeline, language model, or AI agent is performing correctly. NiriZan exists to provide continuous, reproducible evaluation infrastructure that enables teams to measure quality, detect regressions, compare experiments, and build confidence in production AI systems.
NiriZan is an open-source framework to provide automated judge drift attribution, fixed anchor sets with repeatable, on-demand rescoring, rigorous statistical gating (Mann-Whitney + Holm-Bonferroni), trust-weighted health scoring, and CI/CD-integrated regression gating in a single, architecturally disciplined Python package.
Installation
pip install nirizan
Package: pypi.org/project/nirizan
What NiriZan Does
- Automated judge-drift attribution.
AttributionEngineproduces a five-state verdict —NONE,SYSTEM_DRIFT,JUDGE_DRIFT,JOINT_DRIFT, orINCONCLUSIVE— distinguishing a quality drop in the system under test from a change in the judge measuring it, and separately flagging when both shifted at once or when there wasn't enough data to tell. - Fixed evaluation anchors, rescored on demand. A versioned
AnchorSetis never edited in place; updating it means creating a newanchor_set_id, so historical comparisons stay meaningful. - Statistically rigorous regression gating. Mann-Whitney U tests with Holm-Bonferroni correction for multiple comparisons, Cohen's d effect sizes, and bootstrap confidence intervals (5,000 resamples), not a bare threshold on a single score.
- Trust-weighted health scoring.
compute_system_health_scorediscounts the aggregate score when the attribution verdict signals judge unreliability, not just system degradation. - An 8-layer, unidirectional architecture,
instrumentation → orchestrator → metrics → trust → storage → regression → gate → reporting, enforced byimport-linterin CI, not just documented as a diagram.
What is NiriZan?
NiriZan is an open-source continuous evaluation infrastructure for production AI systems. It enables engineers and researchers to systematically measure, benchmark, validate, and monitor the quality of:
- Retrieval-Augmented Generation (RAG) pipelines
- AI agents
- Large Language Model (LLM) applications
- Custom AI workflows
Unlike orchestration frameworks that focus on building AI applications, NiriZan focuses on engineering confidence in AI systems. It provides:
| Capability | Description |
|---|---|
| Reproducible evaluation pipelines | Consistent, repeatable test runs across environments |
| Benchmark execution | Standardized quality benchmarking for AI systems |
| Regression detection | Automated flagging of quality drops between versions |
| Experiment tracking | Full history of runs, configs, and results |
| Quality reporting | Clear, actionable reports on system performance |
| Deployment-aware validation | Checks tuned to pre-, during-, and post-deployment stages |
Vision
The long-term vision of NiriZan is to become the engineering quality layer for production AI, ensuring that every AI application can be continuously measured before, during, and after deployment.
Where the Name Comes From
NiriZan is a fusion of two words from two languages, each contributing a core idea behind the project.
| Niri | Zan |
|---|---|
| Origin: নিরীক্ষা (Nirikkha) - Bangla/Bengali | Origin: ميزان (Mīzān) - Arabic |
| Meaning: Inspection · Evaluation · Verification · Audit | Meaning: Scale · Balance · Measurement · Criterion |
Together, Niri + Zan captures the essence of the project: inspecting AI systems and measuring them against a balanced standard of quality.
Read More
For the complete user guide, see the NiriZan User Manual. link
For architecture, contracts, module reference docs, and the evaluation results behind the claims above, see docs/.
See CHANGELOG.md for release history and notable changes between versions.
The Ruler Can Change Too: Navigating Judge Drift in Production AI Evaluation, the first NiriZan engineering post, covering the judge-drift problem and the fixed-anchor, statistical-attribution approach this project takes to it.
License
Copyright (C) 2026 Redwan Rahman
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
Author
Redwan Rahman github.com/Red1-Rahman
Metadata
Release files for nirizan 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| nirizan-0.2.0.tar.gz | 1.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| nirizan-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.3 MB
Release files / nirizan-0.2.0.tar.gz
| Download URL | nirizan-0.2.0.tar.gz |
|---|---|
| Size | 1.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
3804d0d9295211959662e12548eb1f2865da1f414d3ae431206f877b98842eed
|
|
BLAKE2b-256 checksum How to use checksums |
e91c8a70955a7b937d1529a392e84f70ffca85a6443c694aaeb5186914365f3a
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.
Transparency logRelease files / nirizan-0.2.0-py3-none-any.whl
| Download URL | nirizan-0.2.0-py3-none-any.whl |
|---|---|
| Size | 63.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2c85c89ac28f49b9e5afcf832c2fea9a3dce2da84aa1c03049683d2f8f682711
|
|
BLAKE2b-256 checksum How to use checksums |
c7829ee1de1d014c5ee09bbb45e44bcebcdd2345fde2e52f14d4978701b9f408
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 28, 2026.
Transparency log