Skip to main content

pytest for ML — unified testing across the entire ML lifecycle

Project description

mltk - ML Test Kit

pytest for ML -- unified testing across the entire ML lifecycle.

License Python Tests Rust

Installation note (temporary): pip install mltk is not yet available — the PyPI name is pending a transfer claim. Until resolved, install directly from GitHub:

pip install git+https://github.com/Liorrr/mltk

All imports (import mltk), CLI commands (mltk scan), and functionality are identical.

Install

Homebrew (macOS/Linux):

brew tap Liorrr/mltk
brew install mltk

pip:

pip install mlspec  # PyPI name — 'mltk' pending transfer

Why mltk?

ML systems fail silently. A model can train on corrupt data, produce confident predictions from stale features, and pass every unit test while being completely wrong in production. Traditional testing does not catch these failures.

mltk gives you 241 assertions covering the entire ML lifecycle -- data quality, model validation, drift detection, fairness testing, inference benchmarking, training bug detection, LLM evaluation, behavioral consistency, NER-based PII detection, red team security, multimodal evaluation, observability, and production monitoring. One toolkit, one pip install, native pytest integration. No more gluing together 5 different tools.

Quick Start

import pandas as pd
import pytest
from mltk.data import (
    assert_schema, assert_no_nulls, assert_no_pii,
)
from mltk.model import (
    assert_metric, assert_no_regression, assert_no_bias,
)
from mltk.inference import assert_latency, assert_throughput

@pytest.mark.ml_data
def test_training_data():
    df = pd.read_csv("data/training.csv")
    assert_schema(
        df,
        {"id": "int64", "text": "object", "label": "int64"},
    )
    assert_no_nulls(df, columns=["label"])
    assert_no_pii(df, columns=["text"])

@pytest.mark.ml_model
def test_model_quality(y_true, y_pred):
    assert_metric(
        y_true, y_pred, metric="f1", threshold=0.85,
    )
    assert_no_regression(
        y_true, y_pred, baseline=0.90, tolerance=0.02,
    )
    assert_no_bias(
        y_true, y_pred,
        sensitive_feature=gender,
        method="demographic_parity",
    )

@pytest.mark.ml_inference
def test_inference_performance():
    assert_latency(model.predict, X_test, p95=50.0)
    assert_throughput(model.predict, X_single, min_rps=100)

Run with HTML report:

pytest --mltk-report

What's Included (v0.12.4)

Behavioral consistency testing — 7 assertions (paraphrase invariance, format invariance, output stability, semantic equivalence, directional expectation, retrieval consistency, ParaphraseGenerator) that catch models memorizing phrasing instead of learning concepts. Multi-method evaluation: lexical (token F1), embedding, NLI, LLM-as-Judge. No other ML testing tool ships these as pytest assertions.

NER PII detectionassert_no_pii(method="ner"|"gliner"|"hybrid") adds Named Entity Recognition to PII scanning. Presidio + spaCy catches names, organizations, and locations that regex cannot find. GLiNER provides zero-shot detection for domain-specific entities (healthcare MRN, legal case numbers). Hybrid mode runs regex + NER with intelligent deduplication. Install with pip install mltk[ner].

YAML test definitions — write ML tests in YAML, no Python required. Run with mltk test tests.yaml. Supports 11 data assertions with env:VAR_NAME data source for CI/CD.

EU AI Act compliance reportsmltk compliance results.json auto-maps test results to EU AI Act articles, classifies risk level (unacceptable/high/limited/minimal), and generates an HTML evidence table with gap analysis.

mltk doctor — 9 diagnostic checks that inspect your Python version, dependencies, config, Rust extension, and pytest plugin, with actionable fix hints for each failure.

Data contracts — define column types, ranges, and uniqueness rules in YAML. Use mltk contract generate-tests to auto-generate a pytest file from the contract — zero boilerplate.

LLM evaluationassert_semantic_similarity, assert_no_toxicity, assert_no_hallucination, assert_ttft, assert_itl, RAG evaluation (faithfulness, context precision/recall, RAGAS composite), agentic testing (task completion, tool selection), multi-turn conversation, coherence, BERTScore, and text quality assertions for GenAI systems.

Training bug detection — catch data leakage (P0), gradient failures (P1), numerical instability, plus P2 bugs: augmentation corruption, checkpoint integrity, distributed training sync, memory leaks, and training-serving skew detection.

Server platformmltk server starts a self-hosted result-tracking server with REST API, live dashboard, API key auth, webhooks, run comparison, and GitHub CI integration (PR comments + check runs).

Cloud monitoring — AWS SageMaker/CloudWatch, GCP Vertex AI, Azure ML, and Prometheus/Triton adapters for production endpoint health, latency, and error rate assertions.

FDA audit trailmltk fda-audit generates a device-grade audit trail from test results for regulatory submissions.

Compliance PDFmltk compliance-pdf exports any HTML compliance report (EU AI Act, OWASP) to PDF for auditors.

Model Card Generatormltk model-card results.json generates a Google-format Model Card in Markdown from test results.

Chat interfacemltk chat provides interactive Q&A about test results with no LLM or external API required.

Environment variable config — all config keys available as MLTK_* env vars, highest priority in the cascade. Works out of the box in CI/CD pipelines.

JSON export--mltk-export-json flag exports full test results to JSON for downstream tooling.

Feature Matrix (241 assertions)

Module Assertions Purpose
Data Schema assert_schema, assert_no_nulls, assert_dtypes Validate DataFrame structure
Data Distribution assert_range, assert_unique, assert_no_outliers Verify statistical properties
Data Statistics assert_column_mean, assert_column_median, assert_column_stdev, assert_quantiles Statistical bounds checking
Data Validation assert_datetime_format, assert_values_in_set, assert_no_conflicting_labels, assert_feature_label_correlation_stable Value-level data checks
Data Drift assert_no_drift (KS, PSI, KL, Chi2) Detect distribution shifts
Data Drift (Advanced) assert_no_embedding_drift + 7 methods (KS, PSI, KL, Chi2, JS, Wasserstein, auto) Embedding + tabular drift
Data PII assert_no_pii, scan_pii (40+ regex patterns + Presidio NER + GLiNER zero-shot + hybrid) Find leaked personal data
Data Freshness assert_freshness, assert_row_count Verify data recency and size
Data Labels assert_label_balance, assert_label_coverage Validate label quality
Data Contracts validate_data, generate_tests_from_contract YAML → auto-test generation
Data Quality Preset assert_data_quality, data_quality_report One-call full data audit
Data Lineage assert_lineage_complete Track data provenance
Model Metrics assert_metric (9 metrics: accuracy, F1, AUC, MSE, R2...) Quality gates
Model Regression assert_no_regression, save_baseline Catch silent degradation
Model Slicing assert_slice_performance, assert_calibration Subgroup fairness + ECE
Model Bias assert_no_bias (5 methods) EU AI Act / four-fifths rule
Model Adversarial assert_robust Perturbation stability
Model Overfitting assert_no_overfitting, assert_label_drift Training vs. validation gaps
Training Bugs (P0) assert_no_train_test_overlap, assert_temporal_split, assert_no_target_leakage Data leakage
Training Bugs (P1) assert_gradient_flow, assert_no_vanishing_gradient, assert_no_exploding_gradient, assert_loss_finite Gradient health
Training Numerical assert_no_nan_inf, assert_loss_decreasing, assert_no_loss_divergence, assert_softmax_valid Numerical stability
Training Bugs (P2) assert_augmentation_preserves_signal, assert_no_augmentation_on_test, assert_checkpoint_complete, assert_resume_loss_continuous, assert_effective_batch_size, assert_gradient_sync, assert_loss_is_detached, assert_no_memory_leak Augmentation, checkpoint, distributed, memory
Training Skew assert_no_training_serving_skew Training-serving distribution mismatch
Inference assert_latency, assert_cold_start, assert_throughput, assert_api_contract Performance SLAs
Pipeline assert_reproducible, assert_checksum, assert_pipeline Determinism + E2E
Monitor assert_no_degradation, assert_sla, assert_no_output_drift Production health
Cloud Monitor AWS (assert_endpoint_healthy/latency/error_rate), GCP (assert_endpoint_healthy, assert_prediction_latency), Azure (assert_endpoint_healthy/latency), Prometheus/Triton Multi-cloud monitoring
LLM Evaluation assert_semantic_similarity, assert_no_toxicity, assert_no_hallucination, assert_ttft, assert_itl GenAI testing
LLM RAG assert_faithfulness, assert_context_relevancy, assert_context_precision, assert_context_recall, assert_answer_relevancy, assert_ragas_score RAG pipeline evaluation (RAGAS)
LLM Agentic assert_task_completion, assert_tool_selection, assert_tool_call_correctness Agent/tool-use testing
LLM Text Quality assert_text_length, assert_output_format, assert_readability Output quality gates
LLM Multi-turn assert_knowledge_retention, assert_turn_relevancy, assert_conversation_completeness Conversation evaluation
LLM Coherence assert_coherence Logical consistency
LLM BERTScore assert_bertscore (Rust-accelerated token matching) Semantic similarity via BERT
NLP assert_bleu, assert_rouge, assert_ner_f1, assert_no_prompt_injection Text + security
NLP Sentiment assert_sentiment_positive, assert_no_sentiment_drift Sentiment analysis + drift
Speech assert_wer, assert_cer, assert_rtf, assert_accent_coverage Recognition + fairness
CV assert_iou, assert_map, assert_frame_accuracy, assert_temporal_consistency, assert_topk_accuracy Video analytics
CV Tracking assert_mota, assert_motp, assert_idf1 Multi-object tracking (CLEAR-MOT)
CV Face assert_face_far False accept rate (NIST FRVT)
Tabular assert_feature_drift, assert_feature_importance_stable, assert_class_balance Feature validation
Compliance EU AI Act report, OWASP LLM Top 10, FDA audit trail, compliance PDF export Regulatory evidence
Retrieval assert_ndcg, assert_mrr, assert_recall_at_k, assert_map_at_k Ranking metrics (RAG)
LLM-as-Judge assert_llm_judge_score, assert_llm_judge_pairwise Vendor-neutral LLM eval
Summarization assert_summary_coverage, assert_summary_compression, assert_summary_faithfulness Summary quality
Recommendation assert_hit_rate, assert_diversity, assert_novelty, assert_coverage, assert_serendipity RecSys testing
Long-Context LLM assert_needle_in_haystack, assert_context_utilization, assert_no_lost_in_middle Long-context eval
Composable Suite MltkSuite, SuiteResult, export to JSON/HTML/JUnit Run without pytest
Code Generation assert_code_executes, assert_code_passes_tests, assert_no_code_vulnerabilities, assert_code_complexity LLM codegen testing
Healthcare HIPAA compliance mapping, assert_hipaa_coverage Regulated industry
Finance SR 11-7 model risk, assert_sr117_coverage Bank model governance
Enterprise RBAC, audit log, custom compliance frameworks SOC 2 + governance
Observability TestScheduler, anomaly detection, impact analysis, Grafana dashboards Production monitoring
Testing Patterns assert_matches_golden, flaky detection, retry, test selection Advanced test strategies
Reports Visual diff, test history summarizer, bias report, model card Analysis + reporting
Integrations Jira, MLflow, GitHub App, W&B, DVC, Kubeflow, SageMaker, Slack, webhooks Ecosystem connectors

Installation

# Core (data + model assertions)
pip install mltk

# With CLI
pip install mltk[cli]

# With HTML reports
pip install mltk[report]

# With server platform
pip install mltk[server]

# NER PII detection
pip install mltk[ner]         # Presidio + spaCy NER

# Domain kits
pip install mltk[cv]          # Computer Vision
pip install mltk[nlp]         # NLP
pip install mltk[speech]      # Speech Recognition
pip install mltk[contracts]   # Data contracts (YAML)
pip install mltk[torch]       # Gradient inspection
pip install mltk[llm]         # LLM evaluation

# Integrations
pip install mltk[integrations]  # Jira adapter
pip install mltk[mlflow]        # MLflow logging
pip install mltk[aws]           # AWS SageMaker/CloudWatch
pip install mltk[gcp]           # GCP Vertex AI
pip install mltk[azure]         # Azure ML monitoring
pip install mltk[pdf]           # Compliance PDF export
pip install mltk[server]        # Self-hosted server platform

# Everything
pip install mltk[all]

CLI

# Core
mltk version                              # Show version
mltk init                                 # Scaffold config + tests
mltk scan data.csv                        # Quick data quality scan
mltk drift ref.csv cur.csv                # Drift analysis
mltk score                                # ML Test Score
mltk doctor                               # Diagnose environment

# Test execution & reporting
mltk test tests.yaml                      # Run YAML-defined tests
mltk model-card results.json              # Generate Google Model Card
mltk compliance results.json              # EU AI Act compliance report
mltk fda-audit results.json              # FDA device audit trail export
mltk compliance-pdf report.html          # Export compliance report to PDF

# Data contracts
mltk contract init                        # Scaffold contract YAML
mltk contract validate data.csv          # Validate against contract
mltk contract generate-tests             # Generate pytest from contract

# Documentation
mltk docs serve                           # Serve docs locally (hot reload)
mltk docs build                           # Build static HTML docs
mltk docs open                            # Build, serve, and open in browser

# Test registry
mltk registry push my_fixtures            # Save test files to registry
mltk registry pull my_fixtures            # Restore from registry
mltk registry list                        # List saved collections

# Notifications
mltk notify slack --results-json r.json  # Send results to Slack

# Server platform
mltk server                               # Start result-tracking server
mltk server-create-key --project myproj  # Generate API key for server

# Interactive
mltk chat --results-json results.json    # Q&A about test results

pytest Plugin

Auto-registered on install. Provides ML-specific markers and --mltk-report:

pytest -m ml_data                 # Run only data quality tests
pytest -m ml_model                # Run only model quality tests
pytest -m "not ml_slow"           # Skip long-running tests
pytest --mltk-report              # Generate HTML report
pytest --mltk-export-json         # Export results to JSON

Markers: ml_data, ml_model, ml_drift, ml_inference, ml_smoke, ml_slow, ml_gpu, ml_nondeterministic

Rust Acceleration

Optional Rust backend for 10-100x speedup on drift detection (KS test, PSI). Falls back to scipy/numpy automatically.

Comparison

Feature Great Expectations Deepchecks Evidently DeepEval Giskard Arize Phoenix mltk
Data validation Yes Yes Limited No Yes No Yes
Model testing No Yes Limited No Yes (tabular) No Yes
Drift detection No Yes Yes No No No Yes (Rust)
Streaming drift No No No No No No Yes (ADWIN/CUSUM)
Bias/fairness No No No No Yes No Yes
Inference testing No No No No No No Yes
pytest native No No No Yes Yes No Yes
CLI Yes No No Yes Yes Yes Yes
Domain kits No Partial No No No No Yes
ML Test Score No No No No No No Yes
Rust acceleration No No No No No No Yes
YAML test defs No No No No No No Yes
Compliance frameworks No No No No No No Yes (8 frameworks)
Training bug detection No No No No No No Yes
Conformal prediction No No No No No No Yes
Composable TestSuite No No No No No No Yes
Code Generation No No No No No No Yes (4 assertions)
LLM evaluation No Yes (LLM-only) No Yes (50+ metrics) Yes (OWASP scanner) Yes (tracing) Yes (241 assertions)
Behavioral consistency No No No No No No Yes (7 assertions)
NER PII detection No No No No No No Yes (4 methods)
Agent trace testing No No No Yes (basic) No Yes (tracing) Yes (9 assertions)
Multi-agent testing No No No No No No Yes
Synthetic data validation No No No No No No Yes (4 assertions)
Safety taxonomy No No No Yes (red-teaming) Yes (OWASP) No Yes (per-category)
LLM observability No No No No No Yes No
OWASP LLM scanning No No No No Yes No Yes
License Elastic AGPL Apache 2.0 Apache 2.0 Apache 2.0 Elastic Elastic 2.0
Core deps Heavy Heavy Medium Heavy (LLM) Heavy Heavy 2 (numpy, pandas)

Documentation

License

Elastic License 2.0 — free for internal use; commercial licensing available at lior1cc@gmail.com. See LICENSE-COMMERCIAL.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mlspec-0.13.0.tar.gz (603.0 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

mlspec-0.13.0-cp312-cp312-win_amd64.whl (1.4 MB view details)

Uploaded CPython 3.12Windows x86-64

mlspec-0.13.0-cp312-cp312-manylinux_2_34_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.34+ x86-64

mlspec-0.13.0-cp312-cp312-manylinux_2_34_aarch64.whl (1.4 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.34+ ARM64

mlspec-0.13.0-cp312-cp312-macosx_11_0_arm64.whl (1.4 MB view details)

Uploaded CPython 3.12macOS 11.0+ ARM64

File details

Details for the file mlspec-0.13.0.tar.gz.

File metadata

  • Download URL: mlspec-0.13.0.tar.gz
  • Upload date:
  • Size: 603.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mlspec-0.13.0.tar.gz
Algorithm Hash digest
SHA256 89eb84848632fa90086de85f32f249564fa9fec1781db9b54ca904d256b2304b
MD5 32c3ecf7b20c2b2cb41d19657f3a228b
BLAKE2b-256 b25d1459090e9f35c4abd9a9e95c3cab5538b7dea32055e9d96f7e034c2d74dc

See more details on using hashes here.

Provenance

The following attestation bundles were made for mlspec-0.13.0.tar.gz:

Publisher: release.yml on Liorrr/mltk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mlspec-0.13.0-cp312-cp312-win_amd64.whl.

File metadata

  • Download URL: mlspec-0.13.0-cp312-cp312-win_amd64.whl
  • Upload date:
  • Size: 1.4 MB
  • Tags: CPython 3.12, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mlspec-0.13.0-cp312-cp312-win_amd64.whl
Algorithm Hash digest
SHA256 8e245775d6105dbc860c1150e12180c687cb8d3b4120f07fffa1dd2c68ca5b7a
MD5 362893f031f727992a1e7df99ebcae96
BLAKE2b-256 a4e9b90877bd83e7f944f4b5fb90c5fd0be0da37bbc667c38041bf41431d9399

See more details on using hashes here.

Provenance

The following attestation bundles were made for mlspec-0.13.0-cp312-cp312-win_amd64.whl:

Publisher: release.yml on Liorrr/mltk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mlspec-0.13.0-cp312-cp312-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for mlspec-0.13.0-cp312-cp312-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 5dee86a46f82648cf93045d244e165a2f7d2353babeccad1f349dd3d5a368d10
MD5 2791f539617d845b101a193b3f825a18
BLAKE2b-256 a9b544ffe221aa99410c406a533b9035d379fdfa2080c61471f3fc9e8f02f963

See more details on using hashes here.

Provenance

The following attestation bundles were made for mlspec-0.13.0-cp312-cp312-manylinux_2_34_x86_64.whl:

Publisher: release.yml on Liorrr/mltk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mlspec-0.13.0-cp312-cp312-manylinux_2_34_aarch64.whl.

File metadata

File hashes

Hashes for mlspec-0.13.0-cp312-cp312-manylinux_2_34_aarch64.whl
Algorithm Hash digest
SHA256 bb603f72084114813cadf6841280009a5e2314ba6a2d3b31a80567b70878cdd2
MD5 3d81b49d07284da7bf4ba258ec407142
BLAKE2b-256 99de2d1d40b2e170c75554e117b467ee24a389840741d77329af7fa9903938a8

See more details on using hashes here.

Provenance

The following attestation bundles were made for mlspec-0.13.0-cp312-cp312-manylinux_2_34_aarch64.whl:

Publisher: release.yml on Liorrr/mltk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mlspec-0.13.0-cp312-cp312-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for mlspec-0.13.0-cp312-cp312-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 ed35eac8c004921b24e3c4ebaf97b219df66d2b6c23608da47b5bfbee989acf9
MD5 79d8872c2889a205d44e60ca167a50b9
BLAKE2b-256 d2abdcefab3978a53f97daa8eb8129d8c2bb3bc86d60e981a995fe40009a968b

See more details on using hashes here.

Provenance

The following attestation bundles were made for mlspec-0.13.0-cp312-cp312-macosx_11_0_arm64.whl:

Publisher: release.yml on Liorrr/mltk

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page