PULSE
Open-source pytest-style compliance testing and evaluation framework for AI voice and chat agents.
⚠️ Technical signal only. Not legal certification. Consult qualified counsel before relying on PULSE results for regulatory purposes.
PULSE is for teams building custom-coded AI voice and chat agents (LiveKit, Pipecat, raw Python/Node) who need to validate compliance (DPDP, RBI KYC, TRAI DLT, HIPAA) and quality (WER, TTFA, barge-in) before shipping. It is not a SaaS tool. Zero cloud, zero dashboard, zero account. Like pytest is for Python testing, PULSE is for agent compliance testing.
Quickstart — pytest-native (primary path)
# test_my_agent.py
from pulse_eval.testing import AgentTestCase, assert_compliant, assert_quality
def test_kyc_flow_hindi():
"""KYC verification in Hindi must pass DPDP + RBI KYC + TRAI DLT."""
case = AgentTestCase(
agent=my_livekit_agent,
scenario="kyc_verification_hindi", # loaded from tests/generated/
)
assert_compliant(case, packs=["dpdp", "rbi_kyc", "trai_dlt"])
def test_voice_quality_hindi():
"""Voice quality metrics must meet minimum thresholds."""
case = AgentTestCase(agent=my_livekit_agent, scenario="kyc_verification_hindi")
assert_quality(case, metrics=["VOICE-013", "VOICE-021"], min_score=0.7)
Run it exactly like any other pytest:
pytest tests/test_my_agent.py -v
Failures include full judge reasoning and the compliance disclaimer, right in the pytest output.
CLI Batch Workflow
For full compliance reports (CI, pre-release audits):
pip install pulse-eval
# 1. Configure PULSE — enter your LiteLLM model strings and API keys
pulse init
# 2. Generate test scenarios from your agent's system prompt
pulse generate \
--prompt interview_agent_prompt.txt \
--domain general \
--language en-IN \
--count 11
# 3. Run the full evaluation
pulse run config.yaml
# 4. Open the HTML report in your browser
# Reports saved to .pulse/runs/<run_id>/report.html
Example terminal output
PULSE Run config.yaml
─────────────────────────────────────────────────────────
Agent: my_agent.handler:respond (generic)
Judge: claude-sonnet-4-6
Scenarios:11
Running 11 scenarios... ████████████████ 100% 12.3s
┌──────────────────────────────────────────────────────┐
│ Pass Rate: 81.8% │ Scenarios: 11 │ Passed: 9 │
│ Failed: 2 │ Critical: 1 │ │
└──────────────────────────────────────────────────────┘
Latency p50=520ms p75=620ms p95=990ms p99=1450ms
⚠️ Critical Failures:
[CRITICAL] DPDP-001 (DPDP Act 2023 §6(1))
Reason: Agent collected phone number before obtaining explicit consent
Turn: 2 | Confidence: 0.91
⚠️ Technical signal only. Not legal certification...
Cohort Breakdown:
hi-IN north-indian 6 scenarios 93.3% WER: 6.2%
en-IN neutral 5 scenarios 90.0% WER: 3.1%
📄 Reports → .pulse/runs/a3f12b9c/
Judge Calibration
PULSE is the first framework in this category to publish a judge agreement rate. Before launching your compliance program, calibrate your judge against human-labeled transcripts:
pulse calibrate labeled_cases/
Output:
Agreement Rate: 87.3% (42 cases) | Cohen's κ: 0.812
Per-Rule Agreement:
DPDP-001 91.2%
RBI-001 85.7%
TRAI-002 88.1%
HIPAA-001 83.9%
📝 README snippet: .pulse/calibration/summary.md
Put the agreement rate in your README. It's your credential — nobody else in this category publishes it.
Calibration note: Agreement rate measures whether the judge's verdict matches human reviewer verdicts. Cohen's kappa corrects for chance agreement. This is not a measure of legal accuracy.
PULSE vs. DeepEval vs. Promptfoo
| PULSE | DeepEval | Promptfoo | |
|---|---|---|---|
| Voice-native metrics (WER, TTFA, barge-in) | ✅ | ❌ | ❌ |
| India compliance (DPDP, RBI KYC, TRAI DLT) | ✅ | ❌ | ❌ |
| HIPAA compliance pack | ✅ | Partial | Partial |
| Fully air-gapped (local models via Ollama) | ✅ | Partial | Partial |
| pytest-native API | ✅ | ✅ | ❌ |
| Zero cloud, zero accounts | ✅ | ❌ (cloud dashboard) | ❌ |
| Judge calibration rate published | ✅ | ❌ | ❌ |
| Open-source, MIT, no paid tier | ✅ | Partial | ✅ |
If your agent is text-only and you don't need India compliance, use DeepEval or Promptfoo — excellent tools. PULSE exists specifically for:
- The voice layer — metrics that require audio timestamps, WER, STT confidence
- India-specific compliance — DPDP Act 2023, RBI KYC, TRAI DLT/DND — none of the others touch this
What PULSE covers
Compliance Assertion Packs
| Pack | Regulations | Key Assertions |
|---|---|---|
dpdp |
DPDP Act 2023 | Consent before PII, purpose stated, right to withdraw, no third-party without disclosure |
rbi_kyc |
RBI Master Circular on KYC | Identity before account info, consent to record, interest rate disclosure, no guaranteed returns |
trai_dlt |
TRAI DLT/DND 2018 | Registered sender ID, consent for outbound, DND registry respected, opt-out honored |
hipaa |
HIPAA Privacy Rule | Identity before PHI, minimum necessary, no PHI in tool calls, no unauthorized diagnosis |
Quality Metrics
Universal (voice + chat): Task Completion, Context Retention, Factual Accuracy, Entity Extraction, Out-of-Scope Handling, Escalation Appropriateness, Tone, Repetition Rate, Persona Consistency, Harmful Advice Detection, PII Leakage
Voice-only: TTFA, Interruption Recovery, Filler Word Rate, Dead Air, Barge-in Handling, Speech Rate, WER (via jiwer), Domain Term Accuracy, Noisy Environment Robustness, Disfluency Handling, Intent Classification Accuracy & Confidence, Multi-Intent Handling, Clarifying Questions, Voice Response Length, Latency Consistency (p50/p75/p95/p99), Tail Latency, Avg Turns to Completion
Chat-only: Response Length, Message Coherence, Formatting Appropriateness, Chat Response Latency
Configuration
config.yaml
agent:
type: voice
adapter: generic # generic (default) | livekit | pipecat
entrypoint: "my_agent.handler:respond"
# For LiveKit:
# adapter: livekit
# livekit_url: "ws://localhost:7880"
# room_name: "test-room"
# api_key: "devkey"
# api_secret: "secret"
judge:
primary_model: "claude-sonnet-4-6" # critical/high assertions
fast_model: "gemini/gemini-3-flash" # medium assertions
# Fully local/air-gapped:
# primary_model: "ollama/llama3.1"
# fast_model: "ollama/llama3.1"
simulator:
model: "gemini/gemini-3-flash"
persona_style: realistic # realistic | adversarial | polite
metrics:
enabled: true
voice:
ttfa_warn_ms: 1500
ttfa_fail_ms: 3000
wer_warn: 0.15
wer_fail: 0.30
scenarios:
- name: "kyc_verification_hindi"
language: hi-IN
accent: north-indian
domain: fsi
max_turns: 9
expected_intent: "home_loan_eligibility_inquiry"
domain_terms: [EMI, CIBIL, NACH, KYC, PAN]
persona: "first time home loan applicant, nervous, speaks Hindi, salaried employee"
assertions:
- rbi_kyc.identity_before_disclosure
- rbi_kyc.interest_rate_disclosed
- dpdp.consent_before_pii
- dpdp.purpose_stated
- trai_dlt.consent_for_outbound
metrics:
- VOICE-013 # WER
- VOICE-021 # intent accuracy
- VOICE-031 # latency consistency
Bring Your Own Model (BYO)
PULSE uses LiteLLM to route ALL LLM calls. Any LiteLLM-supported model string works:
# Anthropic
primary_model: "claude-sonnet-4-6"
# Google
primary_model: "gemini/gemini-2.5-pro"
# OpenAI
primary_model: "gpt-4o"
# Fully local — air-gapped mode (customer transcripts never leave your infra)
primary_model: "ollama/llama3.1"
fast_model: "ollama/llama3.1"
Air-gapped mode (Ollama/vLLM) is directly relevant to DPDP and RBI data-localization requirements — customer transcript data never leaves your infrastructure.
Adapters
| Adapter | Use Case | Dependency |
|---|---|---|
generic (default) |
Any custom-coded agent | None (always available) |
livekit |
LiveKit voice agents | pip install 'pulse-eval[livekit]' |
pipecat |
Pipecat pipeline agents | pip install 'pulse-eval[pipecat]' |
Generic Adapter — Mode A (function call)
# my_agent/handler.py
def respond(text: str) -> str:
return your_agent_logic(text)
# config.yaml
agent:
adapter: generic
entrypoint: "my_agent.handler:respond"
Generic Adapter — Mode B (HTTP/WebSocket)
agent:
adapter: generic
entrypoint: "ws://localhost:8000/agent" # WebSocket
# or: entrypoint: "http://localhost:8000/agent" # HTTP POST
Protocol: {"role": "user", "content": "..."} → {"role": "agent", "content": "..."}
Dogfood: Mock Interview Voice Agent (LiveKit)
PULSE was built alongside a Mock Interview voice agent running on LiveKit. This is the primary internal validation target:
# tests/test_interview_agent.py
from pulse_eval.testing import AgentTestCase, assert_compliant, assert_quality
from my_agents.interview_agent import InterviewLiveKitAgent
def test_interview_agent_compliance():
agent = InterviewLiveKitAgent(room="test-room")
case = AgentTestCase(
agent=agent,
scenario={
"name": "mock_interview_en_in",
"language": "en-IN",
"domain": "general",
"max_turns": 12,
"persona": "fresh graduate applying for software engineer role",
}
)
# Checks tone, context retention, no harmful advice
assert_compliant(case, packs=["dpdp"])
assert_quality(case, metrics=["VOICE-013", "VOICE-005", "VOICE-001"], min_score=0.7)
Convert Production Failures to Tests
pulse convert-failure failed_call_2026_07_31.json
# ✅ Regression test created: tests/regression/failed_consent_flow_fsi.yaml
# It will run automatically on every future pulse run.
Local Storage
.pulse/ ← NEVER commit (in .gitignore)
config.toml ← API keys
runs/
<run_id>/
report.json
report.html
transcripts/
<scenario_name>.json
calibration/
report.json
summary.md
No database. No SQLite. No cloud. Pure local files.
What's NOT in PULSE
- ❌ Production monitoring — PULSE is pre-deployment testing only. Monitoring is a separate product with different infrastructure requirements.
- ❌ PDF reports — Planned for v2 once calibration data exists to back "audit-ready" claims.
- ❌ Vapi / Retell / Bland.ai — Low-code platforms, wrong side of the "coded, not low-code" filter.
- ❌ Web dashboard, accounts, payments — Zero, forever, MIT.
License
MIT. No gated packs, no license key, no paid tier. PULSE is free, forever.
The business model is trust — built here, monetized elsewhere in a separate downstream product if at all.
⚠️ Disclaimer: Every compliance assertion result carries the label "Technical signal only. Not legal certification. Consult qualified counsel before relying on this result for regulatory purposes." This is not a footnote — it appears on every result in every output format.
Metadata
Release files for pulse-eval 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pulse_eval-0.1.0.tar.gz | 57.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pulse_eval-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 119.5 kB
Release files / pulse_eval-0.1.0.tar.gz
| Download URL | pulse_eval-0.1.0.tar.gz |
|---|---|
| Size | 57.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
41ff7ca6ed0e5f709db4506679c4ef01fdc8b495088aa90e300f9d521588d204
|
|
BLAKE2b-256 checksum How to use checksums |
340b3214bb0dce5e955a7c5fe0e81da3278bff7115e4a5a4478011416adf06c9
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.13 {"installer":{"name":"uv","version":"0.12.13","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / pulse_eval-0.1.0-py3-none-any.whl
| Download URL | pulse_eval-0.1.0-py3-none-any.whl |
|---|---|
| Size | 61.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
101d14f53ee134cc74031aaf5de507ecb278acc553f8af4f351c1b2b84daa02d
|
|
BLAKE2b-256 checksum How to use checksums |
f8219e128d60bfec4fb9ad87e8ade33fed3f4c54f00b3f7f9b4383a1583d3744
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.13 {"installer":{"name":"uv","version":"0.12.13","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|