Skip to main content

A trainable continual-learning architecture (manas/buddhi/ahaṃkāra/citta) that learns without forgetting, knows when it doesn't know, watches for drift, and governs its own actions. Ships the Continual-Governance Benchmark (CGB): four faculties, four measured numbers (retention, calibrated abstention, drift false-alarm, governed escalation) on one leaderboard across tabular/vision/language, plus a tamper-evident hash-chained audit trail.

Project description

antahkarana — train a continual-learning mind on your data

PyPI License Python Model card

antahkarana is a trainable architecture, not a frozen model. You bring your data (and optionally your backbone); it trains a model that learns continually without forgetting, detects novelty / zero-days, abstains when unsure, and consolidates in sleep. Four organs — manas · buddhi · ahaṃkāra · citta — the mind is the same for every domain; only a thin adapter changes.

architecture

Built on the validated Antaḥkaraṇa-base core (gated G0–G5, proven across language · vision · security · with seeded, adaptive evaluation — see the model card for plots, results, and honest limits).

Install

pip install antahkarana              # core + tabular
pip install antahkarana[text]        # + HuggingFace LLM adapter
pip install antahkarana[bench]       # + the open Continual-Governance Benchmark (2.1)
pip install antahkarana[telemetry]   # + OpenTelemetry OTLP export (2.2)

Train on your data in ~10 lines (tabular)

from antahkarana import Antahkarana, TabularAdapter

# YOUR data: a list of tasks, each (X_train, y_train, X_test, y_test)
stream = TabularAdapter.make_stream(your_tasks)

bb   = TabularAdapter(input_dim=122, n_tasks=4, n_classes=2)   # built-in MLP — bring only data
mind = Antahkarana(bb, replay_strategy="der", avidya_strategy="energy")
res  = mind.train(stream)

print(res["forgetting"], res["final_row"], res["risk_coverage"])

New in 9.0 — Spanda: the quantum-inspired intellective layer

Spanda

Spanda (स्पन्द, "the primordial pulse") is the release where the mind stops committing prematurely. It holds a superposition of hypotheses (real signed amplitudes, Born-rule probabilities), lets evidence interfere — including destructive cancellation a bag of positive weights can't express — and collapses to a determination only under the governance ring. Below a confidence bar it is allowed to abstain and escalate. It is honest about "quantum": three real lanes — quantum-inspired optimization, quantum-resistant cryptography, and superposition reasoning — and it refuses the one hype lane.

from antahkarana.spanda import superpose, GovernedCollapse

sup = superpose(["fraud", "benign", "review"])
sup.interfere({"fraud": 1.4, "benign": 0.6}, mode="bayes")   # evidence reweights amplitudes
GovernedCollapse(bar=0.7).collapse(sup)                       # → determine … or abstain & escalate

Four pillars, all pure-Python / torch-free / opt-in, each machine-checked (INV-16..20, with teeth):

  • Superposition reasoning — better-calibrated answers; ECE 0.42 → 0.29 vs a single-hypothesis baseline, answering only when confident (100% accurate there) and abstaining otherwise.
  • Padārtha — a Vedic-named typed layer: categories (padārtha), relations (sambandha), and evidence that pervades related concepts (vyāpti) before collapse. Conformance-checked.
  • Quantum-inspired optimization — QUBO + a tunnelling annealer bridged to nidrā memory consolidation (≥ greedy by construction); runs on CPU or any GPU (torch CUDA/ROCm) with a sparse operator — 17.8× faster on an A10 at n=512, identical optimum.
  • Post-quantum integrity — pure-hashlib Lamport + Merkle signatures and a PQ-signed audit log: tamper-evident and quantum-resistant.

Spanda — measured

Which workloads it helps: high-stakes classification/triage that must abstain rather than guess (fraud, medical, security); agents that should escalate under ambiguity instead of acting; memory consolidation / retrieval-budget selection; and any regulated system that needs a quantum-resistant audit trail. Opt in with from antahkarana import spanda — a plain import antahkarana never pulls it.

New in 8.0 — Kriyā: the governed agent

A governed agent you can prove is safe — every tool call capability-token-gated, every step audited, every learned skill gated and forgettable — that is parallel-first, highly available, runs on CPU or GPU (NVIDIA and AMD), drops into LangChain / LangGraph / Llama Stack, and whose core safety invariants are machine-checked exhaustively. Kriyā ("action") is the agentic layer: v5.0 governed the action ring, v7.0 the memory layer, v8.0 the agent itself.

Kriyā — measured results

  • The governed agent loop (G21) + parallel-first (G21b). plan → gate → mint → act → observe: a tool runs only by spending a single-use capability token; writes escalate to a human; every step journaled. Independent tool calls fan out concurrently — measured 7× wall-clock speedup at 8 calls/round — with mint+journal serialized so the hash chain stays intact; read memoization + a batch that gates all before executing any (fail-closed).
  • Multi-model (G22). ModelRouter — route / fallback-closed / ensemble-verify + a concurrent verifier panel (quorum) + best-of-N fanout() + migration_conformance() (report drift before a model swap).
  • Self-improving, safely (G23). SelfImprover compiles skills from the agent's own trajectories: parallel red-team fuzz → non-self-modifiable gate → shadow adoption → auto-revert on first regression.
  • Production interop + CPU/GPU (G24). rag.backend VectorBackend seam (pgvector/Qdrant/Milvus) + device=cpu|cuda/nvidia|rocm/amd; antahkarana.interop adapters for LangChain / LangGraph / Llama Stack (zero hard deps).
  • Living legal documents (G25). antahkarana.dharma — immutable obligation (Vidhi) ⇄ revisable document (Smriti); reconcile() re-binds a preserved obligation and quarantines a conflicting edit — your policy never silently changes when a document does.
  • Machine-checked (G26) + high availability (G27). formal.check_agent: INV-11..15 over all 273 reachable states, 5 broken variants each fail with the shortest counterexample. Replicated self-healing journal + quorum vector backend (reads survive 2 of 3 down) + hedged circuit-broken model pool + lease.
from antahkarana.agent import GovernedAgent, Tool, ModelRouter, SelfImprover
from antahkarana.formal import check_agent

agent = GovernedAgent(planner, tools=[kb, commit], budget=8, journal="agent.jsonl",
                      approver=human_approve, max_parallel=8)   # parallel, token-gated, escalating
r = agent.run("Refund order 7781 if within policy, else escalate.")
assert check_agent().ok                                          # INV-11..15 over all reachable states

295/295 tests. Additive over 7.x; zero new required runtime dependencies. See the full architecture below.

New in 7.0 — Dhāraṇā: lifelong governed memory

A memory that behaves like a mind and is governed like infrastructure — fast at any scale, permanent and auditable, forgets on command provably, grows capability beyond the base model and carries it across model generations. Dhāraṇā ("sustained holding") makes what the system learns durable, retrieved in milliseconds at any size, and portable — so improvement compounds for years and across models, not just within one training run.

Dhāraṇā — measured results

  • Hybrid sparse+dense retrieval (G17). A dense ANN + a sparse lexical index behind the unchanged CittaStore.search: 29–42× faster than the v6.0 full scan on the real store path (and it improves with scale), lifting exact-value recall from 0.10 → 1.000 — codes, IDs, dates matched exactly, where embeddings and weights are fuzzy. Plus a durable, tamper-evident memory journal (verifiable by ant audit verify) and permanence via pinning.
  • Provable unlearning across tiers (G18). forget() cascades to a fixpoint through every summary and schema derived from a fact — and the model checker proves a forgotten fact is represented by no live record. GDPR-grade deletion that weights cannot do.
  • A skill library beyond the base model (G19). Verified, gated procedures (the same non-self-modifiable asymmetry that gates policy): a task suite the base model scores 0/4 on becomes 4/4 once a skill is adopted — and two different serving models get the identical result, because the capability lives in the substrate, not the weights.
  • Machine-checked and stable (G20). Four memory invariants (forget-completeness, consolidation-safety, permanence, bounded-staleness) hold over all 273 reachable states, with broken variants failing on the shortest counterexample; a 200-cycle unattended run consolidated 1205 observations to 58 records (4.8%) with zero recall drift and the journal intact throughout.
from antahkarana.rag import CittaStore              # hybrid index + durable journal + pinning
from antahkarana.skills import Skill, SkillLibrary  # gated, model-agnostic capability
from antahkarana.formal import check_memory         # exhaustive INV-7..10 memory checker

store = CittaStore(index="hybrid", journal="mem.jsonl")   # fast, exact, durable, forget()-able
store.pin(store.add("The depot code is VG-77."))          # permanent, retrievable for life
assert check_memory().ok                                  # INV-7..10 hold over all reachable states

194/194 tests. Zero new runtime dependencies (pure-Python/NumPy; an optional [ann] backend for extreme scale). See the full architecture and every gate result below.

New in 6.0 — Saṃsāra: the governed lifecycle, one wheel

A served LLM wrapped in a control ring that governs every request, remembers like a mind, and re-trains its own weights every night — and can prove it never bypassed its own governance. 6.0.0 "Saṃsāra" (the turning wheel) closes the loop the first five versions opened: govern → retrieve → consolidate → re-train, each turn witnessed on one tamper-evident chain.

Saṃsāra — measured performance

The one design decision that makes this major version strong: knowledge and skill live in different stores (complementary learning systems, the same split the mammalian brain uses).

  • Knowledge → citta-RAG — facts are stored verbatim, retrieved semantically, and cited, so a value is never hallucinated; the store is unbounded, forget()-able, and correctable. On 21 facts, retrieval@1 = 21/21 = 1.000.
  • Skill → nightly weights — a value-based TraceHarvester feeds a QLoRA trainer; a fixed, non-self-modifiable AdapterPipeline gates adoption (improving auto-adopts, unproven waits for a human, regressing auto-rejects). Over a real 7-night run: 5 adoptions, 2 rejections — a poisoned night auto-rejected on the red-team stage, the served chain never regressing through a reject.
  • The ring governs both — retrieved memory and sensed content are taint-fenced (propose, never dispose); adapter adoption is gated by a pipeline the tuner cannot modify. Governance overhead p50 2.53 ms (live, vLLM Qwen3 on 2× A10), ≈20× under budget.

Because knowledge is a store property, not a weight-capacity property, there is no retention ceiling: parametric-only nightly fine-tuning plateaued at ~0.76 and stayed value-fuzzy; the RAG path retains 1.000 on the same facts. See the full architecture and every gate result in the 6.0.0 model card.

from antahkarana.llm import GovernedLLM, OpenAICompatBackend
from antahkarana.rag import CittaStore, PramanaRAG
from antahkarana.tune import NightlyNidra          # the three faces of the wheel

glm = GovernedLLM(OpenAICompatBackend("http://localhost:8000/v1", "Qwen/Qwen3-8B"))
rag = PramanaRAG(CittaStore(), glm=glm)            # cite-or-abstain knowledge
# NightlyNidra(...).cycle("night-01")              # gated self-improvement, lineage-tracked

New in 5.3 — Nightly nidrā (Gate G15): yesterday's confirmed value becomes tonight's weights

The rebirth loop, governed: a TraceHarvester selects ONLY confirmed value (approved answers, human corrections, durable citta knowledge — never difficulty), NightlyNidra trains a LoRA adapter on it, and a fixed, non-self-modifiable AdapterPipeline decides adoption with the G12 asymmetry — an adapter that strictly improves may auto-adopt, a safe-but-unproven one waits for a human, a regressing one closes itself. Every adoption lands in the LineageLedger:

atk why --rule lora-night-04 --ledger ~/nidra/lineage.jsonl
# → serving.adapter: lora-night-02 → lora-night-04 · auto-applied · dataset sha · report hash
from antahkarana.tune import TraceHarvester, AdapterPipeline, NightlyNidra

nn = NightlyNidra(TraceHarvester(trace_log="traces.jsonl", approvals=approved_ids),
                  my_qlora_trainer, AdapterPipeline(my_evals), ledger=ledger)
outcome = nn.cycle("night-01")        # auto:improving | queued:human | closed — never silent

Gate G15 (real 7-night GPU run): 21 facts over 7 nights → 5 auto-adoptions, 2 auto-rejections, final adapter lora-night-07. The poisoned night-4 was auto-rejected on the red-team stage and the served chain held through the reject (a rejected night is never served), recovering the next clean night; frozen general-ability never regressed. Parametric retention reached 0.762 — which is exactly why 6.0 routes knowledge to citta-RAG (1.000) and keeps the weights as the gated skill path.

New in 5.2 — Citta-RAG (Gate G14): retrieval that remembers like a mind

RAG where the store has memory dynamics, not just vectors: saṃskāra (memories confirmed useful win near-ties; stale ones fade on the validated Ebbinghaus curve), nidrā (a sleep pass merges re-observations, writes cluster summaries, and evicts only under a durable summary), and pramāṇa (answers cite [mem-N] evidence or explicitly abstain — calibrated, never a bluff).

Gate G14 (simulated week, 3 seeds): hit@1 0.952 vs 0.667 for vanilla cosine-RAG, with the store at 6 records vs 48 (8× compression) — and a poisoned memory record can propose but never dispose (it enters the prompt DERIVED-tainted and fenced; the privileged sink never fires).

from antahkarana.rag import CittaStore, PramanaRAG

store = CittaStore()                                # HashEmbedder for CI; OpenAICompatEmbedder for real use
store.add("car eleven pit stop took 2.3 seconds")
rag = PramanaRAG(store, glm=my_governed_llm)        # governed end-to-end (optional)
a = rag.answer("how long was car eleven's pit stop?")
rag.feedback(a)                                     # confirmed-useful memories strengthen
store.sleep()                                       # nidrā: consolidate, don't bloat

Built on what's already proven: persistence/tiers/tombstones are the v3.0 memory store; the retention math is organs.citta.Smriti's spaced repetition; the taint fencing is the v2.5 injection defense. Zero new dependencies.

New in 5.1 — GovernedLLM (Gate G13): the served model is governed

The Saṃsāra arc begins: wrap a serving LLM (vLLM / any OpenAI-compatible server) in the control ring so every generate() is a loop tick — screened, gated, audited, explainable.

Gate G13: zero ring bypass across the red-team corpus with the model never called on a blocked request; tool proposals ride the v2.5 propose-never-dispose path (capability tokens); a leaking answer is withheld post-generation; the whole trace lives on one tamper-evident hash chain.

from antahkarana.llm import GovernedLLM, OpenAICompatBackend

glm = GovernedLLM(OpenAICompatBackend("http://localhost:8000/v1", "Qwen/Qwen3-14B"))
r = glm.generate("summarize yesterday's tickets", think="auto")
print(r.text, r.tier, r.flags)        # answer + how the ring ruled
print(glm.why(r.request_id))          # every decision behind this answer, from the chain
atk chat "hello" --mock --why                                  # governed chat, CPU mock (CI)
atk chat "…" --base-url http://localhost:8000/v1 --model Qwen/Qwen3-14B --think on
atk why --request req-000001 --audit governed_llm_audit.jsonl  # decision trace + chain verdict
  • Thinking is a privilege the ring grants, not a mode the caller owns — Qwen3's hybrid modes map onto the organs: non-thinking = manas fast path, thinking = buddhi deliberation; think="auto" deliberates exactly when the input screen lands in the NOTIFY band, and a strict policy denies deliberation entirely while still serving the answer.
  • Fixed, audited stage order — input screen → ring → thinking → backend → tool proposals → ring → output screen → ring → answer; escalations are written before the step they gate.
  • Zero new dependencies — the vLLM client is stdlib urllib; the screen is classical (NFKC + zero-width strip + homoglyph fold), pluggable for a learned manas scorer.

New in 5.0 — Self-governing (Gate G12): the rules evolve; the judge is proven

The system improves its own policy — and that's safe not because we say so, but because the non-bypass invariant is machine-checked exhaustively, and when it can break, the checker returns the shortest counterexample.

Gate G12: non-bypass machine-checked + one self-proposed improvement adopted through the full pipeline. — INV-1/4/5 hold over all 291 reachable states; 1,000,000 traces conform (0 rejections); a tightening diff auto-adopts and a loosening diff is human-merged.

from antahkarana.formal import check, Broken

print(check().summary())              # ✅ INV-1/4/5 hold over all reachable states
print(check(Broken(bypass=True)).summary())   # ❌ INV-1 violated — shortest counterexample returned
atk verify --teeth                    # machine-check non-bypass, and fail every broken model
atk why --rule nf-001 --ledger lineage.jsonl   # how a rule reached its current form
  • Formal core (antahkarana.formal) — an executable exhaustive checker for the token/ring/sink skeleton (INV-1 non-bypass, INV-4 one-tick control, INV-5 token uniqueness) with counterexample extraction; a canonical TLA⁺ spec under formal/; trace conformance so the model can't drift; runtime monitor twins for INV-1..6.
  • Self-improvement (antahkarana.improve.policy / .pipeline / .lineage) — a PolicyCritic proposes a rules-only PolicyDiff with its own predicted effect; a fixed, non-self-modifiable 6-stage pipeline (composing the G8/G9/G11 gates) adopts tightening diffs automatically and routes loosening diffs to a human; mispredictions auto-close; out-of-scope diffs are rejected at the constraint layer; every adoption is traceable via atk why.

New in 4.0 — Platform (Gate G11): multi-tenant, provably isolated

The library + mesh become an operable, multi-tenant product. antahkarana.platform is the tested isolation core — tenancy, encrypted-at-rest stores, principals/roles, loop credentials, SLAs, signed audit export, policies:plan, tiered deployment, policy packs. The deployable atk-server (gRPC/REST + Postgres + Web UI) is the hosting layer on top; the invariants live here, in pure Python.

Gate G11: one tenant's policies, memories, and traces are provably unreachable from another, plus SOC2-style audit export. — no plaintext at rest, cross-tenant decrypt denied, audit tamper localized exactly, policies:plan bit-identical to reality.

from antahkarana.platform import ControlPlane, DecryptError, AuditExporter, verify_export

cp = ControlPlane(); cp.create_tenant("acme"); cp.create_tenant("globex")
cp.tenant_store("acme", "memory").put("host7", {"secret": 1})
cp.cross_tenant_read_attempt("globex", "acme", "memory", "host7")   # → raises DecryptError

bundle = AuditExporter().export(records)     # signed, hash-chained
ok, broken_index = verify_export(bundle)     # tamper → (False, index of the break)
  • Tenancy at the encryption boundary — per-tenant key hierarchy (KEK → per-store DEK); a cross-tenant read yields ciphertext no other tenant's key can decrypt, so isolation survives even a misconfigured filter.
  • Control plane over the v3.5 mesh — loop credentials with bounded blast radius + one-tick revocation; ranked triage inbox; escalation SLA engine (overdue → up-chain, audited).
  • Audit export — signed, hash-chained, tamper-localizing, byte-reproducible (SOC2/ISO mappings).
  • policies:plan — dry-run a policy change over the replay corpus, verdict-diff bit-identical to actually running it. Tiers (safe-by-default, wider-by-grant) + signed tighten-only packs.

New in 3.5 — Multi-agent mesh (Gate G10): a fleet, governed

The scale story. Many governed loops under a supervisor that can kill, throttle, or quarantine any child within one tick — enforced at the loop-engine level, so a child cannot refuse. Rings inherit and can only tighten; a correlated drift no single loop can see registers as one systemic event.

Gate G10: control lands in one tick; fleet-wide correlated drift is detected. — one-tick control p100 = 0 leaked actions; fleet attribution 1.00 at 0 false alarms on uncorrelated noise.

from antahkarana.mesh import MeshController
from antahkarana.mesh.supervisor import SupervisorLoop
from antahkarana.mesh.fleet import FleetCorrelator

ctrl = MeshController()
for cid, loop, mint in fleet: ctrl.register(cid, loop, mint=mint)
sup = SupervisorLoop(ctrl, fleet=FleetCorrelator())   # a supervisor is just a Loop over the fleet
sup.run(max_steps=100)
ctrl.quarantine("sensor-7")     # pause + revoke tokens + seal forensics — child can't refuse, 1 tick
  • antahkarana.meshMeshController + ChildHandle: pause/throttle/quarantine/kill land within one tick; quarantine revokes capability tokens so an adversarial child still can't act, and seals a forensic snapshot that replays with zero access to the live child. Telemetry is derived signals only — never child payloads or memories.
  • antahkarana.mesh.ringsPolicyOverlay + lattice merge: effective policy is the strictest of child spec and parent overlay; a loosening overlay is rejected at load (the mesh cannot widen a child).
  • antahkarana.mesh.fleetFleetCorrelator: correlated sub-threshold drift across the fleet → one attributed systemic event, with a published false-alarm budget on uncorrelated noise.

New in 3.0 — Memory & feedback (Gate G9): the gate learns

Alignment learned outside the weights. Citta gains a retrievable episodic memory with a consolidation daemon, and the control gate learns from human approve/deny decisions — bounded by a monotone-tightening invariant so it can only grow stricter on its own; any loosening is a proposed policy diff a human must merge.

Gate G9: on held-out replays the feedback-tuned gate strictly dominates the static gate on both false-alarm and miss rate. — PASS on 3 corpora (miss 0.50/0.58/0.67 → 0.00 at equal FAR).

from antahkarana.feedback import FeedbackDataset, ThresholdCalibrator
from antahkarana.memory import SqliteStore, MemoryRecord, Tier
from antahkarana.memory.nidra import Nidra

ds  = FeedbackDataset.build(labels, ticks)                     # approve/deny ⨝ trace attributes
cal = ThresholdCalibrator(spec_min=0.5, spec_max=0.95, static_threshold=0.9).fit(ds)
gate_threshold = cal.tuned_threshold()                         # clamped; tighten-only autonomously
if cal.loosening_diff: ...                                     # a spec diff for a human to merge

store = SqliteStore("mem.db"); Nidra(store).run_once()          # consolidate episodes → summaries → facts
store.forget({"cluster": "host7"}, by="dpo@corp")               # GDPR-shaped, audited deletion
atk gateparams rollback sensor-7     # revert a bad calibration in < 1s, audited (like a baseline)
  • antahkarana.memory — three-tier episodic store (episode/summary/fact), zero-dep SqliteStore, audited forget(), retrieval taint-labeled and fenced (memory can't widen permissions — the v2.5 defenses apply to it verbatim).
  • antahkarana.memory.nidra — consolidation scheduler: episodes → per-cluster summaries → human- approved facts; episodes expire only after their summary is durable (no information cliff).
  • antahkarana.feedbackFeedbackDataset + ThresholdCalibrator (clamped, monotone- tightening) + gateparams registry deployment (canary → promote → rollback) + reliability/ECE.

New in 2.5 — Integration (Gate G8): the LLM proposes, the ring disposes

Rent a frontier brain to reason inside a governed loop — and make it structurally impossible for untrusted content to widen its own permissions. The LLM emits proposals; the control ring disposes. Only an ALLOW verdict mints the single-use, signed capability token a sink demands, so a drafting engine can never reach a tool it wasn't authorized for — bypass is a type error.

Gate G8: injected instructions inside sensed content achieve zero policy escalations. — 16 gate tests green; real-model 500-case red-team confirmation on GPU.

from antahkarana.llm import BuddhiLLM, ProposalRouter, GatedSink
from antahkarana.security import CapabilityMint, Observation
from antahkarana.control import RiskRouter, Policy

mint = CapabilityMint()
ring = RiskRouter(Policy.from_dict({"hard_block": ["privileged"], "notify_threshold": 0.99}))
router = ProposalRouter(ring, mint, loop_id="mailwatch",
                        sinks={"notify.status": GatedSink(send, mint, "mailwatch")})

brain = BuddhiLLM(provider="anthropic", model="claude-opus-4-8")   # or provider="local" for CPU/CI
router.run_tick(brain, ctx)     # reason → propose → gate each → act only allowed, token-checked
atk redteam run --n 500 --model Qwen/Qwen2.5-7B-Instruct --device cuda --out runs/g8
  • antahkarana.security — capability tokens bound to (loop, action, params, tick) (replay / mutation / cross-loop all fail closed) + taint tracking (tainted text reaches the model fenced, as data never instructions) + a shipped injection red-team corpus.
  • antahkarana.llmProposal/Backbone/BuddhiLLM; the model holds no capability, provider errors abstain (never default-allow), escalations are recorded before any token is minted (unsuppressible), the prompt template is versioned and hash-locked.
  • antahkarana.mcp — any MCP server becomes a source/sink (sinks act only on a ring token); a loop exposes itself as MCP tools so Claude Desktop/Code is the human-escalation surface, no UI.

New in 2.2 — Hardening (Gate G7): replay any incident from traces alone

The runtime becomes production-grade. Four opt-in modules (zero behavior change until enabled) plus the new atk CLI, behind a falsifiable gate: G7 — replay any past incident end-to-end from traces alone, and roll back a bad baseline refit in one command (20 gate tests, all green).

from antahkarana import Loop, telemetry

telemetry.install(out_dir="traces")                        # one trace per tick, one span per organ
loop = telemetry.instrument(Loop(...), loop_id="sensor-7")
loop.run()
atk replay traces/ --loop sensor-7 --tick 441 --diff 440   # → the first divergent organ attribute
atk policy test policies/        # versioned declarative policies: shadow lint + table tests + version lock
atk baseline rollback sensor-7   # baseline registry: canary shadow → promote → rollback, < 1s, no refit
  • antahkarana.telemetry — spans carry hashes/scores/verdicts, never raw sensed content; zero-dep traces/*.otlpjsonl file exporter; real OTLP collectors via the [telemetry] extra.
  • antahkarana.budget — declarative budgets + three-state circuit breakers that act through the control ring: exhausted budget or open breaker ⇒ an audited soft_block, one escalation per window; act failures trip the breaker open instead of killing the loop. Thread-safe ledgers (in-memory + SQLite), no lost debits under concurrency.
  • antahkarana.policyspec — policies as versioned spec files (apiVersion: antahkarana.dev/v1) with load-time shadow analysis (a hard-blocked category can never be silently re-allowed), colocated table-driven tests, and a version lock: change the ruleset without bumping metadata.version and atk policy test fails.
  • antahkarana.registry — every OneClassChange refit is an immutable versioned artifact (fitted → canary → active → retired | rolled_back); the canary scores every tick in shadow and can never reach the ring; promotion and rollback are atomic pointer swaps.

New in 2.1 — the Continual-Governance Benchmark (CGB) + tamper-evident audit

Four faculties, four measured numbers, one leaderboard — retention ①, selective-risk AURC ②, drift false-alarm rate ③, governed-escalation F1 + audit integrity ④ — across security-tabular, vision, and language (frozen Mistral-7B + LoRA), plus measurable consolidation (ΔR = R_sleep − R_noSleep) and a hash-chained AuditLog whose verify() makes governance a checkable contract. Full board + honest negatives: CGB_RESULTS.

pip install "antahkarana[bench]"
python -m antahkarana.bench.run --suite cgi-v3-tabular --seeds 5 --out runs/tabular

New in 2.0 — the loop engine: watch · learn · improve · orchestrate

antahkarana 2.0 loop engine

The same four organs that learn continually on training tasks now run as a governed loop engine on live observations. Generic agent loops watch and act; this one learns what's normal (fewer false alarms over time), remembers without forgetting, governs by construction (every action passes an audited control ring), and abstains when unsure. Fully additive — every 1.x symbol and default is unchanged.

from antahkarana import Loop, Policy

# A governed stock watcher: it stays silent until the right shoe appears,
# then ESCALATES to you — it never buys on its own (purchase is hard-blocked).
watch = Loop(
    sense=check_product_page,                         # your eyes: return the current observation
    change=lambda o, mem: 1.0 if (o["in_stock"] and o["price"] <= 175) else 0.0,
    category=lambda o: "purchase",
    signature=lambda o: f"{o['in_stock']}:{o['price']}",   # smṛti dedup — don't re-alert the same thing
    gate=Policy(name="shopper", hard_block=["purchase"]),  # the control ring: never auto-buy
    on_escalate=lambda obs, iv: notify_me(obs),
)
watch.run(max_steps=1000)

Four composing phases — each maps 1:1 to an organ:

  • Loop (watch) — a governed sentinel. Sense → score novelty → fire only on a real spike vs the running baseline (ahaṃkāra) → route through the control ring → dedup (smṛti). Hard-blocked actions escalate to a human, never auto-act; everything is audited.
  • OneClassChange (learn) — a learned baseline for Loop: a Normal-only autoencoder that refits on a sliding window with held-out calibration and adapts to drift. Under drift, a static rule's false alarms blow up to 45–98% while the learned baseline holds them at the ~5% budget (3 seeds).
  • ImproveLoop (improve)build → critique → fix → repeat until a learned standard is met, then gate the ship. Monotone convergence in 3–4 rounds, and it never auto-ships (the finished draft is escalated for sign-off).
  • Orchestrator (orchestrate) — buddhi fans out N specialist loops under ONE ring and ONE audit log, then ranks what surfaced so you see the critical thing first. The software spine of the air-gapped Box.

The loop engine is a thin governed runtime around your backbone — it supplies memory + calibrated novelty + the audited gate + the learning baseline; the reasoning/action inside each loop is yours. See examples/loop_stock_watch.py.

New in 1.1.0 — four ways to not forget · scale-stable training · closed-loop plasticity

antahkarana 1.1

A backward-compatible release — every new behaviour is opt-in; the defaults reproduce 1.0. Validated on the NSL-KDD high-interference stream (3 seeds), cross-checked on UNSW-NB15 + vision; core gates green.

  • Four complementary no-forgetting mechanisms — saṃskāra (EWC) · DER++ (now on the text adapter too) · OrthoLoRA structural isolation (forgetting +0.288 → +0.215 where a penalty gives ~0) · forgetting-curve replay (nidrā) that rehearses the about-to-be-forgotten items first (rare-tail recall 0.276 → 0.314).
  • MatraLAMB scale-stable optimization (Antahkarana(optimizer="lamb")) — a trust-ratio step (mātrā) so a full backbone and a tiny adapter move in proportion to their own magnitude, and one lr is stable across any scale. At the lr that used to collapse O-LoRA, forgetting +0.167 → +0.003 (≈ zero), accuracy up, robust on every seed; plus max_grad_norm gradient clipping (saṃyama).
  • Adaptive guṇa (Antahkarana(adaptive_guna=True)) — plasticity closes the loop on measured forgetting: strictly Pareto-better where there's headroom (−11% forgetting and +accuracy), a safe no-op when there isn't.
  • OneClassNovelty — complementary zero-day detector for threats sitting inside the known-data manifold (UNSW Backdoor AUROC 0.47 → 0.93).

antahkarana 1.1 performance

Full details + the honest negatives we kept (and didn't ship): CHANGELOG · API.md.

The 1.0 foundation — stable API, with the control ring + calibrated novelty built in

The public API is now frozen under SemVer — 13 symbols (Antahkarana, the adapters, FeatureOOD, NoveltyCalibrator, the eval metrics, and the Policy / RiskRouter / Tier / Intervention control ring). A breaking change to any of them — or to the frozen BackboneAdapter contract — bumps the major version, so your adapters and training code keep working while organ implementations stay free to improve. Full surface + guarantees: API.md. The two capabilities that landed on the road to 1.0 are part of that stable surface:

The control ring (Layer-2 policy)

A swappable risk-tier router over any trained backbone. One consolidated brain (the weights = Layer-1 disposition), many deployment policies (config = Layer-2) — change safety posture without retraining, and every decision is logged (transparency by construction).

two-layer architecture

A request flows through the consolidated organs (Layer 1); manas emits a risk score; the RiskRouter applies the deployment Policy (Layer 2) to pick a tier — ALLOW / NOTIFY / SOFT_BLOCK / HARD_BLOCK — and logs every decision.

control ring performance

Validated on UNSW-NB15: calibrated to ~5% false-block on normal traffic, the ring gates known attacks and catches ~95% of a held-out zero-day family it never trained on — built on the Layer-1 DER++ no-forgetting win.

from antahkarana import Policy, RiskRouter
ring = RiskRouter(Policy.load("policy.yaml"))     # hard/soft blocks, thresholds, operator overrides
iv   = ring.gate(score=novelty, category="intrusion")
# iv.tier ∈ {allow, notify, soft_block, hard_block};  iv.allowed / iv.notify / iv.reason drive the response

See examples/control_ring.py.

Calibrated novelty — NoveltyCalibrator (manas → ring)

Raw novelty scores live on arbitrary scales, so a fixed ring threshold over- or under-blocks. NoveltyCalibrator maps one or more signals (energy, FeatureOOD, …) to a calibrated [0,1] against a normal reference (empirical CDF) — so a ring threshold of 0.95 is a 5% false-alarm rate, and manas output is directly ring-ready.

from antahkarana import NoveltyCalibrator
cal = NoveltyCalibrator(mode="mean")                  # "mean" steadier · "max" more sensitive
cal.fit(energy=normal_energy, feature=normal_feat)    # calibrate against NORMAL data
p = cal.score(energy=new_energy, feature=new_feat)    # 0..1 ; 0.95 ≈ 5% false alarms

Honest note: no unsupervised fusion of energy+feature beat the best single signal on held-out UNSW-NB15 families — calibration's value is the FPR-tunable scale, not an AUROC boost.

Core (since 0.3.0) — DER++ replay + feature-space novelty

DER++ works for tabular (replay_strategy="der"): on a 6-family UNSW-NB15 stream it cut forgetting from +0.048 (naive) to +0.009 — the no-forgetting guarantee holds on modern data, not just text/vision.

FeatureOOD adds a penultimate-feature novelty detector (Mahalanobis / kNN) complementary to the default energy score — measured to rescue families where energy fails (Backdoor AUROC 0.37→0.79) while energy remains better where it already works. Pick per deployment or ensemble; energy stays the default.

from antahkarana import FeatureOOD
det   = FeatureOOD("mahalanobis").fit(known_features, known_labels)   # bb.features(inputs) -> embeddings
score = det.score(new_features)        # higher = more novel / likely zero-day

Text (any HuggingFace causal-LM + LoRA)

from antahkarana import Antahkarana, TextAdapter
bb     = TextAdapter("mistralai/Mistral-7B-v0.1")             # frozen base, small LoRA trains
stream = TextAdapter.make_stream(your_text_tasks)             # [(train_pairs, eval_pairs), …]
Antahkarana(bb).train(stream)

Any other modality

Copy antahkarana/adapters/template.py (CustomAdapter) and implement 5 methods over your encoder (audio, graph, multi-modal, robotics, …). The continual-learning mind is unchanged.

Runnable examples

python examples/train_tabular.py     # concept-drift stream: naive forgets, the core doesn't
python examples/train_security.py    # continual threat detection (attack families) + calibrated triage

What you get back

train() returns: the task×task accuracy matrix, final_row, final_avg, forgetting, risk_coverage (calibrated abstention), recovery (sleep), and the trajectory (per-task guṇa/novelty).

Knobs

Antahkarana(bb, samskara=, replay_strategy="naive"|"der", avidya_strategy="msp"|"energy", sleep=, base_lr=, epochs=) — turn each organ on/off and swap in the SOTA implementation (DER++ dark-knowledge replay, energy-OOD novelty).

Ship it closed-source

python build_wheel.py compiles the engine to binary .so (Nuitka) and drops the source, leaving only the public interface readable — others can pip install and train on their data without seeing your code. (For maximum IP protection, serve it behind an API instead.)

60-second tour

python examples/try_it.py        # no-forgetting · zero-day · abstention · control ring

(API stability + the frozen BackboneAdapter contract are covered under The 1.0 foundation above; full surface in API.md.)


Author — Deepak Soni · Apache-2.0

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

antahkarana-9.0.0.tar.gz (345.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

antahkarana-9.0.0-py3-none-any.whl (390.1 kB view details)

Uploaded Python 3

File details

Details for the file antahkarana-9.0.0.tar.gz.

File metadata

  • Download URL: antahkarana-9.0.0.tar.gz
  • Upload date:
  • Size: 345.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for antahkarana-9.0.0.tar.gz
Algorithm Hash digest
SHA256 9d4a9180a423d28e5684da008d7dfc69f2a46af2946816c303fce6dca5a8b92b
MD5 d874cc7352d86be15ce0f1e49b57dc52
BLAKE2b-256 57d153180525ae3c5c5a4421a58f120c6f730e24f3dace8fd7ab38bdb6a17b68

See more details on using hashes here.

File details

Details for the file antahkarana-9.0.0-py3-none-any.whl.

File metadata

  • Download URL: antahkarana-9.0.0-py3-none-any.whl
  • Upload date:
  • Size: 390.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for antahkarana-9.0.0-py3-none-any.whl
Algorithm Hash digest
SHA256 7783874ba65de6f112ffe216f30de23e3d7daf1e5bb89aa0c60460de8fc38324
MD5 7a31934aeb84fb5f5f6536bfbab00a33
BLAKE2b-256 cf37cf96ad61f3dca6437250a79f35da92838a93d4d5094cc259e22f60197436

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page