Skip to main content

A declarative language for LLM-driven state machines (reference interpreter).

Project description

mklang

CI Docs License

A declarative language for LLM-driven state machines. A .mk file (mk = machine) describes an agent as a set of states; an LLM is the runtime that executes generative steps. The document is the program; the host supplies the interpreter, optional tools, and optional code-hook gates.

mklang : LangGraph  ::  a declarative spec : Python code

The idea

Each state has four faces:

Face Answers Example
structure what shape? "The output is an email reply, max 150 words"
prompt what to think? "Write a reply to {{ticket.body}}…"
execution how to act? "Do not invent policies not in the KB facts"
gates when to exit? see below

Real side effects (search, send, calc) are tool: states — host callables, not prose in execution. See examples/react.mk and examples/triage.mk.

The output of a state is stored in the shared context under its output: key, so later states read it via {{key}}. Four optional faces unlock richer reasoning: reason (traced chain-of-thought), accumulate (append to a list), fan-out (sample: N / over: {{list}}), and call (run another machine) — see Reasoning architectures.

Gates are the transitions. A state's gates list is its transition table: each gate is a natural-language condition the LLM judges, plus what happens next.

gates:
  - when: the reply resolves the request and is in the required tone
    then: ok
    to: send
  - when: information from the KB is missing
    repair: 2 # re-run this state with feedback, up to 2 times
    to: gather
  - when: the request needs a human
    escalate: true
    to: human_review

Policies: ok (advance), repair(N) (self-correct with feedback), escalate (route to a handler), fail (abort). A global step budget prevents runaway loops.

Design commitments

  • Document-first — readable without the interpreter; prose-first for the common path. Production machines still need developer judgment for tools, hooks, and untrusted inputs (see SPEC threat model).
  • LLM-as-runtime — non-deterministic by design; gates (prose + optional code hooks + budgets + trace) are the reliability mechanism. Prose-gate accuracy is an empirical claim, not a free lunch.
  • Prose, not typesstructure and gate conditions are natural language, judged by the LLM at runtime; optional hook: gates add host bool checks.
  • Provider-agnostic — a .mk never names a provider or model. States route by capability tier (fast / balanced / reasoning); the runtime maps each tier to a concrete model. Portability of the document is syntactic; whether different providers fire the same gates on the same run is measurable (see scripts/gate_divergence.py).
  • Spec + conformance — an implementation-neutral conformance suite pins interpreter semantics so a second runtime can match the language contract.
  • Language-agnostic runtime — the spec assumes only "some host with an LLM".

Files

Runtime configuration

The .mk picks a tier; a host-side config picks the model. This is the whole of "make it multi-provider":

active: deepseek # deepseek | anthropic | openai | google | openrouter | xai | mistral | local
providers:
  deepseek:
    base_url: https://api.deepseek.com
    tiers:
      {
        fast: deepseek-chat,
        balanced: deepseek-chat,
        reasoning: deepseek-reasoner,
      }
  anthropic:
    tiers:
      {
        fast: claude-haiku-4-5,
        balanced: claude-sonnet-5,
        reasoning: claude-opus-4-8,
      }
  local:
    base_url: http://localhost:11434/v1
    tiers: { fast: qwen3:8b, balanced: qwen3:32b, reasoning: deepseek-r1:70b }

The example config defaults to DeepSeek (the path we live-test against). Flip active: anthropic (or openai / local / …) and every example runs unchanged. Blocks ship for Anthropic, OpenAI, Google, DeepSeek, OpenRouter, xAI (Grok), Mistral, and local (Ollama/vLLM) — every non-Anthropic one is OpenAI-compatible, so a single adapter serves them all. OpenRouter is a meta-provider: its vendor/model ids let each tier target a different vendor through one endpoint. Per-tier params (Anthropic adaptive-thinking + effort, OpenAI/xAI reasoning_effort, …) live under params. Full map: config/runtime.example.yaml.

Reasoning architectures

Every modern reasoning/agentic pattern maps onto the core (states + gates + prose + tiers + the optional faces). Full skeletons in SPEC.md §10; operating guidance in docs/patterns.md.

Eight of these ship as ready, general-purpose std_* machines — parameterized by context, callable from your machines (call: std_refine), runnable by name:

mklang run std_self_consistency --set task="Estimate the risk of X"

See the stdlib catalog (ADR 0012). The patterns that need host tools/hooks or static call: targets (ReAct, router, exact policy) stay as authored examples.

Architecture mklang constructs
Chain-of-Thought reason: true
ReAct think → tool state (host callable) → observation accumulated
Reflexion / self-refine produce → self-judge gate → repair
Self-consistency sample: N → reducer state (majority)
Tree-of-Thought sample: k → score/select reducer → loop (depth via budget)
Plan-and-Execute planner parse: list (0.3) → over: {{steps}} → reducer
Debate / ensemble over: {{personas}} → synthesizer
Map-Reduce over: {{chunks}} → reducer
Router-of-experts classify → call specialists
Speculative cascade tier: fast draft → escalatetier: reasoning
Exact policy checks gate hook: host (ctx, output) -> bool (no LLM)

Install

pip install mklang

Editor validation for .mk files works out of the box via the JSON Schema — point yaml-language-server at https://raw.githubusercontent.com/gianlucamazza/mklang/main/schema/mklang.schema.json.

Quickstart (reference interpreter)

cp .env.example .env            # set DEEPSEEK_API_KEY=… (or another provider key)
uv run mklang check examples/self_consistency.mk
uv run mklang lint examples/self_consistency.mk   # + static analysis
uv run mklang run examples/self_consistency.mk \
  --set question.text="What is the capital of Australia?"
# default provider is deepseek; override with --provider anthropic|openai|…

# pause on budget, resume later (exit code 3 = suspended):
uv run mklang run examples/self_consistency.mk --max-tokens 300 --checkpoint ck.json
uv run mklang resume ck.json --max-tokens 5000

# human-in-the-loop: escalate gates suspend; resume with the human decision:
uv run mklang run examples/expense_approval.mk --checkpoint ck.json --hitl
uv run mklang resume ck.json --set human.reply="approved, cost center 42"

The .mk picks tiers; config/runtime.example.yaml maps them to models (active: deepseek by default); the key comes from .env. Same machine, any provider.

Test your machine without API keys

mklang test runs your machine against a script of named scenarios with a scripted LLM (produce texts, judge picks) and scripted tools/hooks — fully deterministic, no provider or key. It pins the paths you care about before you spend a token on a live run.

uv run mklang test examples/triage.mk --script examples/triage.test.yaml
# PASS happy-path
# PASS kb-empty-escalates

Each scenario declares a scripted llm:/tools:/hooks: and an expect: (status, error, result, at, trace skeleton, context keys) — the same case format the conformance suite uses. A mismatch prints a minimal diff (the first differing key, expected vs actual) and exits 1. See examples/triage.test.yaml.

MCP server (agentic hosts)

Agent hosts that speak MCP (Claude Code and other clients) can commission a machine instead of embedding the library (ADR 0011): the host requests a run and gets back the result with full provenance (trace + usage).

pip install 'mklang[mcp]'
claude mcp add mklang -- mklang-mcp --config /abs/path/to/runtime.yaml

The server exposes exactly two tools: run (machine as inline .mk source or a path; inputs merged into the context) and resume (opaque single-use handle + e.g. {"human.reply": "…"} for HITL). Suspended runs stay in an in-memory session store — the blackboard never touches the server's disk. Provider keys resolve server-side from the environment, never over the wire.

Console (interactive)

mklang console (extra mklang[console]) is the agent-first front door: type what you want, the console's agent authors or picks a machine, commissions it, and streams the run state-by-state — escalations and tool consent come back to you inline. The agent itself is a machine (agent.mk) — read it, lint it, swap it with --agent your_brain.mk (ADR 0015).

pip install 'mklang[console]'
mklang console

Status

Language v0.2 / package 0.5.4 — core complete: states + gates + prose, tiers, reason / accumulate / fan-out / call / tool states / code-hook gates; multi-provider interpreter with entry-point plugins (tools, hooks, providers); resumable checkpoints (mklang resume, ADR 0007); human-in-the-loop (--hitl, ADR 0008); mklang check / lint / test (scripted scenarios, no API keys); implementation-neutral conformance suite (ADR 0009). Gate judging follows the state tier by default; see CHANGELOG 0.5.2 for the observable change and 0.5.3 for authoring tooling.

  • Live: DeepSeek (default) and OpenAI green (2026-07-16), including examples on the current OpenAI tier map (fast: gpt-5.4-mini, balanced/reasoning: gpt-5.5 — latest chat-completions models on this account; gpt-5.5-pro is Responses-API only and not mapped). Anthropic unit-tested; live blocked by provider billing. Gate-divergence (deepseek×openai): agreement 1.0 — see docs/experiments/gate-divergence.md.
  • Release policy: DeepSeek + OpenAI smoke and three-run gate agreement are blocking; other configured providers are reported without blocking. PyPI publication uses GitHub OIDC Trusted Publishing from the release workflow.
  • Open: Anthropic live once the account has credit; ADR 0010 (LLM lint) when ready.
  • Roadmap and full release notes: ROADMAP.md, CHANGELOG.md.

License

Apache-2.0. Contributions welcome — see CONTRIBUTING.md.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mklang-0.7.0.tar.gz (264.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mklang-0.7.0-py3-none-any.whl (104.2 kB view details)

Uploaded Python 3

File details

Details for the file mklang-0.7.0.tar.gz.

File metadata

  • Download URL: mklang-0.7.0.tar.gz
  • Upload date:
  • Size: 264.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mklang-0.7.0.tar.gz
Algorithm Hash digest
SHA256 dbbbae76f7f57647be3b187e2aeb0537ccbcb45ab41678323f23e0da15e870a9
MD5 7d5bfef155aa1b6a1c9855cbf3005a04
BLAKE2b-256 fb8e43fa603284bb00da6f2820376dce6078848f578ef45ed70b6e708d1988dd

See more details on using hashes here.

Provenance

The following attestation bundles were made for mklang-0.7.0.tar.gz:

Publisher: release.yml on gianlucamazza/mklang

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mklang-0.7.0-py3-none-any.whl.

File metadata

  • Download URL: mklang-0.7.0-py3-none-any.whl
  • Upload date:
  • Size: 104.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mklang-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 a68fc506cc1a03e38be0cb6c437c4c57c53bcbfcad52f36cedcd864aaeffad55
MD5 561029cc0146e1cd1bcda29a65a4b667
BLAKE2b-256 d1fa21bfcd76de7d40f6144fb7ce43c7a33992c240213b3de0d2c0a3139aea87

See more details on using hashes here.

Provenance

The following attestation bundles were made for mklang-0.7.0-py3-none-any.whl:

Publisher: release.yml on gianlucamazza/mklang

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page