Skip to main content

A declarative language for LLM-driven state machines (reference interpreter).

Project description

mklang

CI Docs License

A declarative language for LLM-driven state machines, with an agent-first console to author and run them. A .mkl file (mklang) describes an agent as a set of states; an LLM is the runtime that executes generative steps. The document is the program; the mklang console is how you drive it. The host supplies the interpreter, optional tools, and optional code-hook gates.

mklang : LangGraph  ::  a declarative spec : Python code

Two things to look at first: the language — states with prose faces and natural-language gates as transitions — and the console — an agent-first TUI that authors, commissions, and traces machines for you. Everything else (CLI, MCP, scenario tests) is scaffolding around those two.

See it in action

The two product surfaces — the console and the language:

Console: agent-first TUI, commissioning, trace Language: gates, tools, reasoning loop
Live mklang agent demo mklang language demo

Both recordings run the real surfaces against DeepSeek (the agent demo also hits the live web). See the full WebM recordings, transcripts, and reproducibility notes.

The language

Each state has four faces:

Face Answers Example Ref. interpreter
structure what shape? "The output is an email reply, max 150 words" system
execution how to act? "Do not invent policies not in the KB facts" system
prompt what to think? "Write a reply to {{ticket.body}}…" user (+ {{…}})
gates when to exit? see below separate judge call

Sticky policy goes in execution; turn data and {{context}} go in prompt. There is no system: keyword — see Best practices §3.

Real side effects (search, send, calc) are tool: states — host callables, not prose in execution. See examples/react.mkl and examples/triage.mkl.

The output of a state is stored in the shared context under its output: key, so later states read it via {{key}}. Four optional faces unlock richer reasoning: reason (traced chain-of-thought), accumulate (append to a list), fan-out (sample: N / over: {{list}}), and call (run another machine) — see Reasoning architectures.

Gates are the transitions. A state's gates list is its transition table: each gate is a natural-language condition the LLM judges, plus what happens next.

gates:
  - when: the reply resolves the request and is in the required tone
    then: ok
    to: send
  - when: information from the KB is missing
    repair: 2 # re-run this state with feedback, up to 2 times
    to: gather
  - when: the request needs a human
    escalate: true
    to: human_review

Policies: ok (advance), repair(N) (self-correct with feedback), escalate (route to a handler), fail (abort). A global step budget prevents runaway loops.

The console

mklang console is the agent-first front door and the primary way to use the language: type what you want, and the console's agent authors or picks a machine, commissions it, and streams the run state-by-state as a live trace tree — escalations and tool consent come back to you inline (human-in-the-loop is a first-class part of the flow, not an afterthought).

pip install mklang        # the console ships in the core package since 0.15.0
mklang console

The agent itself is a machine (agent.mkl) — read it, lint it, swap it with --agent your_brain.mkl (ADR 0015). It has no privileged powers the language lacks: it commissions the same .mkl machines you write, over the same gates and tiers. Sessions persist (--continue), agent replies render as Markdown, and the brain declares host clocks today / now for wall-clock questions. Details: docs/guides/console.md.

Design commitments

  • Document-first — readable without the interpreter; prose-first for the common path. Production machines still need developer judgment for tools, hooks, and untrusted inputs (see SPEC threat model); the runtime delimits untrusted context structurally (SPEC §6), it does not judge it for you.
  • LLM-as-runtime — non-deterministic by design; gates (prose + optional code hooks + budgets + trace) are the reliability mechanism. Prose-gate accuracy is an empirical claim, not a free lunch.
  • Prose, not typesstructure and gate conditions are natural language, judged by the LLM at runtime; optional hook: gates add host bool checks.
  • Provider-agnostic — a .mkl never names a provider or model. States route by capability tier (fast / balanced / reasoning); the runtime maps each tier to a concrete model. Portability of the document is syntactic; whether different providers fire the same gates on the same run is measurable (see scripts/gate_divergence.py).
  • Spec + conformance — an implementation-neutral conformance suite pins interpreter semantics so a second runtime can match the language contract.
  • Language-agnostic runtime — the spec assumes only "some host with an LLM".

Reasoning architectures

Every modern reasoning/agentic pattern maps onto the core (states + gates + prose + tiers + the optional faces). Full skeletons in SPEC.md §10; operating guidance in docs/guides/patterns.md.

Ten of these ship as ready, general-purpose std_* machines — parameterized by context, callable from your machines (call: std_refine), runnable by name (from the CLI or the console's /run):

mklang run std_self_consistency --set task="Estimate the risk of X"

See the stdlib catalog (ADR 0012). The patterns that need host tools/hooks or static call: targets (ReAct, router, exact policy) stay as authored examples.

Architecture mklang constructs
Chain-of-Thought reason: true
ReAct think → tool state (host callable) → observation accumulated
Reflexion / self-refine produce → self-judge gate → repair
Self-consistency sample: N → reducer state (majority)
Tree-of-Thought sample: k → score/select reducer → loop (depth via budget)
Plan-and-Execute planner parse: list (0.3) → over: {{steps}} → reducer
Debate / ensemble over: {{personas}} → synthesizer
Map-Reduce over: {{chunks}} → reducer
Router-of-experts classify → call specialists
Speculative cascade tier: fast draft → escalatetier: reasoning
Exact policy checks gate hook: host (ctx, output) -> bool (no LLM)

Runtime configuration

The .mkl picks a tier; a host-side config picks the model. This is the whole of "make it multi-provider":

active: deepseek # deepseek | anthropic | openai | google | openrouter | xai | mistral | local
providers:
  deepseek:
    base_url: https://api.deepseek.com
    tiers:
      {
        fast: deepseek-chat,
        balanced: deepseek-chat,
        reasoning: deepseek-reasoner,
      }
  anthropic:
    tiers:
      {
        fast: claude-haiku-4-5,
        balanced: claude-sonnet-5,
        reasoning: claude-opus-4-8,
      }
  local:
    base_url: http://localhost:11434/v1
    tiers: { fast: qwen3:8b, balanced: qwen3:32b, reasoning: deepseek-r1:70b }

The example config defaults to DeepSeek (the path we live-test against). Flip active: anthropic (or openai / local / …) and every example runs unchanged. Blocks ship for Anthropic, OpenAI, Google, DeepSeek, OpenRouter, xAI (Grok), Mistral, and local (Ollama/vLLM) — every non-Anthropic one is OpenAI-compatible, so a single adapter serves them all. OpenRouter is a meta-provider: its vendor/model ids let each tier target a different vendor through one endpoint. Per-tier params (Anthropic adaptive-thinking + effort, OpenAI/xAI reasoning_effort, …) live under params. Full map: config/runtime.example.yaml.

Install

pipx install 'mklang[mcp]'   # console TUI is in the core package; [mcp] adds the MCP server
mklang init --user           # scaffold config, .env, and a sample machine
# set DEEPSEEK_API_KEY (or another provider key) in the .env that init reported, then:
mklang console

pip install mklang works too. The full walk-through — including a first run that needs no API key — is the Getting started guide. A one-shot scripts/install.sh and an Arch Linux PKGBUILD are also available.

Editor validation for .mkl files works out of the box via the JSON Schema — point yaml-language-server at https://raw.githubusercontent.com/gianlucamazza/mklang/main/schema/mklang.schema.json.

The CLI (for scripting and CI)

The console is the interactive surface; the mklang CLI is the scriptable one — same interpreter, same machines. Drive a checkout through uv with no install step:

git clone https://github.com/gianlucamazza/mklang && cd mklang
cp .env.example .env            # set DEEPSEEK_API_KEY=… (or another provider key)
uv run mklang check examples/self_consistency.mkl
uv run mklang lint examples/self_consistency.mkl   # + static analysis
uv run mklang run examples/self_consistency.mkl \
  --set question.text="What is the capital of Australia?"
# default provider is deepseek; override with --provider anthropic|openai|…

# pause on budget, resume later (exit code 3 = suspended):
uv run mklang run examples/self_consistency.mkl --max-tokens 300 --checkpoint ck.json
uv run mklang resume ck.json --max-tokens 5000

# human-in-the-loop: escalate gates suspend; resume with the human decision:
uv run mklang run examples/expense_approval.mkl --checkpoint ck.json --hitl
uv run mklang resume ck.json --set human.reply="approved, cost center 42"

Every command, flag, and exit code: CLI reference. The .mkl picks tiers; config/runtime.example.yaml maps them to models (active: deepseek by default); the key comes from .env. Same machine, any provider.

Test your machine without API keys

mklang test runs your machine against a script of named scenarios with a scripted LLM (produce texts, judge picks) and scripted tools/hooks — fully deterministic, no provider or key. It pins the paths you care about before you spend a token on a live run.

uv run mklang test examples/triage.mkl --script examples/triage.test.yaml
# PASS happy-path
# PASS kb-empty-escalates

Each scenario declares a scripted llm:/tools:/hooks:, optional host input: context (untrusted by provenance, SPEC §6), and an expect: (status, error, result, at, trace skeleton, context keys) — the same case format the conformance suite uses. A mismatch prints a minimal diff (the first differing key, expected vs actual) and exits 1. See examples/triage.test.yaml.

MCP server (agentic hosts)

Agent hosts that speak MCP (Claude Code and other clients) can commission a machine instead of embedding the library (ADR 0011): the host requests a run and gets back the result with full provenance (trace + usage).

pip install 'mklang[mcp]'
claude mcp add mklang -- mklang-mcp

The server auto-discovers config and keys through the same chain as the CLI (project → user host → /etc/mklang → bundled example, ADR 0023); pass --config only to pin a specific file. It exposes commissioning tools (run / resume), discovery (list_machines / describe_machine), and check (ADR 0011 + 0013). run accepts inline .mkl source or a path, with inputs merged into the context (host inputs are untrusted by provenance and reach prompts fenced — SPEC §6); resume takes an opaque handle or checkpoint file (e.g. {"human.reply": "…"} for HITL, injected values equally tainted). Live engine events stream as mklang.event logging notifications (ADR 0019). In-memory sessions hold suspensions unless you pass checkpoint_path. Provider keys resolve server-side from the environment, never over the wire.

Files

Stack

  • Language spec: .mkl = YAML validated by a JSON Schema; semantics fixed by SPEC.md and an implementation-neutral conformance suite.
  • Reference interpreter: Python ≥ 3.11, dependencies pyyaml, jsonschema, python-dotenv, openai (the OpenAI-compatible adapter serves every non-Anthropic provider), rich, and textual (the console).
  • Providers: DeepSeek / OpenAI / Google / OpenRouter / xAI / Mistral / local via one OpenAI-compatible adapter, plus a native Anthropic adapter (extra).
  • Surfaces: the textual console TUI, the mklang CLI, and an optional stdio mklang-mcp server (extra mklang[mcp]).
  • Quality: ruff, mypy (zero suppressions, a growing strict tier), pytest + pytest-cov (coverage gate, offline via MockLLM/scripted LLM), and the conformance suite — on an ubuntu 3.11–3.13 + macOS + Windows matrix.
  • Packaging: hatchling; published to PyPI via GitHub OIDC Trusted Publishing; Arch PKGBUILD.

Status

Language v0.3 / package 1.0.3 — core complete: states + gates + prose, tiers, reason / accumulate / fan-out / call / tool / parse: list / code-hook gates; multi-provider interpreter with entry-point plugins (tools, hooks, providers, machines); resumable checkpoints + HITL; mklang check / lint (--llm optional) / test / doctor; conformance suite; machine stdlib (std_*); MCP host; console TUI (bundled by default); structured web search (offline stub by default); host tool stub architecture for search / search_kb / send_reply (ADR 0020); host clock conventions context.today / context.now; sectioned produce system prompts from structure+execution; output anti-cutoff + context budgets (ADR 0016–0019); untrusted-context delimiting — provenance taint + <data-NONCE> fences in produce and judge prompts (SPEC §6, ADR 0025); best practices. Gate judging follows the state tier by default.

  • Live: DeepSeek (default) and OpenAI green through the 1.0.x release matrices, including the blocking cross-provider gate-agreement check at 1.0 — the release gate runs the single gate_divergence machine; the four-machine suite measured 1.0 agreement per machine (DeepSeek + OpenAI ×3). Anthropic unit-tested; live e2e still blocked (key/credits).
  • Release policy: DeepSeek + OpenAI smoke and three-run gate agreement are blocking; other configured providers are reported without blocking. PyPI publication uses GitHub OIDC Trusted Publishing from the release workflow. Arch: AUR mklang tracks the PyPI sdist pin.
  • 1.0 posture: SemVer from 1.0.0, spec 0.3 frozen (ADR 0026); freeze is provisional on evidence (ADR 0028). Authoring-loop blind_spot = 0.0167 (no test_machine yet).
  • Open / later: Anthropic live when the account has credit; five-reader distribution test; on_truncate=continue stitching; language-level context zones (ROADMAP).
  • Roadmap and full release notes: ROADMAP.md, CHANGELOG.md.

License

Apache-2.0. Contributions welcome — see CONTRIBUTING.md.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mklang-1.0.3.tar.gz (456.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mklang-1.0.3-py3-none-any.whl (156.8 kB view details)

Uploaded Python 3

File details

Details for the file mklang-1.0.3.tar.gz.

File metadata

  • Download URL: mklang-1.0.3.tar.gz
  • Upload date:
  • Size: 456.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for mklang-1.0.3.tar.gz
Algorithm Hash digest
SHA256 90f6307cf3cbf5d893d454ef6a8771e411999ec226b846273dd34f323e5fbb4a
MD5 a76a59c379a67366d1e9af289ed309cb
BLAKE2b-256 65e56e816c4cb37cfbcf3d0368a9a9b731fbba2f791f5528c69512ac258567f0

See more details on using hashes here.

Provenance

The following attestation bundles were made for mklang-1.0.3.tar.gz:

Publisher: release.yml on gianlucamazza/mklang

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mklang-1.0.3-py3-none-any.whl.

File metadata

  • Download URL: mklang-1.0.3-py3-none-any.whl
  • Upload date:
  • Size: 156.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for mklang-1.0.3-py3-none-any.whl
Algorithm Hash digest
SHA256 aa1098b006a540558b787a16db59dc1d430754820629689e2eacf272102a3f17
MD5 ab5c053dfe105e33f4febc445f15f8fe
BLAKE2b-256 4948ac39c6b00c25bb3bf82153699a14162efe25ed59c08e16a18bed47895f34

See more details on using hashes here.

Provenance

The following attestation bundles were made for mklang-1.0.3-py3-none-any.whl:

Publisher: release.yml on gianlucamazza/mklang

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page