Skip to main content

LatticeAG ForgeDistill 🔨

License CI Stars Issues Top language Python

Correctness-by-construction distillation for agentic tool-calling models.
Distillation Studio first slice: run lineage ledger + packaged CLI, not a hosted trainer.

Why This Exists · How It Works · Quick Start · Configuration · Dataset Format · Verification · Known Issues


ForgeDistill (part of the LatticeAG Forge series - distillation, training, and models) is an open-source harness for generating high-structure training data for agentic tool-calling models - multi-turn, sequentially dependent tool calls with real error-recovery - using any teacher model, weak or strong.

Most distillation frameworks ask a teacher to demonstrate correct agentic behavior. Most models can't do that reliably, which is why so much synthetic agentic data has shallow one-call trajectories and fabricated values. ForgeDistill inverts the problem: the structure is built deterministically, the teacher only writes prose.

Built by LatticeAG.

Why This Exists

Framework Approach Weakness
distilabel (Argilla) General-purpose pipeline framework No built-in correctness guarantees - you write your own steps
APIGen / xLAM (Salesforce) Post-hoc 3-stage verification Verifies after generation, still throws away garbage, needs frontier teachers
Glaive function-calling Generator-trusted, mostly single-turn No dependency enforcement
AgentInstruct (Microsoft) Multi-agent flows No tool-dependency correctness

The differentiator: correctness by construction + grounding gates.

  1. We build the tool-call chain deterministically with real, unskippable dependencies - e.g. send_email(to="$0.result.email") where the address is an opaque value only obtainable by calling a prior tool. Guessing fails with a 400.
  2. The teacher is used only for prose (thoughts + final answer), never for tool calls - malformed JSON is impossible.
  3. Semantic grounding gates reject final answers that fabricate values not present in the real tool results.

Result: structurally-perfect, prose-grounded training traces from any OpenAI-compatible teacher endpoint.

Distillation Studio this cycle is CLI-first lineage plus packaged console scripts, not a hosted trainer and not Unsloth.

How It Works

flowchart LR
  A[Phase 1 - Chain Builder<br/>deterministic plan templates] -->|executed trajectory| B[Phase 2 - Teacher<br/>writes thoughts + final answer]
  B -->|prose only| C[Gates<br/>format + grounding]
  C -->|validated| D[Training trace<br/>reversed-v2 JSONL]

Phase 1 - deterministic chain construction (zero teacher tokens). agentic_plans.py picks a plan template (multi_hop, branch, recovery, join, fanout, reorder, digest, schema, idempotent, disambiguate, stop), fills variables, resolves $S.result references against prior step results, and executes against a deterministic mock executor. Sequential dependencies are unskippable by construction:

  • send_email(to="$0.result.email") - the teacher cannot guess the opaque address (e.g. carol.y.9024@internal.corp); it must come from a prior get_user result
  • db_query(...WHERE plan = '$0.result.plan') - the exact plan string must be learned from the user profile
  • Error-recovery plans deliberately fail the first step (bad id / bad city), then correct - teaching the observe-error-and-retry loop

Phase 2 - teacher prose (the only teacher spend). The teacher receives the full execution record (prompt + every call + exact results) and writes ONLY:

  • N <thought> blocks (reasoning before each call, including recovery thoughts)
  • FINAL_ANSWER: grounded in the real tool results

The harness assembles the final trace with its own validated tool calls. Weak models can do this; strong models do it better - both produce structurally-correct data.

Quick Start

# 1. Clone + editable install (verified path; this is what CI runs)
git clone https://github.com/LatticeAG/ForgeDistill.git
cd ForgeDistill
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

# 2. Or from PyPI (published release)
pip install latticeag-forge-distill

# 3. Configure your multi-provider OpenAI-compatible roster
cp configs/roster.example.yaml configs/roster.yaml
#    - set base_url / key_env per provider
#    - export your keys, e.g. export MY_PROVIDER_KEY=sk-...

# 4. Sanity check the modules (no sys.path hacks after install)
python -c "import distill_tools, agentic_plans, prose_writer; print('OK')"

# 5. Pilot run (10 traces)
ulimit -n 65536
distill --count 10 --pilot

# 6. Audit the output on disk (don't trust stdout)
cat data/raw/traces_*.jsonl | python -m json.tool --json-lines | head -20

After pip install -e ".[dev]", console scripts on PATH: distill, eval_card, export_sft, dpo_pairs, eval_live, archive_data, dataset_publish, forge-status. Direct python src/X.py still works.

Trace jsonl is append-guarded: the harness refuses to run over existing traces unless --wipe (which archives via archive_data, never unlinks). Derived --out paths (eval_card.json, export_sft, dataset_publish bundle files) replace on re-run.

Configuration

Copy the example roster and point it at your OpenAI-compatible endpoints:

cp configs/roster.example.yaml configs/roster.yaml
export MY_PROVIDER_KEY=sk-...   # keys live in env vars ONLY, never in the file
distill --count 10 --pilot --roster configs/roster.yaml

Roster shape (verified from configs/roster.example.yaml):

providers:
  example-provider:
    base_url: "https://api.example.com/v1"
    key_env: "EXAMPLE_PROVIDER_KEY"   # env var holding the key
    concurrency: 32
    strict_tool_protocol: false       # true only if your router enforces tool_call_id pairing
    models:
      example-model-thinking: {mode: thinking, max_tokens: 2000, weight: 2}
      example-model-concise:  {mode: concise, max_tokens: 1000, weight: 1}

  local-router:                       # keyless local proxy
    base_url: "http://127.0.0.1:PORT/v1"
    key_env: ""                       # empty = no auth
    concurrency: 64
    models:
      local-model-a: {mode: concise, max_tokens: 1000, weight: 1}
Key Description
mode thinking (reasoning models, keep max_tokens >= 2000) or concise (fast models, max_tokens ~1000).
weight Relative sampling weight — higher means more traces from that model.
concurrency Per-provider semaphore width.
strict_tool_protocol Set true only if your router enforces OpenAI tool_call_id pairing on tool messages.

distill flags: --providers / --models (comma-separated subsets), --curriculum {off,uniform,linear}, --holdout-frac (default 0.15, 0 disables), --verify-sample / --no-verify, --cross-teacher / --cross-teacher-rate, --dpo / --dpo-rate, --mp / --shard (multiprocess; --mp child i uses seed + i × 7919), --wipe, --out-dir (default data/raw).

Dataset Format

Reversed-v2 format - each JSONL line:

{
  "seed_class": "agentic",
  "prompt": "...user prompt...",
  "teacher": "provider/model",
  "teacher_mode": "thinking|concise",
  "plan_id": "plan_<template_id>",
  "distill_version": "reversed-v2",
  "forge_spec": "0.2",
  "plan_tier": "easy",
  "chain_steps": [
    {"tool": "get_user", "args": {"user_id": 123}, "result": {"status": 200, "result": {...}}, "expect": "success"}
  ],
  "eval": {
    "format_ok": true,
    "grounding_ok": true,
    "chain_ok": true,
    "dependency_ok": true,
    "n_rounds": 2,
    "n_tool_calls": 2,
    "skills": ["multi_hop", "stop"],
    "tier": "easy",
    "repaired": false,
    "cross_teacher": false
  },
  "messages": [
    {"role": "system", "content": "..."},
    {"role": "user", "content": "..."},
    {"role": "assistant", "content": "<thought>...</thought>\n<tool_call>[{\"name\": ..., \"arguments\": {...}}]</tool_call>"},
    {"role": "tool", "content": "{\"status\": 200, \"result\": {...}}"},
    {"role": "assistant", "content": "final answer..."}
  ]
}
  • Assistant turns: thought + exact tool call (deterministic, validated JSON)
  • Tool turns: real executor results (success or error payloads)
  • No artificial "provide the final answer" nudge in exported messages
  • Final answer must contain key facts from the last successful tool result (grounding gate)

Verification

CI (.github/workflows/ci.yml) runs on Python 3.11 and 3.12: compileall, pytest -q, CLI --help smoke, eval_card --input tests/fixtures --require-gates, 300-chain invalid count == 0, python -m build + twine check.

300-chain stress test (after editable install; printed number is the invalid count, want 0):

python -c "from agentic_plans import build_chain, validate_chain; import random; r=random.Random(0); print(sum(1 for _ in range(300) if validate_chain(build_chain(r)['steps'])))"

Dataset (structural) audit

Operator-measured (not in this git tree)

Latest operator run: 500 traces, multi-provider teacher fleet, reversed-v2. Source file data/raw/eval_card.json is gitignored; a clone cannot reproduce these cells without a live teacher run.

eval_card --input data/raw --out data/raw/eval_card.json --require-gates
Metric Source path Value
Trace count n_traces 500
Prose gate pass rate gates.validate_prose_trace_pass 1.0
Grounding gate pass rate gates.validate_answer_grounding_pass 1.0
Chain gate pass rate gates.validate_chain_pass 1.0
Dependency fidelity pass rate gates.dependency_fidelity_pass 1.0
Nudge leak rate gates.nudge_leak_rate 0.0
Malformed tool-call rate gates.malformed_tool_call_rate 0.0
Multi-round rate structure.multi_round_rate 0.988
Unique prompt rate structure.unique_prompt_rate 0.996
Unique trajectory hash rate structure.unique_traj_hash_rate 1.0
Send-email learned address rate structure.send_email_learned_address_rate 1.0
Cross-teacher split rate structure.cross_teacher_split_rate 0.0 (opt-in, off by default)
Cross-teacher fallback rate structure.cross_teacher_fallback_rate 0.0
300-chain invalid commands.stress_300 stdout 0 (300/300 valid)
Plan templates defined skills.n_templates_defined 47
Skill tags defined skills.n_tags_defined 15
Easy tier plans skills.tier_counts.easy 9
Medium tier plans skills.tier_counts.medium 12
Hard tier plans skills.tier_counts.hard 21
Expert tier plans skills.tier_counts.expert 5

--require-gates exits 0 on that operator run. Numbers were copied from an operator-local data/raw/eval_card.json.

CI-measured (in this git tree)

eval_card --input tests/fixtures --require-gates

tests/fixtures/mini_traces.jsonl and tests/fixtures/traces_fixture.jsonl are byte-identical. CI globs traces_*.jsonl under that directory.

Metric Source Value
Trace count n_traces 3
Prose / grounding / chain / fidelity gates.*_pass 1.0
Nudge leak / malformed tool-call gates.*_rate 0.0
observability.n_rejects_files fixture dir has no reject sidecars 0

Historical note (2026-08-13, v0.1 pilot, n=10): 10/10 passed format + grounding gates; 100% multi-round; 0 malformed tool calls; 0 unique-prompt collisions; all send_email calls used in-context learned emails.

Plan counts source of truth:

python -c "from agentic_plans import PLANS, SKILLS; print(len(PLANS), len(SKILLS), sorted(SKILLS))"

Forge Live Tool Eval (student), not BFCL v3

Category Score Command
multiple no student checkpoint in this tag python src/eval_live.py --endpoint $STUDENT_URL --model $STUDENT_MODEL --holdout data/raw/holdout_plan_ids.json --n 50 --out data/raw/live_eval.json --key-env STUDENT_KEY_ENV
parallel no student checkpoint in this tag same as above
multi_turn no student checkpoint in this tag same as above
irrelevance unscored same as above

CI uses eval_live --replay (ReplayStudent) separately; those scores are not student-checkpoint rows in this table. ReplayStudent fixture score: tests/test_eval_live.py::test_replay_student_scores_one_on_ten_built_chains.

Export quickstart

export_sft \
  --input data/raw \
  --out data/export/nanbeige.jsonl \
  --format messages \
  --check-mask

--template is optional; the default is packaged forge_assets/templates/nanbeige.json. Pass --template configs/templates/nanbeige.json (or chatml.json) to override. Clone-path JSON and packaged JSON are byte-identical.

configs/templates/nanbeige.json and configs/templates/chatml.json set assistant loss: true and all other roles loss: false. --check-mask validates message-level assistant-only spans. Token-level --check-tokenizer is pending a published NANBEIGE_TOKENIZER checkpoint and an optional transformers install (not a default dependency); do not claim tokenizer verification until that env is set. --out replaces the destination file. The command above writes data/export/nanbeige.jsonl; data/export/mask_audit.txt records the first 3 traj_hash values and per-example trainable-span counts from that run. Do not type span counts by hand.

Curriculum

--curriculum {off,uniform,linear} on distill:

  • uniform: 0.25 weight per tier (easy / medium / hard / expert)
  • linear: mix_for_progress at progress 0: easy 0.50, medium 0.30, hard 0.15, expert 0.05; at progress 1: easy 0.10, medium 0.20, hard 0.40, expert 0.30

--holdout-frac 0.15 writes holdout_plan_ids.json beside traces. Pilots use --holdout-frac 0.

Coverage note from eval_card.COVERAGE_NOTE: skills.coverage is computed against train_ids only; holdout_plan_ids lists excluded ids so readers can reproduce. Tier coverage assertions in CI run with --holdout-frac 0 fixtures.

Features

Category Feature
Structure 47 plans, 15 tags: arithmetic, branch, calendar, digest, disambiguate, fanout, idempotent, join, multi_hop, nested, recovery, reorder, schema, search, stop. Tiers: easy 9, medium 12, hard 21, expert 5
Dependencies $S.result.field refs create unskippable sequential dependencies
Recovery Deliberate failure + correction plans teach observe-error-and-retry
Grounding Semantic gate rejects fabricated values in final answers
Teacher-agnostic Multi-provider teacher fleet - any OpenAI-compatible endpoint
Resilience Per-provider health state machine (healthy / backoff / quarantined), 429 quarantine, Retry-After honoring, exponential backoff with jitter
Fleet management Per-provider semaphores, concurrency scaling, weighted model sampling, dead-route re-probing
Robustness Trajectory-hash dedup, checkpoint/resume per provider, per-worker RNG, token accounting
Safety Traces append-guarded; --wipe archives; derived --out paths replace
Curriculum --curriculum {off,uniform,linear} tier mixing
Export export_sft.py renders SFT jsonl with assistant-only loss masks
DPO dpo_pairs.py offline preference pairs from assembled traces
Eval card eval_card.py --require-gates structural audit json
Live hook eval_live.py Forge Live Tool Eval (ReplayStudent in CI)
Lineage lineage.py writes lineage_{prov}.jsonl (kept) and rejects_{prov}.jsonl (dropped); eval_card.observability
Studio CLI Eight console scripts after install (see Quick Start)
Public bundle dataset_publish scrubs traces, DPO pairs, and lineage for HuggingFace

Lineage

Each kept trace gets an additive lineage_id (also on the sidecar row). Sidecars:

  • data/raw/lineage_{provider}.jsonl - kept records (lineage_spec 1.0): plan id, teacher route, tokens for that trace, eval copy, optional dpo_pair_id. No messages, no chain_steps, no API keys.
  • data/raw/rejects_{provider}.jsonl - one row per dropped attempt (attempt_seq + ts so retries do not collide).

eval_card loads those globs when --input is a directory and adds observability (n_kept, n_rejects_files, reject_reasons, tokens_per_kept_trace, n_lineage_ids). dataset_publish writes lineage/train.jsonl with teacher fields stripped.

Repository Layout

src/agentic_plans.py     Phase 1: deterministic chain builder + plan templates
src/prose_writer.py      Phase 2: teacher prose contract + format/grounding validators
src/distill_tools.py     Async worker loop: fleet health, semaphores, checkpointing, CLI
src/mock_tools.py        Shared deterministic tool executor (11 tools)
src/archive_data.py      Archive data/raw to data/archive/<timestamp>_<label>/ - never deletes
src/eval_card.py         Structural eval card json from traces_*.jsonl (+ observability)
src/export_sft.py        SFT export with template-driven loss masks
src/eval_live.py         Forge Live Tool Eval student hook
src/verifier.py          Optional LLM verifier for prose repair
src/curriculum.py        Tier mix, holdout split, plan picking
src/dpo_pairs.py         Offline DPO pair builder
src/dataset_publish.py   Public-safe HuggingFace bundle scrubber
src/lineage.py           Run lineage ledger helpers (kept/reject/summarize)
src/status.py            forge-status: count traces_*.jsonl
src/seed_bank.py         Frozen legacy-v1 seeds (packaged, generator not called)
src/harvest_prompts.py   Unpackaged legacy HF harvest (operator-only)
src/dedup_filter.py      Unpackaged legacy prompt filter
src/forge_assets/        Installed template pack + roster.example.yaml
configs/templates/*.json Chat templates (nanbeige, chatml); byte-match forge_assets
configs/roster.example.yaml  Teacher fleet config template (copy to roster.yaml)
tests/                   pytest suite
.github/workflows/ci.yml CI: 3.11+3.12, pytest, gates, build/twine
pyproject.toml           Package metadata and console scripts (0.4.2)
safe_launch.sh           Archive-first launcher (FORGE_PYTHON / .venv / python3)
SECURITY.md              Vulnerability reporting; keys env-only
AGENT.md                 Contributor conventions
LICENSE                  MIT

harvest_prompts.py and dedup_filter.py are unpackaged legacy (excluded from py-modules). forge_assets is the wheel template pack.

Publishing

Scrub internal fields before a public upload:

dataset_publish --input data/raw --out /tmp/hf-bundle --check

That writes sft/train.jsonl, dpo/train.jsonl, and lineage/train.jsonl without teacher identities. Export HF_TOKEN for the commands below; do not rely on a cached huggingface-cli login; if unset at tag time the upload is a later operator step.

huggingface-cli upload LatticeAG/ForgeDistill-agentic data/raw/eval_card.json --repo-type dataset
huggingface-cli upload LatticeAG/ForgeDistill-agentic data/export/nanbeige.jsonl --repo-type dataset --path-in-repo sft/nanbeige.jsonl
huggingface-cli upload LatticeAG/ForgeDistill-agentic data/raw/dpo_pairs_offline.jsonl --repo-type dataset --path-in-repo dpo/dpo_pairs.jsonl

Include holdout_plan_ids.json on the dataset card so the 0.15 holdout split is visible to consumers. Operators may upload it alongside eval_card.json:

huggingface-cli upload LatticeAG/ForgeDistill-agentic data/raw/holdout_plan_ids.json --repo-type dataset

Configure teachers via a multi-provider OpenAI-compatible roster (configs/roster.yaml); keys live in env vars only.

Locked design decisions (v0.3, still in 0.4.0)

search.query. Intersection over whitespace tokens; hit order is DOC_BY_QUERY[tokens[0]]; empty and unknown tokens behave as implemented. Do not change this semantics.

price. Price values come from VAR_POOLS["price"] = [20, 35, 50], not a CRM field.

Opaque ids. Identifiers such as evt.* and doc.* may be quoted when they appeared in tool payloads; inventing one fails grounding; students are not required to recite them.

roles.answer pin. A roster pin for roles.answer wins over sampling; rate still gates whether a split happens; sampling never overrides a valid pin.

--mp token accounting. Multiprocess runs write .token_usage_{shard_i}.json sidecars merged by the parent; do not publish all-zero token totals on multiprocess runs.

License

Released under the MIT License. Built by LatticeAG.

Known Issues

  • Live teacher endpoints required for generation - the harness ships no teacher; without reachable base_urls and exported key env vars, distill cannot produce traces. Fixture-only verification (eval_card --input tests/fixtures --require-gates) works offline.

  • Token-level mask check needs a checkpoint - --check-mask validates message-level spans, but --check-tokenizer stays pending until a published NANBEIGE_TOKENIZER checkpoint plus an optional transformers install.

  • No student checkpoint in this tree - eval_live category rows are unscored without a $STUDENT_URL endpoint; CI covers the ReplayStudent path only.

Metadata

Release files for latticeag-forge-distill 0.4.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for latticeag-forge-distill 0.4.2
File Size Uploaded
latticeag_forge_distill-0.4.2.tar.gz 120.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for latticeag-forge-distill 0.4.2
File Interpreter ABI Platform
latticeag_forge_distill-0.4.2-py3-none-any.whl Python 3 none any Details

Total release size: 213.0 kB

Release files / latticeag_forge_distill-0.4.2.tar.gz

Download URL latticeag_forge_distill-0.4.2.tar.gz
Size 120.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f12d0f7f3ef109e51306ee72cea49164805d90077feb8cb00da6678e2d0a49ac
BLAKE2b-256 checksum
How to use checksums
7b2e966bd27fb41aa65d3f33e722c6f564f3bee7dd646d71adedf7dd16f982ca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / latticeag_forge_distill-0.4.2-py3-none-any.whl

Download URL latticeag_forge_distill-0.4.2-py3-none-any.whl
Size 92.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4226eae03c923f3a357fb38c043542a01e8d118ec2596f9c034ad85899f9ac6a
BLAKE2b-256 checksum
How to use checksums
214c40be2bcf2cb7a12b195f7ff19d9b14db3172abaf2b271454205d531176f4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

0.5.1

2 release files

0.5.0

2 release files

This release

0.4.2 This release

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page