Skip to main content

toolsynth

Six tool-use / function-calling data-synthesis papers, one reproducible framework

CI Coverage PyPI Python Methods License: MIT

English | 简体中文

toolsynth terminal demo
pip install toolsynth
export OPENAI_API_KEY=sk-... OPENAI_BASE_URL=https://api.deepseek.com TOOLSYNTH_MODEL=deepseek-chat
python -m toolsynth run self-instruct --target 176

Independent, unofficial research reproductions: the Kimi K2 agentic data pipeline, the four methods it builds on (Self-Instruct, AgentInstruct, ACEBench, ToolACE), and Kimi K3 §4.2.2 KG-guided task synthesis — each paper a self-contained package over shared LLM / schema / generation / verification / evaluation components, behavior pinned by golden tests.

Methods · Installation · Quickstart · Architecture · Data formats


Highlights

  • 6 papers, one framework — every method lives in toolsynth/methods/<paper>/ with its own config, pipeline, prompts and paper docs; the algorithmic machinery is shared, not duplicated.
  • Extensible by design — adding a paper means one MethodConfig subclass + one BasePipeline subclass (plus its stages) and a single @register_method; the rest of the framework needs zero changes.
  • One LLM client for everything — a single OpenAI-compatible client with retry/backoff, n-sampling fallback, JSON self-repair and reasoning-model handling; a scripted FakeLLMClient drives all offline tests.
  • Unified run tracing & logging — every run records per-stage timing and Alg-X provenance tags to debug/stages.jsonl; console verbosity is one --log-level knob.
  • Fidelity first — prompts and hyperparameters carry source annotations ([paper], [INFERRED], [UNRESOLVED]); paper-silent scalars stay required with no invented defaults.

Methods

Method Paper Pipeline form
self-instruct Self-Instruct (arXiv 2212.10560) non-linear bootstrap loop, 4 Alg modules
agent-instruct AgentInstruct (arXiv 2407.03502) SkillRegistry (17 skills, 3 wired) + three flows
acebench ACEBench (arXiv 2501.12851) single pipeline, five modes (construct / special / agent / grade / overall)
toolace ToolACE (arXiv 2409.00920) TSS → SDG (G4 loop) → DLV cascade
k2-agentic Kimi K2 (arXiv 2507.20534) 4-stage fusion pipeline, hybrid environment routing
k3-kg Kimi K3 §4.2.2 (KG-guided task synthesis) Phase A concept-DAG build (background thread) + Phase B–F sample → retrieve → synthesize → gate loop

Each method ships its paper and implementation notes under toolsynth/methods/<method>/docs/. The exact on-disk data formats of all artifacts (with a training-oriented comparison across methods) are documented in docs/data-formats.md.

Installation

From PyPI (recommended):

pip install toolsynth

Or from source:

python3.12 -m venv .venv && source .venv/bin/activate
pip install -e .        # deps: openai>=1.50, pydantic>=2.5, httpx>=0.27, numpy>=1.26

API keys are never hardcoded — read from the environment:

export DEEPSEEK_API_KEY=sk-...        # or OPENAI_API_KEY (+ OPENAI_BASE_URL, TOOLSYNTH_MODEL)

The generation endpoint defaults to DeepSeek deepseek-v4-flash (a reasoning model; reasoning_content is handled by the unified client). Local NLL scoring (ToolACE Eq. 1 / k2 student model) runs a Qwen checkpoint forward pass from ToolACE/models/ and does not use the API.

Quickstart

CLI

python -m toolsynth list                     # registered methods
python -m toolsynth config <method> --json   # resolved config (inspect precedence)
python -m toolsynth run <method> [flags]     # end-to-end; per-method flags: --help
python -m toolsynth run <method> --log-level DEBUG

Python API

from toolsynth import AutoPipeline, list_methods
from toolsynth.cli import configure_logging

print(list_methods())      # ['acebench', 'agent-instruct', 'k2-agentic', ...]
configure_logging("INFO")  # programmatic callers configure logging themselves

# self-instruct works with all defaults; overrides are top-priority explicit params
pipe = AutoPipeline.create("self-instruct", overrides={"target": 181})
result = pipe.run()        # → RunResult
result.run_dir             # runs/<method>/<id>/ (incl. debug/stages.jsonl)
result.artifacts           # on-disk artifact paths (e.g. machine_tasks.jsonl)
result.payload             # method-defined structured result

Per-method entry points

Method Command
self-instruct python -m toolsynth run self-instruct --target 176 --seed 0 --out runs/self-instruct/machine_tasks.jsonl (target counts the 175 seeds toward the pool total; default 176 is the historical small-scale run)
agent-instruct python -m toolsynth run agent-instruct --skill "Tool Use" --kind code --strategy hypothesize ... (--list-skills lists all 17 skills)
acebench python -m toolsynth run acebench --mode {construct|special|agent|grade|overall} ...
toolace python -m toolsynth run toolace --raw-docs <json> --target-api-count N --n-dialogs K --out ..., or --api-pool <json> to inject a prepared API pool and skip TSS (TOOLACE_* env overrides supported)
k2-agentic python -m toolsynth run k2-agentic --seed <seed.json> --targets N --agents 1 --max-turns T ... (smoke switches: --smoke-targets-real, --embed-backend)
k3-kg python -m toolsynth run k3-kg --n-tasks N --sequential --store ./kg_store --out runs/k3-kg/tasks.jsonl (needs EMBED_MODEL + SEARCH_BACKEND in config.local.py; chat endpoint rides the default LLMSettings block)

Per-method runtime requirements (for a real run)

Method Configure Notes
self-instruct / agent-instruct / acebench one OpenAI-compatible chat endpoint (DEFAULT_API_KEY / DEFAULT_BASE_URL / DEFAULT_MODEL in config.local.py, or OPENAI_API_KEY + OPENAI_BASE_URL + TOOLSYNTH_MODEL env) all LLM calls are real
toolace chat endpoint + the local Qwen checkpoint (ToolACE/models/) Eq. 1 complexity scoring is a local forward pass on your machine — no API call
k2-agentic chat endpoint + the required [UNRESOLVED] scalars + QWEN_STUDENT_PATH (local checkpoint) watch the smoke defaults — see Testing & scope
k3-kg chat endpoint + EMBED_MODEL (+ EMBED_BASE_URL / EMBED_API_KEY when the chat provider has no embedding API, e.g. DeepSeek + local Ollama) + SEARCH_BACKEND (JSON web-search backend, e.g. SerpApi) Jina-Reader page fetching is a public endpoint — keyless, real

CLI runs always build real clients: the Fake* doubles live only in tests and are injectable solely via constructor arguments — the CLI path never touches them.

Configuration

Precedence: CLI flags / overrides > environment variables > root config.local.py > method defaults.

  • Environment: OPENAI_API_KEY / OPENAI_BASE_URL / TOOLSYNTH_MODEL, plus per-field TOOLSYNTH__<FIELD>.
  • config.local.py (gitignored): an UPPERCASE = ... key binds the same-named config field (e.g. QWEN_STUDENT_PATH).
  • python -m toolsynth config <method> --json prints the fully resolved config.
  • Note: k2-agentic has several paper-silent hyperparameters that are required with no defaults — supply them via CLI / overrides / config.local.py.
  • Note: k3-kg retrieves from the real web — set SEARCH_BACKEND (a JSON web-search backend descriptor) and EMBED_MODEL (optionally EMBED_BASE_URL / EMBED_API_KEY for a separate embedding endpoint) in config.local.py.

Architecture

flowchart TD
    CLI["python -m toolsynth — CLI"]
    CORE["core — AutoPipeline registry · BasePipeline/BaseStage · MethodConfig · RunContext"]
    METHODS["methods/ — one package per paper: self_instruct · agent_instruct · acebench · toolace · k2_agentic · k3_kg"]
    SHARED["llm · schemas · generation · verification · evaluation · environment"]

    CLI --> CORE
    CORE -->|"create(name, overrides)"| METHODS
    METHODS -->|reuse| SHARED

The kernel (toolsynth/core) defines four abstractions:

Abstraction Role
core.MethodConfig one config subclass per method; unified precedence
core.BasePipeline / BaseStage the abstract process: run() executes the assembled stage sequence; non-linear flows override run()
toolsynth.AutoPipeline the method registry: AutoPipeline.create("toolace", overrides={...})
core.RunContext per-run context: stage trace (debug/stages.jsonl), metrics, artifact registry, event streams (ctx.append / ctx.metric)
toolsynth/
├── core/           # config (MethodConfig/LLMSettings), registry (AutoPipeline),
│                   #   RunContext, io, BasePipeline/BaseStage
├── llm/            # the single OpenAI-compatible client (retry/backoff +
│                   #   n-sampling fallback + JSON self-repair), embeddings
│                   #   (OpenAICompatEmbedder, separate-endpoint capable),
│                   #   FakeLLMClient, local NLL scoring (Eq. 1 reference)
├── schemas/        # canonical pydantic contracts (ToolSpec/State/Obs/Rubric/Trajectory/...)
│                   #   + ApiDef format adapters (ToolACE bidirectional / ACEBench / AgentInstruct)
├── generation/     # unified tool-call loop, user-simulator finish-token protocol,
│                   #   context tree, diversity operators
├── verification/   # DLV rule layer R1–R4, short-circuiting CompositeVerifier,
│                   #   n-judge voting, ROUGE-L filtering
├── evaluation/     # ACEBench 4-format parsing spine, EA/PA (LCS + monotonic
│                   #   matching strategies), graders, sqrt-weighted overall
├── environment/    # WorldModelSimulator (k2 6-step), ACEBench sandbox (4 scenarios),
│                   #   RealSandbox (K8s), static-response environments
└── methods/        # one package per paper: config.py + pipeline.py + algorithm
                    #   modules + docs/

Logging

Two layers, one knob:

  • Human console logs go through stdlib logging. run provides --log-level {DEBUG,INFO,WARNING,ERROR} (default: $TOOLSYNTH_LOG_LEVEL or INFO) — ToolACE pipeline progress, ACEBench stage lines ([construct]/[agent]/..., failures as WARNING with DEBUG tracebacks) and AgentInstruct per-seed progress (DEBUG unless AGENTINSTRUCT_VERBOSE / --verbose) all answer to this single knob. Programmatic callers use toolsynth.cli.configure_logging("INFO").
  • Machine-readable run logs are always on, independent of verbosity: every run writes per-stage timing and Alg-X provenance tags to runs/<method>/<id>/debug/stages.jsonl, plus event streams (ctx.append, e.g. rollouts.jsonl) and metrics (ctx.metric).

Testing & scope

Method-only reproduction — no training, no benchmark numbers; "runs" means the mechanisms and wiring are correct, not paper-scale data or quality metrics.

Test layers

  • Golden teststests/golden/ drives every method with a scripted FakeLLMClient and compares outputs against committed fixture JSON (tests/golden/*.json); cross-process nondeterministic fields (e.g. hash ids) are masked explicitly, and k2-agentic rollouts compare with zero masking. k3-kg has its own FakeOps golden (tests/k3_kg/fixtures/tasks.jsonl, byte-exact).
  • Full suitepython -m pytest tests/ = core-primitive unit tests + per-method unit tests + golden tests.

Which methods have had their full flow tested

Method Offline tests Real-endpoint end-to-end run
self-instruct golden ✅ real DeepSeek (deepseek-v4-flash, small budget, Aug 2026)
toolace golden ✅ real DeepSeek
k2-agentic golden (rollouts zero-mask) ✅ real DeepSeek
agent-instruct golden ⚠️ partial: Text Modification ✅ (deepseek-v4-flash-0731) and Tool Use ✅ (deepseek-v4-pro, full OpenAI-wire conversations with paired tool_call_id, artifact in runs/real/deepseek_pro_agent_instruct/, Aug 2026); still 3 of 17 skills wired
acebench 76 fake-client unit tests across the five modes ✅ all five modes ran real (Aug 2026): construct twice — capped driver (deepseek-v4-flash-0731; en passed the full quality gate, zh failed one rule check) AND the CLI-default unbounded tree (deepseek-v4-pro-0813, 605 LLM calls — expect that order of cost by default); special ×3 defect types, grade + overall with real model-outputs (Table-2 report Overall=0.387 in runs/real/deepseek_acebench/); agent-mode task refs remain placeholders (Figure 28 unpublished)
k3-kg FakeOps golden — Knowledge-QA path only ✅ Knowledge path end-to-end on real endpoints (Aug 2026): Phase A built a 20-node concept DAG (kimi-k3 + real kinfra embeddings), then one B–F iteration on glm-5.1 (12 LLM / 2 search / 16 fetch / 32 summarize calls) sampled a node, retrieved + summarized material, synthesized and gate-accepted 1 QATask — obfuscated multi-hop question per the WebSailor policy, 4 evidence docs, full Alg-4..14 event trace in runs/real/k3_ext/; the stall path was separately exercised for real (5 gate rejections → SynthesisStalled with call metrics). Caveats: the open web was substituted by a local corpus server (DuckDuckGo/Jina unreachable from this network), and only the Knowledge path ran — Coding/Vision still offline-only

Not covered by any test today: k3-kg's Coding / Vision synthesizers, the Alg-5 dedup attach branches (EQUIVALENT / PARENT / CHILD / RELATED), and the concurrent Phase-A mode (golden runs --sequential). These gaps are tracked as Roadmap items below.

Configured the API but still not real — three exceptions

  1. k2-agentic smoke defaultsk8s_sandbox=False builds an inert FakeK8sSandbox (only coding/SWE domains reach it) and embed_backend="tfidf" uses local TF-IDF; real K8s / sentence-transformers need explicit opt-in (k8s_sandbox=True, embed_backend="sentence-transformers"). The LLM calls are real either way.
  2. k3-kg Vision tasks — the code-exec verifier is a callable and cannot travel through config/CLI; only Python-API callers can pass code_exec_verifier=.... Under the CLI every Vision candidate is deterministically rejected at the gate (never emits an unverifiable task), so CLI output contains only Knowledge / Coding tasks.
  3. Local-compute components are not mocks — toolace / k2 Qwen NLL scoring, k2 TF-IDF and the ACEBench sandbox are genuine computations executed on your machine; they simply make no network/API calls.

Fidelity notes

  • Agent prompts in ToolACE / k2 / ACEBench are largely rebuilt from prose ([INFERRED]), not verbatim restorations; paper-silent hyperparameters are [UNRESOLVED] and required — no invented defaults.
  • requirements.txt is a frozen snapshot of the shared venv; pyproject.toml is authoritative (optional groups: nll / rouge / embed / finetune / sandbox / test).
  • Known semantic divergences between methods are parameterized explicitly, never silently merged:
    • Process Accuracy: LCS (ACEBench) vs monotone ordered-subset (k2 rejection sampling) → MatchingStrategy;
    • Diversity-operator choice: uniform non-empty subset (ToolACE) vs per-item 50% (k2) → strategy;
    • Eq. 1 undefined (zero-token response): None + skip (ToolACE convention) vs 0.0 (k2); the k2 path takes None (documented).

Method lineage

The Kimi K2 pipeline builds on the four methods above — every mechanism is traceable (details in toolsynth/methods/k2_agentic/docs/):

  • Stage 1 tool synthesis ← ToolACE TSS + Self-Instruct dedup
  • Stage 2 agents & tasks ← AgentInstruct personas + ToolACE H_M
  • Stage 3 trajectory generation ← ACEBench state model / user-sim + ToolACE execution
  • Stage 4 quality filtering ← ToolACE DLV + Self-Instruct + ACEBench EA/PA

Kimi K3 (§4.2.2) is a separate lineage: a self-evolving hierarchical concept DAG guides retrieval-grounded task synthesis (Knowledge QA / Coding / Vision), borrowing WebSailor / WebShaper mechanisms — reconstruction with provenance tags in toolsynth/methods/k3_kg/docs/.

Roadmap

  • Kimi K3 integrationmethods/k3_kg/ (registered k3-kg): Phase A concept-DAG build + Phase B–F loop migrated verbatim from the standalone reproduction, with a new shared llm/embedder.py primitive (separate-endpoint capable) and an additive trace hook; pinned by a FakeOps golden test.
  • Close the test-coverage gaps — k3-kg real-endpoint runs for the Coding / Vision synthesizers and against the open web (the validated run used a local corpus server); k3-kg test coverage for the Alg-5 dedup attach branches (EQUIVALENT / PARENT / CHILD / RELATED) and the concurrent Phase-A mode; wiring + real runs for agent-instruct's 14 unwired skills; an ACEBench end-to-end run once real Figure-28 tasks are available.
  • Unified data format — one canonical export for all methods' artifacts (OpenAI-messages wire shape as the baseline: tool_calls with ids, tool turns paired by tool_call_id), plus tool-call serialization for the k2 corpus via an additive sidecar stream (historical record shapes stay pinned by golden tests). Current format families and gaps: docs/data-formats.md.

References

Paper Link
Self-Instruct: Aligning Language Models with Self-Generated Instructions arXiv 2212.10560
AgentInstruct: Toward Generative Data Augmentation with Different Levels of Complexity arXiv 2407.03502
ACEBench: Who Wins the Match Point in Tool Usage? arXiv 2501.12851
ToolACE: Winning the Points of LLM Function Calling arXiv 2409.00920
Kimi K2: Open Agentic Intelligence arXiv 2507.20534
Kimi K3 §4.2.2: KG-guided task synthesis reconstruction notes in toolsynth/methods/k3_kg/docs/

Disclaimer

This project is provided for learning and research purposes only. It is an independent, unofficial reproduction of the papers listed above and is not affiliated with, endorsed by, or connected to their authors or Moonshot AI (Kimi). The code and any synthesized data are offered as-is, without warranty of any kind. Install from PyPI (pip install toolsynth) or from a checkout for development (pip install -e .).

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

toolsynth-0.1.0.tar.gz (538.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

toolsynth-0.1.0-py3-none-any.whl (617.8 kB view details)

Uploaded Python 3

File details

Details for the file toolsynth-0.1.0.tar.gz.

File metadata

  • Download URL: toolsynth-0.1.0.tar.gz
  • Upload date:
  • Size: 538.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for toolsynth-0.1.0.tar.gz
Algorithm Hash digest
SHA256 7834ede34e741a8dccf61d43a0f0a6c08ab4dad3230b98440c92e44c8a1407f8
MD5 b64a0f0eab4ff8307e31e07c79225ccc
BLAKE2b-256 e6a20661edfa841bcc98aaecd4a24c860859027ef3b0be1af8feb34e18f53744

See more details on using hashes here.

Provenance

The following attestation bundles were made for toolsynth-0.1.0.tar.gz:

Publisher: pypi.yml on aaronlyt/toolsynth

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file toolsynth-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: toolsynth-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 617.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for toolsynth-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 5d104afbc648eb69293a4ec4d21965c1706c6114e8475241e25210cbc96a30cf
MD5 7693da752fa80fc9d0b572fe878b05c9
BLAKE2b-256 8e8877a5860a0d87991b9ccecd02d97497753a0e3074940b3e78993b1e63ffd9

See more details on using hashes here.

Provenance

The following attestation bundles were made for toolsynth-0.1.0-py3-none-any.whl:

Publisher: pypi.yml on aaronlyt/toolsynth

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page