Skip to main content

ftgate

Regression checks for a small model you fine-tuned on tool calling: do the bytes the runtime sends match the bytes the model was trained on, and does it still call your tools correctly after export and quantisation?

Status: ftgate template, ftgate data and ftgate tools work. The premise was checked before any tool code was written — the same way pr-witness was built.

What has been measured so far

  • llama-server reproduces the HF training bytes exactly for a tool-calling prompt (Qwen2.5-0.5B-Instruct, 250/250 tokens, text and ids identical) — even though the GGUF's embedded template hashes differently from HF's. Compare renderings, never hashes. → notes/day1-findings.md
  • Ollama 0.40.2 hands the model a Go struct dump instead of the tool schema: {get_weather … {object <nil> <nil> [city] {…}}}. Root cause in template/template.go (templateToolFunction has no String()); known upstream as ollama/ollama#14601 with open PRs #18318 and #18391.
  • Qwen's official Qwen2.5 GGUFs embed the pre-fix chat template — doubled braces {{"name": …}} in the tool-call instruction, fixed on HF on 2024-09-19 but never regenerated in the GGUF repos. Ollama's library blobs are those files. Ollama never renders it; llama.cpp, vLLM or LM Studio serving the file do. The 0.5B model copies the braces into every tool call (invalid JSON; a strict parser gets 0/35 where the HF template gets 22/35); the 1.5B model ignores them. llama-server's parser forgives it, vLLM's json.loads would not. → notes/day4-doubled-braces-impact.md → notes/day3-template-command.md
  • Ollama prepends the Modelfile SYSTEM when a request has none — a system prompt your training rows may never have had.
  • It costs accuracy. Same GGUF blob in every arm, temperature 0, 40 cases: Qwen2.5-0.5B gets exact arguments 39/40 with its own template and 31/40 with the prompt Ollama builds — every failure an invented parameter amount from_currency, read off Go's [amount from_currency to_currency] rendering of required. Qwen2.5-1.5B: 40/40 either way on this easy set. → notes/day2-impact.md

Use

pip install ftgate
ftgate template --model Qwen/Qwen2.5-0.5B-Instruct \
                --llama http://localhost:8080 \
                --ollama qwen2.5:0.5b-instruct \
                --dataset train.jsonl --sample 20 --diff

--model is the HF checkpoint whose tokenizer and chat template the trainer used; --dataset is your own rows (messages[, tools] per line). Each runtime's prompt and token ids are compared with what SFT fed the model, and every difference is classified: schema_not_json, injected_system, special_prefix, structural, spacing, tokenizer. Ollama's column is rendered with Ollama's own template package (build bin/ once) and cross-checked against the live prompt_eval_count.

ftgate data — lint the dataset with the model's own tokenizer and template

ftgate data train.jsonl --model Qwen/Qwen2.5-0.5B-Instruct --max-seq-len 2048 \
            --eval eval.jsonl --fail-on error

Rows in messages, ShareGPT, Alpaca or preference (chosen/rejected) shape. Reports what the trainer will silently do to each row: an assistant turn that renders to zero target tokens, a conversation cut by max_seq_len before the answer (the answer is never trained) or inside it, a row that never ends with EOS, a template that raises on the row. Structural checks: empty or missing assistant target, role order, tool calls against the declared schema (missing required / unknown arguments / undeclared tool / tool_call_id with no call), chosen == rejected, length bias in preference pairs, exact duplicates, near-duplicates of the eval set. --format json for CI; --fail-on error|warning sets the exit code.

ftgate tools — does it still call your tools, on the backend it will run on?

ftgate tools --tools tools.json --cases cases.jsonl \
             --provider ollama:my-finetune:q4_k_m \
             --provider "openai:m@http://localhost:8080/v1" \
             --min-exact 0.8
ftgate tools --from-dataset train.jsonl --holdout 50 --provider ollama:my-finetune

Evaluation runs on promptfoo (installed once with npm install --prefix .pfrunner promptfoo, or set FTGATE_PROMPTFOO). ftgate writes the config — your tools, your cases (a file, or rows held out of your own training set), one provider per arm, temperature 0, one request in flight for local servers — and a deterministic judge that reports per case: the right tool, every expected argument present and equal (extra optional arguments allowed and counted), and two things a generic eval does not separate from "the model was wrong": parser lost — the model emitted a <tool_call> the provider did not surface — and malformed — the call's JSON does not parse. Failure reasons show the produced arguments next to the expected ones.

In CI, in tests, before commit

# GitHub Actions
- uses: talhayme/ftgate@v0.1.2
  with: { dataset: data/train.jsonl, model: Qwen/Qwen2.5-0.5B-Instruct, max-seq-len: "2048" }
# .pre-commit-config.yaml
- repo: https://github.com/talhayme/ftgate
  rev: v0.1.2
  hooks: [{ id: ftgate-data, args: [--fail-on, error] }]
# tests/test_data.py
from ftgate.testing import assert_dataset_clean

def test_training_data():
    assert_dataset_clean("data/train.jsonl", model="Qwen/Qwen2.5-0.5B-Instruct", max_seq_len=2048)

or, with the plugin that installs alongside the package, pytest --ftgate-dataset data/train.jsonl --ftgate-model Qwen/Qwen2.5-0.5B-Instruct.

Reproduce the findings

uv venv --python 3.12 ~/.venvs/ftgate && uv pip install --python ~/.venvs/ftgate/bin/python torch transformers gguf requests
# llama-server built from llama.cpp with --jinja support; Ollama ≥ 0.40 running on :11434
python scripts/hf_side.py                       # training-side bytes
python scripts/llama_side.py http://localhost:8089
python scripts/ollama_side.py                   # needs bin/ollama-render (see bin/README.md)
python scripts/day2_eval.py qwen2.5:0.5b-instruct   # three-arm accuracy comparison

Plan

ftgate template — the three-sided byte comparison as a command. ftgate data — a linter for SFT/DPO/tool-call datasets using the target model's tokenizer and template. Tool-call evaluation builds on promptfoo's tool-call-f1 / trajectory:tool-args-match rather than duplicating them.

Apache-2.0.

Metadata

Release files for ftgate 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ftgate 0.1.2
File Size Uploaded
ftgate-0.1.2.tar.gz 35.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ftgate 0.1.2
File Interpreter ABI Platform
ftgate-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 68.2 kB

Release files / ftgate-0.1.2.tar.gz

Download URL ftgate-0.1.2.tar.gz
Size 35.6 kB
Tags Source
SHA-256 checksum
How to use checksums
f5e61d4337610ff4e764670615fa89df075a6e1c2afb8ba31d12e010ce8fa673
BLAKE2b-256 checksum
How to use checksums
e982a79536b544019d8738d3e238bf02ac360e46e79fe09a691c2fbaab19c1a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / ftgate-0.1.2-py3-none-any.whl

Download URL ftgate-0.1.2-py3-none-any.whl
Size 32.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7c76fdaf034a723076e06e5ea0907322b6a9c954087c50ee88ece9f07576030c
BLAKE2b-256 checksum
How to use checksums
1682a0980f3415bcb90972e86f67089c6d0cd124280fb59f811623efeb46ed3b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

0.1.3

2 release files

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page