ftgate (working name)
Regression checks for a small model you fine-tuned on tool calling: do the bytes the runtime sends match the bytes the model was trained on, and does it still call your tools correctly after export and quantisation?
Status: ftgate template, ftgate data and ftgate tools work. The premise was
checked before any tool code was written — the same way
pr-witness was built.
What has been measured so far
- llama-server reproduces the HF training bytes exactly for a
tool-calling prompt (Qwen2.5-0.5B-Instruct, 250/250 tokens, text and ids
identical) — even though the GGUF's embedded template hashes differently
from HF's. Compare renderings, never hashes. →
notes/day1-findings.md - Ollama 0.40.2 hands the model a Go struct dump instead of the tool
schema:
{get_weather … {object <nil> <nil> [city] {…}}}. Root cause intemplate/template.go(templateToolFunctionhas noString()); known upstream as ollama/ollama#14601 with open PRs #18318 and #18391. - Qwen's official Qwen2.5 GGUFs embed the pre-fix chat template —
doubled braces
{{"name": …}}in the tool-call instruction, fixed on HF on 2024-09-19 but never regenerated in the GGUF repos. Ollama's library blobs are those files. Ollama never renders it; llama.cpp, vLLM or LM Studio serving the file do. The 0.5B model copies the braces into every tool call (invalid JSON; a strict parser gets 0/35 where the HF template gets 22/35); the 1.5B model ignores them. llama-server's parser forgives it, vLLM'sjson.loadswould not. →notes/day4-doubled-braces-impact.md→notes/day3-template-command.md - Ollama prepends the Modelfile SYSTEM when a request has none — a system prompt your training rows may never have had.
- It costs accuracy. Same GGUF blob in every arm, temperature 0, 40
cases: Qwen2.5-0.5B gets exact arguments 39/40 with its own template and
31/40 with the prompt Ollama builds — every failure an invented
parameter
amount from_currency, read off Go's[amount from_currency to_currency]rendering ofrequired. Qwen2.5-1.5B: 40/40 either way on this easy set. →notes/day2-impact.md
Use
uv pip install -e .
ftgate template --model Qwen/Qwen2.5-0.5B-Instruct \
--llama http://localhost:8080 \
--ollama qwen2.5:0.5b-instruct \
--dataset train.jsonl --sample 20 --diff
--model is the HF checkpoint whose tokenizer and chat template the
trainer used; --dataset is your own rows (messages[, tools] per line).
Each runtime's prompt and token ids are compared with what SFT fed the
model, and every difference is classified: schema_not_json,
injected_system, special_prefix, structural, spacing, tokenizer.
Ollama's column is rendered with Ollama's own template package (build
bin/ once) and cross-checked against the live prompt_eval_count.
ftgate data — lint the dataset with the model's own tokenizer and template
ftgate data train.jsonl --model Qwen/Qwen2.5-0.5B-Instruct --max-seq-len 2048 \
--eval eval.jsonl --fail-on error
Rows in messages, ShareGPT, Alpaca or preference (chosen/rejected)
shape. Reports what the trainer will silently do to each row: an assistant
turn that renders to zero target tokens, a conversation cut by
max_seq_len before the answer (trains nothing) or inside it, a row that
never ends with EOS, a template that raises on the row. Structural checks:
empty or missing assistant target, role order, tool calls against the
declared schema (missing required / unknown arguments / undeclared tool /
tool_call_id with no call), chosen == rejected, length bias in
preference pairs, exact duplicates, near-duplicates of the eval set.
--format json for CI; --fail-on error|warning sets the exit code.
ftgate tools — does it still call your tools, on the backend it will run on?
ftgate tools --tools tools.json --cases cases.jsonl \
--provider ollama:my-finetune:q4_k_m \
--provider "openai:m@http://localhost:8080/v1" \
--min-exact 0.8
ftgate tools --from-dataset train.jsonl --holdout 50 --provider ollama:my-finetune
Evaluation runs on promptfoo (installed
once with npm install --prefix .pfrunner promptfoo, or set
FTGATE_PROMPTFOO). ftgate writes the config — your tools, your cases
(a file, or rows held out of your own training set), one provider per
arm, temperature 0, one request in flight for local servers — and a
deterministic judge that reports per case: the right tool, every expected
argument present and equal (extra optional arguments allowed and counted),
and two things a generic eval does not separate from "the model was
wrong": parser lost — the model emitted a <tool_call> the provider
did not surface — and malformed — the call's JSON does not parse.
Failure reasons show the produced arguments next to the expected ones.
In CI, in tests, before commit
# GitHub Actions
- uses: talhayme/ftgate@main
with: { dataset: data/train.jsonl, model: Qwen/Qwen2.5-0.5B-Instruct, max-seq-len: "2048" }
# .pre-commit-config.yaml
- repo: https://github.com/talhayme/ftgate
rev: main
hooks: [{ id: ftgate-data, args: [--fail-on, error] }]
# tests/test_data.py
from ftgate.testing import assert_dataset_clean
def test_training_data():
assert_dataset_clean("data/train.jsonl", model="Qwen/Qwen2.5-0.5B-Instruct", max_seq_len=2048)
or, with the plugin that installs alongside the package,
pytest --ftgate-dataset data/train.jsonl --ftgate-model Qwen/Qwen2.5-0.5B-Instruct.
Reproduce the findings
uv venv --python 3.12 ~/.venvs/ftgate && uv pip install --python ~/.venvs/ftgate/bin/python torch transformers gguf requests
# llama-server built from llama.cpp with --jinja support; Ollama ≥ 0.40 running on :11434
python scripts/hf_side.py # training-side bytes
python scripts/llama_side.py http://localhost:8089
python scripts/ollama_side.py # needs bin/ollama-render (see bin/README.md)
python scripts/day2_eval.py qwen2.5:0.5b-instruct # three-arm accuracy comparison
Plan
ftgate template — the three-sided byte comparison as a command.
ftgate data — a linter for SFT/DPO/tool-call datasets using the target
model's tokenizer and template. Tool-call evaluation builds on promptfoo's
tool-call-f1 / trajectory:tool-args-match rather than duplicating them.
Apache-2.0.
Metadata
Release files for ftgate 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| ftgate-0.1.0.tar.gz | 35.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| ftgate-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 68.2 kB
Release files / ftgate-0.1.0.tar.gz
| Download URL | ftgate-0.1.0.tar.gz |
|---|---|
| Size | 35.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
34de16b57746a209bfd96855ca256680baa1c127660ff0f7d7a2d224484388ff
|
|
BLAKE2b-256 checksum How to use checksums |
6e8e23893cb24d68eb8a34232561b1c1c2c96fb36aa2d4432e9fc65019beb830
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / ftgate-0.1.0-py3-none-any.whl
| Download URL | ftgate-0.1.0-py3-none-any.whl |
|---|---|
| Size | 32.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a41293fe5ccbda7d8a781733fccda63c37d8e57dab1169da9bfc56d20f372e15
|
|
BLAKE2b-256 checksum How to use checksums |
2b75a2a431595fdc9f37c637b7e2e36b86f2db094f9a85e89d12f4b6ad563eb2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|