This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.
Consider using release 0.2.2 instead.
Reason given by maintainers: Inacurate description
githeri
Spec-Forge: a pipeline that trains an AI to bridge the gap between a human's natural-language feature request and a machine-executable specification that feeds the COMMAND_RUNWAY skill.
The central thesis: most time in AI-assisted development is lost before the AI and the human agree on what to build. Githeri attacks that by producing validated, structured YAML specs — which then feed into COMMAND_RUNWAY runbooks that an executor agent follows verbatim.
Pipeline
Human natural-language request
│
� ▼
Spec-Forge (LLM + validator)
Generates:
data/training_data.jsonl (valid prompt + spec_yaml pairs)
data/failed_specs.jsonl (invalid specs, saved for analysis)
│
� ▼
Runbook Scorer (standalone, post-generation)
Scores:
5 categories — Intent, Preconditions, Structure, Testability, Coverage
Hard gate: missing Inspect/Create/Verify = 0.0
│
� ▼
Human Review
"L1-L4 look good. Approve."
│
� ▼
COMMAND_RUNWAY Skill
Consumes:
a validated spec (single-feature YAML)
Produces:
COMMAND_RUNWAY.md
• ordered implementation plan
• exact file paths
• code modifications
• test skeletons
• verification commands (translated from spec local_goals)
• rollback guidance
• completion criteria
│
� ▼
GRG Executor (with COMMAND_RUNWAY pattern integration)
Consumes:
a validated spec OR COMMAND_RUNWAY plan JSON
Produces (all under foreign/ directory):
• implementation source files
• test files
• RUNBOOK.md (human-readable execution log with GRG scores)
• RUNBOOK.json (machine-readable execution data)
• automatic ruff check --fix on generated code
│
� ▼
Completed Feature
Outputs:
• implementation complete
• all tests passing
• OpenAPI updated
• documentation synchronized
• human notified
Spec Enrichment (IMPROVE_SPEC)
Every generated spec now includes optional enrichment fields that make specs machine-executable:
| Field | Location | Purpose |
|---|---|---|
business_rules |
top-level | Invariants & formulas (e.g., "JWT Secret: 256-bit random, rotated quarterly") |
test_fixtures |
top-level | Seed data & setup commands (e.g., .venv/bin/python scripts/seed_admin.py) |
environment |
top-level | Required packages + env vars (e.g., pyyaml>=6.0, JWT_SECRET) |
global_verification |
top-level | Post-execution gate commands (e.g., pytest tests/, bandit -r src/) |
blueprint |
per-goal | Required for type: create (≥100 chars). Code-level outline: class signatures, route decorators, SQLAlchemy models, business logic steps |
acceptance_criteria |
per-goal | List of {test, steps} — executable test cases in pseudo-code |
type |
per-goal | create | update | delete | inspect | verify — drives runbook stage classification |
These fields are validated by scripts/validator.py and consumed by downstream generators (plan, runbook, scorer).
Runbook Scoring System
Every generated spec is scored against runbook-readiness criteria (see docs/scoring_spec.md). Scoring is decoupled from generation — specs are saved first, then scored in a separate pass via make score.
The scorer (scripts/runbook_scorer.py) evaluates five weighted categories:
| Category | Weight | Key Checks |
|---|---|---|
| Intent & Goals | 20% | Summary present, goals have descriptions, endpoint tasks have HTTP verification |
| Preconditions | 15% | depends_on references valid globals/stages, CLI tools declared in context |
| Command Runway Structure | 30% | Hard gate: must have Inspect (file_exists/read CLI), Create/Modify (build CLI), Verify (HTTP/test CLI). Stage order: Inspect → Create → Verify |
| Verification Testability | 25% | Concrete commands, explicit assertions (status/exit_code/content), reproducible URLs |
| Completion Coverage | 10% | Prompt-mentioned status codes, tests, OpenAPI updates reflected in spec |
Hard gate: If any of the three runway stages (Inspect, Create/Modify, Verify) is missing, the spec scores 0.0 and is not runbook-ready.
The scorer now honors explicit type on goals — a goal with type: create counts as Create/Modify even if its verification is file_exists (executor will generate the file). Similarly type: inspect and type: verify map directly to stages.
Quick Start
1. Setup
git clone <repo-url> && cd githeri
python3.11 -m venv .venv && .venv/bin/pip install -r requirements.txt
# For local generation (Ollama):
ollama pull specgen:latest
ollama pull qwen2.5-coder:7b-instruct # base model (fallback + code execution)
# For cloud generation (NVIDIA NIM - host any model like minimaxai/minimax-m3):
export NVIDIA_API_KEY=your_nvidia_api_key
# For LM Studio local execution (GRG pipeline):
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
# Load qwen2.5-coder-14b-instruct-uncensored model
# For training: install on a GPU machine
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"
2. Generate specs
# Generate a validated spec from a fresh NL prompt (Ollama, specgen:latest)
make spec PROMPT="Add a POST /register endpoint that accepts email and password"
# Use NVIDIA NIM (default model: minimaxai/minimax-m3, base URL: https://integrate.api.nvidia.com/v1)
make spec PROMPT="Add a PATCH /users/{id}/settings endpoint" PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY
# Use any OpenAI-compatible endpoint
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra
# Generate N specs from random seed prompts (475 prompts across 21 categories)
make generate N=10
make generate N=100
# Generate ALL seed prompts in sequential order (full corpus generation)
make generate N=all
make generate-all
# Generate 10 random specs (explicit alias, good for short test runs)
make generate N=random
make generate-random
# Validate all specs in the corpus
make validate
# Score all specs against runbook criteria (standalone, post-generation)
make score
make score-failed # score invalid specs in data/failed_specs.jsonl
# Generate COMMAND_RUNWAY plan from a validated spec (two-step: prompt → LLM)
make plan-from-spec SPEC=data/training_data.jsonl#8
# Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out
# Output: ./out/PLAN.md — use PLAN_MODEL=qwen2.5-coder:7b-instruct for stronger planning
3. Convert + fine-tune + export
# Convert training data to chat format (filters by runbook score)
make convert-chat # default MIN_SCORE=0.75
make convert-chat MIN_SCORE=0.9 # only high-quality specs
# GRG Agent code training (generates verified code with real GRG scores)
make generate-code # cycles through all 475 SEED_PROMPTS
make convert-code-chat # converts passing examples to chat format
# LoRA fine-tuning (requires Unsloth + GPU with 8GB VRAM)
make train # outputs models/qwen2.5-coder-7b-specforge/
# Merge adapter + export GGUF for Ollama
make merge # outputs models/qwen2.5-coder-7b-specforge-gguf/
# Evaluate fine-tuned vs base model on held-out prompts
make eval-model # saves data/eval_results.json
# Register fine-tuned model in Ollama
ollama create specgen -f models/specgen/Modelfile
4. Upload to HuggingFace Hub
Model weights are stored in git (not LFS). The upload sends standard model files to HF Hub.
# Set up auth
echo "HF_TOKEN=hf_your_token_here" > .env
# Upload
make upload-hf REPO=githeri/specgen
make upload-hf REPO=githeri/specgen PRIVATE=1 # private repo
5. Install skill for agent use
The skills/spec-forge/ directory is a self-contained Spec-Forge skill. Install it into your Hermes skills directory to use it from any session.
# Install to ~/.hermes/skills/ (default Hermes skills directory)
make install-skill # copies skills/spec-forge/ to ~/.hermes/skills/spec-forge/
# Install to a custom path (e.g. new server, project-local skills dir)
make install-skill-to-server HERMES_SKILLS_DIR=/path/to/hermes/skills
# Remove
make uninstall-skill
After install, load the skill in any Hermes session:
skill_view(name='spec-forge')
The skill provides: make spec PROMPT="..." (fresh NL → validated spec), make spec-and-plan PROMPT="..." (spec + plan prompt), and the bundled validator (scripts/validator.py) + plan assembler (scripts/plan_from_spec.py).
6. Plan from existing specs
Generate a COMMAND_RUNWAY plan from any validated spec in the corpus (two-step: extract + LLM plan generation):
# Default: specgen:latest, ollama, ./out/PLAN.md
make plan-from-spec SPEC=data/training_data.jsonl#8
# Stronger planning model (if available locally)
PLAN_MODEL=qwen2.5-coder:7b-instruct make plan-from-spec SPEC=data/training_data.jsonl#8
7. Interactive coding agent (scripts/githeri.py)
The project root ships a real interactive coding agent (like Codex/Claude Code). It has a persistent conversation, file operations, command execution, and model/provider switching inside the app.
# Interactive mode — start a conversation
.venv/bin/python scripts/githeri.py
# In-session commands:
# /model <name> - Change model (e.g. /model qwen2.5-coder:7b-instruct)
# /provider <name> - Change provider (e.g. /provider ollama)
# /clear - Clear conversation
# /help - Show help
# /quit, /exit - Exit
# Non-interactive: generate spec from NL prompt via specgen, then execute it
# Only two knobs: --prompt and exec model via env vars. Docker + spec model hardcoded.
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py \
--prompt "Add a PATCH endpoint to update user displayName and bio"
# Execute an existing spec.yaml (skip spec generation)
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
# Switch exec model — same prompt, different execution model
GITHERI_MODEL=qwen2.5-coder-14b-instruct-uncensored \
GITHERI_PROVIDER=lmstudio \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
# Overrides (rare):
# --no-docker - Disable Docker (not recommended)
# --workdir <path> - Custom workdir (default: output/sandbox)
# --docker-image <img> - Custom Docker image (default: githeri-sandbox)
# --server - Start server container after agent loop
# --test-cmd "<cmd>" - Run test command inside container via docker exec
Defaults (hardcoded, never change between tests):
--docker: True— every command runs in throwawaygitheri-sandboxcontainer--workdir: output/sandbox— all file ops jailed here--spec-model: specgen:latest— hardcoded (fine-tuned for spec generation)--spec-provider: ollama— hardcoded- Exec model/provider:
GITHERI_MODEL/GITHERI_PROVIDERenv vars only
Pipeline flow: 1) specgen:latest (Ollama) → NL prompt → spec.yaml; 2) exec model reads spec, creates files, installs deps, runs verifications — all inside Docker; 3) optional server container; 4) optional test command via docker exec.
See docs/SPECGEN_HARNESS.md for full architecture, troubleshooting, and output structure.
Provider Configuration
All generation targets accept provider overrides:
# NVIDIA NIM (default: minimaxai/minimax-m3 at integrate.api.nvidia.com/v1)
make spec PROMPT="..." PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY
# OpenAI GPT-4o
make spec PROMPT="..." PROVIDER=openai API_KEY=$OPENAI_API_KEY MODEL=gpt-4o
# Anthropic Claude
make spec PROMPT="..." PROVIDER=anthropic API_KEY=$ANTHROPIC_API_KEY MODEL=claude-3-5-sonnet-20241022
# LM Studio (local, OpenAI-compatible on port 1234)
make spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
# Any OpenAI-compatible endpoint (Fireworks, Together, vLLM, self-hosted NIM)
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra
# Override sampling
make generate N=5 PROVIDER=nvidia TEMPERATURE=0.3 MAX_TOKENS=4096
# NVIDIA NIM models with longer cold start: use bigger TIMEOUT
make generate N=1 PROVIDER=nvidia TIMEOUT=600
Environment variable fallbacks:
PROVIDER→ollama(default)MODEL→ provider-specific defaultAPI_KEY→OPENAI_API_KEY,NVIDIA_API_KEY,ANTHROPIC_API_KEYBASE_URL→OPENAI_BASE_URL,NVIDIA_BASE_URL(defaulthttps://integrate.api.nvidia.com/v1)TEMPERATURE→0.2MAX_TOKENS→2048
Autonomous Execution System
The autonomous execution system takes a natural language prompt and produces a working implementation without human intervention in the loop. It uses the specgen pipeline to generate a validated spec, then uses the Command Runway skills to generate a plan and runbook, and finally executes the runbook using the GRG executor (with self-healing capabilities via GRG quality gates).
For detailed instructions, see docs/SPRINTS_AUTONOMOUS.md.
Quick Reference
# Basic local execution
.venv/bin/python scripts/spexecutor.py --prompt "Add a POST /notifications endpoint"
# Docker-isolated execution
.venv/bin/python scripts/spexecutor.py --prompt "Add user authentication" --docker
# Using Hermes as the execution backend
.venv/bin/python scripts/spexecutor.py --prompt "Implement rate limiting" --executor hermes
# Makefile: generate spec and plan
make spec-and-plan PROMPT="Add file upload endpoint"
# Makefile: full autonomous cycle (spec -> plan -> runbook -> execute -> report)
make autonomous-cycle SPEC=specs/test-endpoint.yaml
# GRG isolated execution (recommended for clean workspace)
make grg-full PROMPT="Add a POST /webhook endpoint that validates signature" PROVIDER=ollama
# Outputs: foreign/src/, foreign/tests/, foreign/RUNBOOK.md, foreign/RUNBOOK.json
Recent Enhancements
2026-08-29 (This Session)
COMMAND_RUNWAY Plan Generation from Existing Specs (make plan-from-spec)
Added make plan-from-spec SPEC=data/training_data.jsonl#<index> — generates a full
COMMAND_RUNWAY plan from any validated spec in the corpus, in two steps:
- Extract — load spec_yaml from JSONL by index (0-based) via
scripts/plan_from_spec.py - Generate — build plan prompt + call LLM (
scripts/plan_from_spec_file.py→generate_plan())
Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out.
Output: ./out/PLAN.md — full COMMAND_RUNWAY document with Feature, Target Environment,
and ordered Execution Stages (Objective, Verification, Completion Condition per stage).
Override for stronger plans: PLAN_MODEL=qwen2.5-coder:7b-instruct PLAN_PROVIDER=ollama.
Tested: 168-line plan with 3 execution stages generated for session-notes spec
(4176 chars, specgen:latest, ~60s). Requires specgen:latest (or any Ollama model)
to be pulled locally.
This unblocks the Sprint 9 workflow (prompt synthesis → bulk spec generation → bulk plan generation → consolidated COMMAND_RUNWAY document) by providing the per-spec plan generation step as a reusable Makefile target.
Scorer Fix: Accept verification.type: content as file_exists Alias
Both qwen2.5-coder:7b-instruct and specgen:latest emit verification.type: "content"
with path + expect.content_contains — structurally identical to file_exists but
not recognized by the scorer's canonical vocab. Fixed scripts/runbook_scorer.py:
- Added
"content"toREQUIRED_EXPECT_KEYS(same keys asfile_exists) _check_inspect()now treatscontentas inspect (read-only check)_check_verify()treatscontentwith content_contains as verify evidence- Verification testability section handles
contentsame asfile_exists
Before: avg score 0.104, 1/9 above 0.75. After: avg score 0.972, 9/9 above 0.75.
Sprint 9 — Prompt Synthesis & Decomposition Pipeline (Planning)
Created docs/SPRINTS.md Sprint 9 (9a-9d) — four sub-sprints to accept raw unfocused
user text end-to-end:
- 9a
scripts/prompt_synthesizer.py— raw text → decomposed clean prompts (JSONL) - 9b
scripts/bulk_generate.py— batch spec generation from prompt file - 9c
scripts/bulk_plan.py— loop over specs → plans → consolidated doc - 9d
make text-to-plan TEXT="..."— single target: synthesize → generate → score → plan
Independent of Sprints 5-8. Default models: specgen:latest for generation + planning.
Sprints 5, 6, 7 Marked Complete
- Sprint 5 (Model provider switch) — already built: run_pipeline.py supports 7 providers
- Sprint 6 (Skill bundling/install) —
install-skilltarget body added to Makefile - Sprint 7 (Fine-tuning pipeline) — train*.py, merge*.py, eval*.py all exist; Makefile targets wired
scripts/githeri.py now supports a full end-to-end pipeline from natural language to tested code. Spec generation is always handled by specgen:latest on Ollama (finetuned for structured YAML output). All execution and tests run inside a Docker sandbox by default — the container is throwaway, so the host .venv/.git are never polluted across multiple test runs.
# Minimal invocation — only the prompt and exec model vary between tests.
# Everything else (docker, workdir, spec model/provider) is defaulted.
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py \
--prompt "Add a PATCH endpoint to update user displayName and bio"
Defaults (set once, never touched between tests):
- --docker: True — every run_command runs in throwaway githeri-sandbox container
- --workdir: output/sandbox — all file ops jailed here; created under repo root
- --spec-model: specgen:latest — hardcoded (finetuned for spec generation only)
- --spec-provider: ollama — hardcoded
- Exec model/provider: env vars GITHERI_MODEL / GITHERI_PROVIDER — only things you change between tests
- --docker-image: githeri-sandbox
Only two knobs per test: --prompt (or --spec) and the exec model via env vars. No flags for docker, workdir, or spec model — they're locked in.
# Switch exec model — same prompt, same defaults, different model
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
# Different provider for execution (specgen stays ollama/specgen)
GITHERI_MODEL=qwen2.5-14b-instruct-latest:2 \
GITHERI_PROVIDER=lmstudio \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"
# Execute an existing spec.yaml instead of generating one
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml
Overrides (rare):
# Disable docker (e.g. when Docker daemon is down) — not recommended for tests
.venv/bin/python scripts/githeri.py --prompt "..." --no-docker
# Custom workdir
.venv/bin/python scripts/githeri.py --prompt "..." --workdir output/custom-sandbox
# Custom docker image
.venv/bin/python scripts/githeri.py --prompt "..." --docker-image my-sandbox
Pipeline flow:
- Specgen LLM (specgen:latest / Ollama) converts NL prompt → spec.yaml (with entrypoint, packages, goals)
- Execution model (set via env vars) reads spec, creates files in output/sandbox, installs deps, runs verifications — all inside Docker
- Server container starts from entrypoint.start_command (if --server)
- Test command runs inside container via docker exec (if --test-cmd)
Environment isolation (critical):
- Docker is the default and recommended mode — every
run_command(includingpip install) runs in a throwaway container. The host.venvis never touched by spec-generated dependencies. --no-dockerdisables this isolation. When used, pip installs and code execution happen in the host environment (the project.venv), which can pollute it with spec-generated packages (Flask, FastAPI, etc.). Use--no-dockeronly when you understand this risk.- If you must run without Docker, create a separate virtualenv for the workdir:
python -m venv /tmp/githeri-isolated source /tmp/githeri-isolated/bin/activate pip install -r requirements.txt # only githeri's own deps GITHERI_MODEL=... GITHERI_PROVIDER=... \ scripts/githeri.py --prompt "..." --no-docker --workdir /tmp/githeri-workdir
Model separation: Spec generation uses specgen:latest (trained for structured YAML output, hardcoded). Execution uses whatever you set via GITHERI_MODEL / GITHERI_PROVIDER. Override via env vars only — the --spec-model / --spec-provider flags are suppressed (backwards-compatible but ignored).
Server lifecycle: After the agent loop completes, githeri parses entrypoint.start_command from the generated spec, starts a named Docker container (docker run -d), and runs --test-cmd inside it. Container stays running for manual inspection: docker exec bash.
See docs/SPECGEN_HARNESS.md for full architecture, troubleshooting, and output structure.
2026-08-15 (This Session)
GRG Agent Skill — Hermes Native Integration
The GRG agent skill is now installed as a native Hermes skill (~/.hermes/skills/autonomous-ai-agents/grg_agent/) with full multi-provider support:
- Skill Installation — Copied from
skills/grg_agent/and installed via editable pip install - Lightweight GRG Dependency — Uses local
grg-0.1.0-py3-none-any.whlwheel (no karakana dependency) providingAlphaMomentumTrackerandcompute_structural_alpha - Hermes Proxy Support — Skill accepts
providerargument:ollama|hermes|auto— uses Hermes's configured providers (NVIDIA Nemotron, Nous Portal, xAI Grok, etc.) instead of local models - Direct GRG Agent Execution —
grg:executecommand now callsself.agent.solve()directly instead of legacyrun_pipeline.pysubprocess - Make Target Integration —
scripts/grg_make_spec.pyupdated to use the skill with provider argument:make grg-spec PROMPT="..." PROVIDER=ollama|hermes|auto - Project Virtual Environment — Runs in project's own
.venv/(not external karakana venv)
Key Benefit: You can now use cloud models via Hermes proxy (hermes proxy start) instead of relying on locally installed Ollama models. The skill routes through Hermes's provider config which supports NVIDIA Nemotron, Nous Portal, xAI Grok, and any OpenAI-compatible endpoint.
LM Studio Local Model — First End-to-End Working Pipeline
LM Studio with qwen2.5-coder-14b-instruct-uncensored is the first local model to complete the full GRG pipeline end-to-end:
# Full pipeline from NL prompt → validated spec → plan → execution → runbook
make grg-full PROMPT="Implement a FastAPI POST /api/health-check endpoint..." \
PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
What works:
- GRG Agent solving — Generates implementation code with GRG quality gates (composite scoring, diversity control, convergence detection)
- Code verification — Execution-based verification (syntax + runtime) in isolated temp files
- Multi-strategy generation — Standard, decompose, test_first, refine strategies with adaptive temperature
- Health-check endpoint example — Generated FastAPI code with SQLAlchemy DB check + Redis cache check, returns 200/503
- All artifacts isolated in
foreign/— Clean workspace separation
Configuration for LM Studio:
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
# Load qwen2.5-coder-14b-instruct-uncensored model
# Then run:
make grg-spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
Why this model works:
- 14B parameter coder model fine-tuned for code generation
- Uncensored variant removes alignment filters that can interfere with code structure
- OpenAI-compatible API in LM Studio works with GRG's
OllamaClient(customapi_keysupport) - Sufficient context window for spec + plan generation tasks
- Produces valid imports, proper error handling, and correct HTTP status codes
2026-07-30 (Latest)
THIRD_IMPROVE_SPEC — 7 Pipeline Fixes
After analyzing a 10-spec batch run, identified 7 recurring failure patterns and fixed all of them:
- Minimal structural skeleton in SYSTEM_PROMPT — eliminated top-level field confusion
- task_id validator check — rejects L11/G18-style IDs, requires descriptive slug like
jwt-auth-login - Verification types & expect keys table — explicit vocabulary reduces hallucinated keys
- Placeholder ban strengthened — concrete examples in prompt (
JWT_SECRET: "test-secret-...",DATABASE_URL: "postgresql://user:***@...") + validator hint - Helpful error hints — "These fields belong inside a goal under
local_goals, not at the spec root" - depends_on validator hints — rejects L/G refs, redirects to task_ids/stage names
- Structural acceptance_criteria template — local goal field, not top-level
All 90 tests passing.
Generation Mode Aliases
Added make generate N=random and make generate N=all flags plus generate-random / generate-all aliases for short tests and full corpus runs. 475 seed prompts across 21 categories available.
Model Default Reverted
OLLAMA_MODEL reverted to qwen2.5-coder:7b-instruct (was regressed to qwen3.5-4b-128k:latest). Better structure compliance after prompt fixes.
2026-07-30 (Earlier)
Spec Enrichment (IMPROVE_SPEC)
Added five top-level enrichment fields (business_rules, test_fixtures, environment, global_verification) and three goal-level fields (blueprint, acceptance_criteria, type). All validated and scored.
Multi-Provider LLM Support
run_pipeline.py now supports Ollama, OpenAI, Anthropic, NVIDIA (via Together AI), and any OpenAI-compatible endpoint. Configured via --provider CLI arg or Makefile variables.
Runbook Scorer Stage Detection
Scorer now honors explicit type: create|inspect|verify on goals, fixing false "missing stage" penalties for file_exists verification on CREATE goals.
Validator Hardening
- Guarded against
expectbeing a string instead of dict (preventsAttributeError: 'str' object has no attribute 'get') - Near-duplicate detection now handles malformed
expectblocks defensively - All 90 tests pass
Model Compatibility Matrix
| Model | Context | Tools | Speed (M1 16GB) | Spec Gen | Recommended Use |
|---|---|---|---|---|---|
| specgen:latest | 128K | Yes | Fast | Best first-attempt pass rate (fine-tuned) | Default local (Ollama) — spec generation |
| qwen2.5-coder:7b-instruct | 32K | Yes | Fast | Works (best structure compliance) | |
| qwen2.5-coder-14b-instruct-uncensored | 32K | Yes | Medium | Works end-to-end (GRG pipeline) | |
| qwen3.5-4b-128k | 128K | Yes | Fast | Works | |
| qwen3.5-9b-code:128k | 128K | Yes | Slow | Excellent | |
| deepseek-r1:7b | 128K | TBD | Fast | YAML syntax errors | |
| Nemotron 3 Ultra | 128K | Yes | Fast | Excellent |
Current recommendation:
- Spec generation (default) →
specgen:lateston Ollama (fine-tuned for structured YAML output, best first-attempt pass rate) - Bulk corpus →
specgen:lateston Ollama (fine-tuned, best first-attempt pass rate) - Specific features (GRG pipeline) →
qwen2.5-coder-14b-instruct-uncensoredon LM Studio withmake grg-full(100% success via execution verification) - Cloud →
minimaxai/minimax-m3via NVIDIA NIM (--provider nvidia)
Training Data Generation Strategies
Based on empirical testing (2026-08-16), here are the reliable approaches:
1. Bulk Corpus Generation (Recommended for Training Data)
# Fast, decent success rate, fine-tuned for this pipeline
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434" TEMPERATURE=0.2 MAX_TOKENS=4096
- Success rate: ~70%+ (fine-tuned specgen:latest)
- Speed: ~50-100s/spec
- Best for: Generating large training corpora quickly
2. High-Quality Individual Specs (GRG Pipeline)
# Best for specific features - multi-strategy + execution verification
make grg-full PROMPT="Your feature" PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
- Success rate: ~100% (execution-verified)
- Speed: ~260s/spec
- Best for: Critical features requiring guaranteed working specs
3. GRG Agent Code Training (NEW — Real GRG Scores + Verified Code)
# Generates verified code with real GRG composite scores (not dummy 0.5)
# Uses qwen2.5-coder:7b-instruct on Ollama for real logprobs
make generate-code # cycles through all 475 SEED_PROMPTS
make convert-code-chat # converts passing examples to chat format
- Success rate: Variable (depends on prompt complexity)
- Speed: ~30-60s/prompt (optimized: 1 strategy, 1 candidate)
- Output:
data/training_data_code.jsonl+data/training_data_code_chat.jsonl - Key difference: Produces executable code with real GRG composite scores (0.47-0.49) because the model provides logprobs
- Verification: Syntax check + execution test (import + basic run)
- Best for: Fine-tuning code generation models with GRG quality signals
Configuration (in scripts/generate_code_training_fast.py):
skill = create_skill(config={
'llm_provider': 'ollama',
'ollama_base_url': 'http://127.0.0.1:11434/v1',
'ollama_default_model': 'qwen2.5-coder:7b-instruct',
'max_iterations': 2, # Must be >=2 for convergence check
'temperature': 0.3,
'top_p': 0.9,
'max_tokens': 1024,
'candidates_per_strategy': 1, # Speed optimization
'max_strategies': 1, # Speed optimization
})
4. Hybrid Workflow (Best of Both)
# 1. Generate bulk corpus with specgen
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434"
# 2. Score and filter high-quality specs
make score MIN_SCORE=0.75
# 3. Re-generate failed critical specs with GRG pipeline
make grg-full PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
# 4. Generate verified code for fine-tuning
make generate-code
make convert-code-chat
Observed Failure Patterns (specgen:latest)
The remaining failures (if any) are primarily:
- Near-duplicate HTTP verifications - multiple goals hitting same endpoint with same method
- Missing
acceptance_criteriafor CREATE goals - Placeholder values in headers (e.g.,
Authorization: *** ***) - YAML block mapping errors - CLI verification indentation issues
For More Information
- docs/SPRINTS_AUTONOMOUS.md — sprint breakdown, model experiments, decisions, next steps
- docs/scoring_spec.md — runbook scoring specification
- docs/IMPROVE_SPEC.md — spec enrichment field specification (v1)
- docs/SECOND_IMPROVE_SPEC.md — enrichment field enforcement (v2)
- docs/THIRD_IMPROVE_SPEC.txt — 7 pipeline failure patterns + fixes (v3)
- MODEL_CARD.md — model card (uploaded to HF Hub)
- skills/spec-forge/SKILL.md — Spec-Forge skill reference
- skills/command-runway-pattern/SKILL.md — Command Runway pattern skill
- skills/grg_agent/SKILL.md — GRG Agent skill reference
Metadata
Release files for githeri 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| githeri-0.2.0.tar.gz | 80.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| githeri-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 137.1 kB
Release files / githeri-0.2.0.tar.gz
| Download URL | githeri-0.2.0.tar.gz |
|---|---|
| Size | 80.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
229f6668a9338f345e77769f11a2fda4170df3745e93486eade8485eb1308b0a
|
|
BLAKE2b-256 checksum How to use checksums |
15792419d6724a271ec3f5f1f15ad27792c79f45a5d70192b6c0bdcc5cf01e18
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.13
|
Release files / githeri-0.2.0-py3-none-any.whl
| Download URL | githeri-0.2.0-py3-none-any.whl |
|---|---|
| Size | 56.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ee42ba34d36d3b2042e1e465326a9c0425d2192bc3bfd165a80cddc97166eee7
|
|
BLAKE2b-256 checksum How to use checksums |
c3f34f832666c9ef5a111f7cd8f7396b45eef1b799712505ee216a8aaea9add8
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.13
|