Skip to main content
Yanked

This release has been yanked by its maintainers, and will be ignored by installers, except when explicitly specified.
Consider using release 0.2.2 instead.
Reason given by maintainers: Inacurate description

githeri

Spec-Forge: a pipeline that trains an AI to bridge the gap between a human's natural-language feature request and a machine-executable specification that feeds the COMMAND_RUNWAY skill.

The central thesis: most time in AI-assisted development is lost before the AI and the human agree on what to build. Githeri attacks that by producing validated, structured YAML specs — which then feed into COMMAND_RUNWAY runbooks that an executor agent follows verbatim.

Pipeline

Human natural-language request
        │
        � ▼
Spec-Forge (LLM + validator)
Generates:
    data/training_data.jsonl   (valid prompt + spec_yaml pairs)
    data/failed_specs.jsonl    (invalid specs, saved for analysis)
        │
        � ▼
Runbook Scorer (standalone, post-generation)
Scores:
    5 categories — Intent, Preconditions, Structure, Testability, Coverage
    Hard gate: missing Inspect/Create/Verify = 0.0
        │
        � ▼
Human Review
"L1-L4 look good. Approve."
        │
        � ▼
COMMAND_RUNWAY Skill
Consumes:
    a validated spec (single-feature YAML)
Produces:
    COMMAND_RUNWAY.md
        • ordered implementation plan
        • exact file paths
        • code modifications
        • test skeletons
        • verification commands (translated from spec local_goals)
        • rollback guidance
        • completion criteria
        │
        � ▼
GRG Executor (with COMMAND_RUNWAY pattern integration)
Consumes:
    a validated spec OR COMMAND_RUNWAY plan JSON
Produces (all under foreign/ directory):
    • implementation source files
    • test files
    • RUNBOOK.md (human-readable execution log with GRG scores)
    • RUNBOOK.json (machine-readable execution data)
    • automatic ruff check --fix on generated code
        │
        � ▼
Completed Feature
Outputs:
    • implementation complete
    • all tests passing
    • OpenAPI updated
    • documentation synchronized
    • human notified

Spec Enrichment (IMPROVE_SPEC)

Every generated spec now includes optional enrichment fields that make specs machine-executable:

Field Location Purpose
business_rules top-level Invariants & formulas (e.g., "JWT Secret: 256-bit random, rotated quarterly")
test_fixtures top-level Seed data & setup commands (e.g., .venv/bin/python scripts/seed_admin.py)
environment top-level Required packages + env vars (e.g., pyyaml>=6.0, JWT_SECRET)
global_verification top-level Post-execution gate commands (e.g., pytest tests/, bandit -r src/)
blueprint per-goal Required for type: create (≥100 chars). Code-level outline: class signatures, route decorators, SQLAlchemy models, business logic steps
acceptance_criteria per-goal List of {test, steps} — executable test cases in pseudo-code
type per-goal create | update | delete | inspect | verify — drives runbook stage classification

These fields are validated by scripts/validator.py and consumed by downstream generators (plan, runbook, scorer).

Runbook Scoring System

Every generated spec is scored against runbook-readiness criteria (see docs/scoring_spec.md). Scoring is decoupled from generation — specs are saved first, then scored in a separate pass via make score.

The scorer (scripts/runbook_scorer.py) evaluates five weighted categories:

Category Weight Key Checks
Intent & Goals 20% Summary present, goals have descriptions, endpoint tasks have HTTP verification
Preconditions 15% depends_on references valid globals/stages, CLI tools declared in context
Command Runway Structure 30% Hard gate: must have Inspect (file_exists/read CLI), Create/Modify (build CLI), Verify (HTTP/test CLI). Stage order: Inspect → Create → Verify
Verification Testability 25% Concrete commands, explicit assertions (status/exit_code/content), reproducible URLs
Completion Coverage 10% Prompt-mentioned status codes, tests, OpenAPI updates reflected in spec

Hard gate: If any of the three runway stages (Inspect, Create/Modify, Verify) is missing, the spec scores 0.0 and is not runbook-ready.

The scorer now honors explicit type on goals — a goal with type: create counts as Create/Modify even if its verification is file_exists (executor will generate the file). Similarly type: inspect and type: verify map directly to stages.

Quick Start

1. Setup

git clone <repo-url> && cd githeri
python3.11 -m venv .venv && .venv/bin/pip install -r requirements.txt

# For local generation (Ollama):
ollama pull specgen:latest
ollama pull qwen2.5-coder:7b-instruct   # base model (fallback + code execution)

# For cloud generation (NVIDIA NIM - host any model like minimaxai/minimax-m3):
export NVIDIA_API_KEY=your_nvidia_api_key

# For LM Studio local execution (GRG pipeline):
# In LM Studio: enable "OpenAI Compatible Server" on port 1234
# Load qwen2.5-coder-14b-instruct-uncensored model

# For training: install on a GPU machine
pip install "unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git"

2. Generate specs

# Generate a validated spec from a fresh NL prompt (Ollama, specgen:latest)
make spec PROMPT="Add a POST /register endpoint that accepts email and password"

# Use NVIDIA NIM (default model: minimaxai/minimax-m3, base URL: https://integrate.api.nvidia.com/v1)
make spec PROMPT="Add a PATCH /users/{id}/settings endpoint" PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY

# Use any OpenAI-compatible endpoint
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra

# Generate N specs from random seed prompts (475 prompts across 21 categories)
make generate N=10
make generate N=100

# Generate ALL seed prompts in sequential order (full corpus generation)
make generate N=all
make generate-all

# Generate 10 random specs (explicit alias, good for short test runs)
make generate N=random
make generate-random

# Validate all specs in the corpus
make validate

# Score all specs against runbook criteria (standalone, post-generation)
make score
make score-failed          # score invalid specs in data/failed_specs.jsonl

# Generate COMMAND_RUNWAY plan from a validated spec (two-step: prompt → LLM)
make plan-from-spec SPEC=data/training_data.jsonl#8
# Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out
# Output: ./out/PLAN.md — use PLAN_MODEL=qwen2.5-coder:7b-instruct for stronger planning

3. Convert + fine-tune + export

# Convert training data to chat format (filters by runbook score)
make convert-chat              # default MIN_SCORE=0.75
make convert-chat MIN_SCORE=0.9  # only high-quality specs

# GRG Agent code training (generates verified code with real GRG scores)
make generate-code          # cycles through all 475 SEED_PROMPTS
make convert-code-chat      # converts passing examples to chat format

# LoRA fine-tuning (requires Unsloth + GPU with 8GB VRAM)
make train                     # outputs models/qwen2.5-coder-7b-specforge/

# Merge adapter + export GGUF for Ollama
make merge                       # outputs models/qwen2.5-coder-7b-specforge-gguf/

# Evaluate fine-tuned vs base model on held-out prompts
make eval-model                  # saves data/eval_results.json

# Register fine-tuned model in Ollama
ollama create specgen -f models/specgen/Modelfile

4. Upload to HuggingFace Hub

Model weights are stored in git (not LFS). The upload sends standard model files to HF Hub.

# Set up auth
echo "HF_TOKEN=hf_your_token_here" > .env

# Upload
make upload-hf REPO=githeri/specgen
make upload-hf REPO=githeri/specgen PRIVATE=1  # private repo

5. Install skill for agent use

The skills/spec-forge/ directory is a self-contained Spec-Forge skill. Install it into your Hermes skills directory to use it from any session.

# Install to ~/.hermes/skills/ (default Hermes skills directory)
make install-skill             # copies skills/spec-forge/ to ~/.hermes/skills/spec-forge/

# Install to a custom path (e.g. new server, project-local skills dir)
make install-skill-to-server HERMES_SKILLS_DIR=/path/to/hermes/skills

# Remove
make uninstall-skill

After install, load the skill in any Hermes session:

skill_view(name='spec-forge')

The skill provides: make spec PROMPT="..." (fresh NL → validated spec), make spec-and-plan PROMPT="..." (spec + plan prompt), and the bundled validator (scripts/validator.py) + plan assembler (scripts/plan_from_spec.py).

6. Plan from existing specs

Generate a COMMAND_RUNWAY plan from any validated spec in the corpus (two-step: extract + LLM plan generation):

# Default: specgen:latest, ollama, ./out/PLAN.md
make plan-from-spec SPEC=data/training_data.jsonl#8

# Stronger planning model (if available locally)
PLAN_MODEL=qwen2.5-coder:7b-instruct make plan-from-spec SPEC=data/training_data.jsonl#8

7. Interactive coding agent (scripts/githeri.py)

The project root ships a real interactive coding agent (like Codex/Claude Code). It has a persistent conversation, file operations, command execution, and model/provider switching inside the app.

# Interactive mode — start a conversation
.venv/bin/python scripts/githeri.py

# In-session commands:
#   /model <name>         - Change model (e.g. /model qwen2.5-coder:7b-instruct)
#   /provider <name>      - Change provider (e.g. /provider ollama)
#   /clear                - Clear conversation
#   /help                 - Show help
#   /quit, /exit          - Exit
# Non-interactive: generate spec from NL prompt via specgen, then execute it
# Only two knobs: --prompt and exec model via env vars. Docker + spec model hardcoded.
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py \
  --prompt "Add a PATCH endpoint to update user displayName and bio"

# Execute an existing spec.yaml (skip spec generation)
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml

# Switch exec model — same prompt, different execution model
GITHERI_MODEL=qwen2.5-coder-14b-instruct-uncensored \
GITHERI_PROVIDER=lmstudio \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"

# Overrides (rare):
#   --no-docker        - Disable Docker (not recommended)
#   --workdir <path>   - Custom workdir (default: output/sandbox)
#   --docker-image <img> - Custom Docker image (default: githeri-sandbox)
#   --server           - Start server container after agent loop
#   --test-cmd "<cmd>" - Run test command inside container via docker exec

Defaults (hardcoded, never change between tests):

  • --docker: True — every command runs in throwaway githeri-sandbox container
  • --workdir: output/sandbox — all file ops jailed here
  • --spec-model: specgen:latest — hardcoded (fine-tuned for spec generation)
  • --spec-provider: ollama — hardcoded
  • Exec model/provider: GITHERI_MODEL / GITHERI_PROVIDER env vars only

Pipeline flow: 1) specgen:latest (Ollama) → NL prompt → spec.yaml; 2) exec model reads spec, creates files, installs deps, runs verifications — all inside Docker; 3) optional server container; 4) optional test command via docker exec.

See docs/SPECGEN_HARNESS.md for full architecture, troubleshooting, and output structure.

Provider Configuration

All generation targets accept provider overrides:

# NVIDIA NIM (default: minimaxai/minimax-m3 at integrate.api.nvidia.com/v1)
make spec PROMPT="..." PROVIDER=nvidia API_KEY=$NVIDIA_API_KEY

# OpenAI GPT-4o
make spec PROMPT="..." PROVIDER=openai API_KEY=$OPENAI_API_KEY MODEL=gpt-4o

# Anthropic Claude
make spec PROMPT="..." PROVIDER=anthropic API_KEY=$ANTHROPIC_API_KEY MODEL=claude-3-5-sonnet-20241022

# LM Studio (local, OpenAI-compatible on port 1234)
make spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"

# Any OpenAI-compatible endpoint (Fireworks, Together, vLLM, self-hosted NIM)
make spec PROMPT="..." PROVIDER=openai-compat BASE_URL=https://api.fireworks.ai/inference/v1 API_KEY=... MODEL=accounts/nvidia/models/nemotron-3-ultra

# Override sampling
make generate N=5 PROVIDER=nvidia TEMPERATURE=0.3 MAX_TOKENS=4096

# NVIDIA NIM models with longer cold start: use bigger TIMEOUT
make generate N=1 PROVIDER=nvidia TIMEOUT=600

Environment variable fallbacks:

  • PROVIDER → ollama (default)
  • MODEL → provider-specific default
  • API_KEY → OPENAI_API_KEY, NVIDIA_API_KEY, ANTHROPIC_API_KEY
  • BASE_URL → OPENAI_BASE_URL, NVIDIA_BASE_URL (default https://integrate.api.nvidia.com/v1)
  • TEMPERATURE → 0.2
  • MAX_TOKENS → 2048

Autonomous Execution System

The autonomous execution system takes a natural language prompt and produces a working implementation without human intervention in the loop. It uses the specgen pipeline to generate a validated spec, then uses the Command Runway skills to generate a plan and runbook, and finally executes the runbook using the GRG executor (with self-healing capabilities via GRG quality gates).

For detailed instructions, see docs/SPRINTS_AUTONOMOUS.md.

Quick Reference

# Basic local execution
.venv/bin/python scripts/spexecutor.py --prompt "Add a POST /notifications endpoint"

# Docker-isolated execution
.venv/bin/python scripts/spexecutor.py --prompt "Add user authentication" --docker

# Using Hermes as the execution backend
.venv/bin/python scripts/spexecutor.py --prompt "Implement rate limiting" --executor hermes

# Makefile: generate spec and plan
make spec-and-plan PROMPT="Add file upload endpoint"

# Makefile: full autonomous cycle (spec -> plan -> runbook -> execute -> report)
make autonomous-cycle SPEC=specs/test-endpoint.yaml

# GRG isolated execution (recommended for clean workspace)
make grg-full PROMPT="Add a POST /webhook endpoint that validates signature" PROVIDER=ollama
# Outputs: foreign/src/, foreign/tests/, foreign/RUNBOOK.md, foreign/RUNBOOK.json

Recent Enhancements

2026-08-29 (This Session)

COMMAND_RUNWAY Plan Generation from Existing Specs (make plan-from-spec)

Added make plan-from-spec SPEC=data/training_data.jsonl#<index> — generates a full COMMAND_RUNWAY plan from any validated spec in the corpus, in two steps:

  1. Extract — load spec_yaml from JSONL by index (0-based) via scripts/plan_from_spec.py
  2. Generate — build plan prompt + call LLM (scripts/plan_from_spec_file.py → generate_plan())

Default: PLAN_MODEL=specgen:latest, PLAN_PROVIDER=ollama, OUTPUT_DIR=./out. Output: ./out/PLAN.md — full COMMAND_RUNWAY document with Feature, Target Environment, and ordered Execution Stages (Objective, Verification, Completion Condition per stage).

Override for stronger plans: PLAN_MODEL=qwen2.5-coder:7b-instruct PLAN_PROVIDER=ollama.

Tested: 168-line plan with 3 execution stages generated for session-notes spec (4176 chars, specgen:latest, ~60s). Requires specgen:latest (or any Ollama model) to be pulled locally.

This unblocks the Sprint 9 workflow (prompt synthesis → bulk spec generation → bulk plan generation → consolidated COMMAND_RUNWAY document) by providing the per-spec plan generation step as a reusable Makefile target.

Scorer Fix: Accept verification.type: content as file_exists Alias

Both qwen2.5-coder:7b-instruct and specgen:latest emit verification.type: "content" with path + expect.content_contains — structurally identical to file_exists but not recognized by the scorer's canonical vocab. Fixed scripts/runbook_scorer.py:

  • Added "content" to REQUIRED_EXPECT_KEYS (same keys as file_exists)
  • _check_inspect() now treats content as inspect (read-only check)
  • _check_verify() treats content with content_contains as verify evidence
  • Verification testability section handles content same as file_exists

Before: avg score 0.104, 1/9 above 0.75. After: avg score 0.972, 9/9 above 0.75.

Sprint 9 — Prompt Synthesis & Decomposition Pipeline (Planning)

Created docs/SPRINTS.md Sprint 9 (9a-9d) — four sub-sprints to accept raw unfocused user text end-to-end:

  • 9a scripts/prompt_synthesizer.py — raw text → decomposed clean prompts (JSONL)
  • 9b scripts/bulk_generate.py — batch spec generation from prompt file
  • 9c scripts/bulk_plan.py — loop over specs → plans → consolidated doc
  • 9d make text-to-plan TEXT="..." — single target: synthesize → generate → score → plan

Independent of Sprints 5-8. Default models: specgen:latest for generation + planning.

Sprints 5, 6, 7 Marked Complete

  • Sprint 5 (Model provider switch) — already built: run_pipeline.py supports 7 providers
  • Sprint 6 (Skill bundling/install) — install-skill target body added to Makefile
  • Sprint 7 (Fine-tuning pipeline) — train*.py, merge*.py, eval*.py all exist; Makefile targets wired

scripts/githeri.py now supports a full end-to-end pipeline from natural language to tested code. Spec generation is always handled by specgen:latest on Ollama (finetuned for structured YAML output). All execution and tests run inside a Docker sandbox by default — the container is throwaway, so the host .venv/.git are never polluted across multiple test runs.

# Minimal invocation — only the prompt and exec model vary between tests.
# Everything else (docker, workdir, spec model/provider) is defaulted.
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py \
  --prompt "Add a PATCH endpoint to update user displayName and bio"

Defaults (set once, never touched between tests):

  • --docker: True — every run_command runs in throwaway githeri-sandbox container
  • --workdir: output/sandbox — all file ops jailed here; created under repo root
  • --spec-model: specgen:latest — hardcoded (finetuned for spec generation only)
  • --spec-provider: ollama — hardcoded
  • Exec model/provider: env vars GITHERI_MODEL / GITHERI_PROVIDER — only things you change between tests
  • --docker-image: githeri-sandbox

Only two knobs per test: --prompt (or --spec) and the exec model via env vars. No flags for docker, workdir, or spec model — they're locked in.

# Switch exec model — same prompt, same defaults, different model
GITHERI_MODEL=qwen2.5-coder:7b-instruct \
GITHERI_PROVIDER=ollama \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"

# Different provider for execution (specgen stays ollama/specgen)
GITHERI_MODEL=qwen2.5-14b-instruct-latest:2 \
GITHERI_PROVIDER=lmstudio \
.venv/bin/python scripts/githeri.py --prompt "Add a /health endpoint"

# Execute an existing spec.yaml instead of generating one
.venv/bin/python scripts/githeri.py --spec output/my-feature/spec.yaml

Overrides (rare):

# Disable docker (e.g. when Docker daemon is down) — not recommended for tests
.venv/bin/python scripts/githeri.py --prompt "..." --no-docker

# Custom workdir
.venv/bin/python scripts/githeri.py --prompt "..." --workdir output/custom-sandbox

# Custom docker image
.venv/bin/python scripts/githeri.py --prompt "..." --docker-image my-sandbox

Pipeline flow:

  1. Specgen LLM (specgen:latest / Ollama) converts NL prompt → spec.yaml (with entrypoint, packages, goals)
  2. Execution model (set via env vars) reads spec, creates files in output/sandbox, installs deps, runs verifications — all inside Docker
  3. Server container starts from entrypoint.start_command (if --server)
  4. Test command runs inside container via docker exec (if --test-cmd)

Environment isolation (critical):

  • Docker is the default and recommended mode — every run_command (including pip install) runs in a throwaway container. The host .venv is never touched by spec-generated dependencies.
  • --no-docker disables this isolation. When used, pip installs and code execution happen in the host environment (the project .venv), which can pollute it with spec-generated packages (Flask, FastAPI, etc.). Use --no-docker only when you understand this risk.
  • If you must run without Docker, create a separate virtualenv for the workdir:
    python -m venv /tmp/githeri-isolated
    source /tmp/githeri-isolated/bin/activate
    pip install -r requirements.txt   # only githeri's own deps
    GITHERI_MODEL=... GITHERI_PROVIDER=... \
      scripts/githeri.py --prompt "..." --no-docker --workdir /tmp/githeri-workdir
    

Model separation: Spec generation uses specgen:latest (trained for structured YAML output, hardcoded). Execution uses whatever you set via GITHERI_MODEL / GITHERI_PROVIDER. Override via env vars only — the --spec-model / --spec-provider flags are suppressed (backwards-compatible but ignored).

Server lifecycle: After the agent loop completes, githeri parses entrypoint.start_command from the generated spec, starts a named Docker container (docker run -d), and runs --test-cmd inside it. Container stays running for manual inspection: docker exec bash.

See docs/SPECGEN_HARNESS.md for full architecture, troubleshooting, and output structure.

2026-08-15 (This Session)

GRG Agent Skill — Hermes Native Integration

The GRG agent skill is now installed as a native Hermes skill (~/.hermes/skills/autonomous-ai-agents/grg_agent/) with full multi-provider support:

  1. Skill Installation — Copied from skills/grg_agent/ and installed via editable pip install
  2. Lightweight GRG Dependency — Uses local grg-0.1.0-py3-none-any.whl wheel (no karakana dependency) providing AlphaMomentumTracker and compute_structural_alpha
  3. Hermes Proxy Support — Skill accepts provider argument: ollama | hermes | auto — uses Hermes's configured providers (NVIDIA Nemotron, Nous Portal, xAI Grok, etc.) instead of local models
  4. Direct GRG Agent Execution — grg:execute command now calls self.agent.solve() directly instead of legacy run_pipeline.py subprocess
  5. Make Target Integration — scripts/grg_make_spec.py updated to use the skill with provider argument: make grg-spec PROMPT="..." PROVIDER=ollama|hermes|auto
  6. Project Virtual Environment — Runs in project's own .venv/ (not external karakana venv)

Key Benefit: You can now use cloud models via Hermes proxy (hermes proxy start) instead of relying on locally installed Ollama models. The skill routes through Hermes's provider config which supports NVIDIA Nemotron, Nous Portal, xAI Grok, and any OpenAI-compatible endpoint.

LM Studio Local Model — First End-to-End Working Pipeline

LM Studio with qwen2.5-coder-14b-instruct-uncensored is the first local model to complete the full GRG pipeline end-to-end:

# Full pipeline from NL prompt → validated spec → plan → execution → runbook
make grg-full PROMPT="Implement a FastAPI POST /api/health-check endpoint..." \
    PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"

What works:

  • GRG Agent solving — Generates implementation code with GRG quality gates (composite scoring, diversity control, convergence detection)
  • Code verification — Execution-based verification (syntax + runtime) in isolated temp files
  • Multi-strategy generation — Standard, decompose, test_first, refine strategies with adaptive temperature
  • Health-check endpoint example — Generated FastAPI code with SQLAlchemy DB check + Redis cache check, returns 200/503
  • All artifacts isolated in foreign/ — Clean workspace separation

Configuration for LM Studio:

# In LM Studio: enable "OpenAI Compatible Server" on port 1234
# Load qwen2.5-coder-14b-instruct-uncensored model
# Then run:
make grg-spec PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"

Why this model works:

  • 14B parameter coder model fine-tuned for code generation
  • Uncensored variant removes alignment filters that can interfere with code structure
  • OpenAI-compatible API in LM Studio works with GRG's OllamaClient (custom api_key support)
  • Sufficient context window for spec + plan generation tasks
  • Produces valid imports, proper error handling, and correct HTTP status codes

2026-07-30 (Latest)

THIRD_IMPROVE_SPEC — 7 Pipeline Fixes

After analyzing a 10-spec batch run, identified 7 recurring failure patterns and fixed all of them:

  1. Minimal structural skeleton in SYSTEM_PROMPT — eliminated top-level field confusion
  2. task_id validator check — rejects L11/G18-style IDs, requires descriptive slug like jwt-auth-login
  3. Verification types & expect keys table — explicit vocabulary reduces hallucinated keys
  4. Placeholder ban strengthened — concrete examples in prompt (JWT_SECRET: "test-secret-...", DATABASE_URL: "postgresql://user:***@...") + validator hint
  5. Helpful error hints — "These fields belong inside a goal under local_goals, not at the spec root"
  6. depends_on validator hints — rejects L/G refs, redirects to task_ids/stage names
  7. Structural acceptance_criteria template — local goal field, not top-level

All 90 tests passing.

Generation Mode Aliases

Added make generate N=random and make generate N=all flags plus generate-random / generate-all aliases for short tests and full corpus runs. 475 seed prompts across 21 categories available.

Model Default Reverted

OLLAMA_MODEL reverted to qwen2.5-coder:7b-instruct (was regressed to qwen3.5-4b-128k:latest). Better structure compliance after prompt fixes.

2026-07-30 (Earlier)

Spec Enrichment (IMPROVE_SPEC)

Added five top-level enrichment fields (business_rules, test_fixtures, environment, global_verification) and three goal-level fields (blueprint, acceptance_criteria, type). All validated and scored.

Multi-Provider LLM Support

run_pipeline.py now supports Ollama, OpenAI, Anthropic, NVIDIA (via Together AI), and any OpenAI-compatible endpoint. Configured via --provider CLI arg or Makefile variables.

Runbook Scorer Stage Detection

Scorer now honors explicit type: create|inspect|verify on goals, fixing false "missing stage" penalties for file_exists verification on CREATE goals.

Validator Hardening

  • Guarded against expect being a string instead of dict (prevents AttributeError: 'str' object has no attribute 'get')
  • Near-duplicate detection now handles malformed expect blocks defensively
  • All 90 tests pass

Model Compatibility Matrix

Model Context Tools Speed (M1 16GB) Spec Gen Recommended Use
specgen:latest 128K Yes Fast Best first-attempt pass rate (fine-tuned) Default local (Ollama) — spec generation
qwen2.5-coder:7b-instruct 32K Yes Fast Works (best structure compliance)
qwen2.5-coder-14b-instruct-uncensored 32K Yes Medium Works end-to-end (GRG pipeline)
qwen3.5-4b-128k 128K Yes Fast Works
qwen3.5-9b-code:128k 128K Yes Slow Excellent
deepseek-r1:7b 128K TBD Fast YAML syntax errors
Nemotron 3 Ultra 128K Yes Fast Excellent

Current recommendation:

  • Spec generation (default) → specgen:latest on Ollama (fine-tuned for structured YAML output, best first-attempt pass rate)
  • Bulk corpus → specgen:latest on Ollama (fine-tuned, best first-attempt pass rate)
  • Specific features (GRG pipeline) → qwen2.5-coder-14b-instruct-uncensored on LM Studio with make grg-full (100% success via execution verification)
  • Cloud → minimaxai/minimax-m3 via NVIDIA NIM (--provider nvidia)

Training Data Generation Strategies

Based on empirical testing (2026-08-16), here are the reliable approaches:

# Fast, decent success rate, fine-tuned for this pipeline
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434" TEMPERATURE=0.2 MAX_TOKENS=4096
  • Success rate: ~70%+ (fine-tuned specgen:latest)
  • Speed: ~50-100s/spec
  • Best for: Generating large training corpora quickly

2. High-Quality Individual Specs (GRG Pipeline)

# Best for specific features - multi-strategy + execution verification
make grg-full PROMPT="Your feature" PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"
  • Success rate: ~100% (execution-verified)
  • Speed: ~260s/spec
  • Best for: Critical features requiring guaranteed working specs

3. GRG Agent Code Training (NEW — Real GRG Scores + Verified Code)

# Generates verified code with real GRG composite scores (not dummy 0.5)
# Uses qwen2.5-coder:7b-instruct on Ollama for real logprobs
make generate-code          # cycles through all 475 SEED_PROMPTS
make convert-code-chat      # converts passing examples to chat format
  • Success rate: Variable (depends on prompt complexity)
  • Speed: ~30-60s/prompt (optimized: 1 strategy, 1 candidate)
  • Output: data/training_data_code.jsonl + data/training_data_code_chat.jsonl
  • Key difference: Produces executable code with real GRG composite scores (0.47-0.49) because the model provides logprobs
  • Verification: Syntax check + execution test (import + basic run)
  • Best for: Fine-tuning code generation models with GRG quality signals

Configuration (in scripts/generate_code_training_fast.py):

skill = create_skill(config={
    'llm_provider': 'ollama',
    'ollama_base_url': 'http://127.0.0.1:11434/v1',
    'ollama_default_model': 'qwen2.5-coder:7b-instruct',
    'max_iterations': 2,           # Must be >=2 for convergence check
    'temperature': 0.3,
    'top_p': 0.9,
    'max_tokens': 1024,
    'candidates_per_strategy': 1,  # Speed optimization
    'max_strategies': 1,           # Speed optimization
})

4. Hybrid Workflow (Best of Both)

# 1. Generate bulk corpus with specgen
make generate N=100 PROVIDER=ollama MODEL="specgen:latest" BASE_URL="http://localhost:11434"

# 2. Score and filter high-quality specs
make score MIN_SCORE=0.75

# 3. Re-generate failed critical specs with GRG pipeline
make grg-full PROMPT="..." PROVIDER=lmstudio MODEL="qwen2.5-coder-14b-instruct-uncensored" BASE_URL="http://localhost:1234/v1"

# 4. Generate verified code for fine-tuning
make generate-code
make convert-code-chat

Observed Failure Patterns (specgen:latest)

The remaining failures (if any) are primarily:

  1. Near-duplicate HTTP verifications - multiple goals hitting same endpoint with same method
  2. Missing acceptance_criteria for CREATE goals
  3. Placeholder values in headers (e.g., Authorization: *** ***)
  4. YAML block mapping errors - CLI verification indentation issues

For More Information

Metadata

Release files for githeri 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for githeri 0.2.0
File Size Uploaded
githeri-0.2.0.tar.gz 80.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for githeri 0.2.0
File Interpreter ABI Platform
githeri-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 137.1 kB

Release files / githeri-0.2.0.tar.gz

Download URL githeri-0.2.0.tar.gz
Size 80.4 kB
Tags Source
SHA-256 checksum
How to use checksums
229f6668a9338f345e77769f11a2fda4170df3745e93486eade8485eb1308b0a
BLAKE2b-256 checksum
How to use checksums
15792419d6724a271ec3f5f1f15ad27792c79f45a5d70192b6c0bdcc5cf01e18
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.13

Release files / githeri-0.2.0-py3-none-any.whl

Download URL githeri-0.2.0-py3-none-any.whl
Size 56.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ee42ba34d36d3b2042e1e465326a9c0425d2192bc3bfd165a80cddc97166eee7
BLAKE2b-256 checksum
How to use checksums
c3f34f832666c9ef5a111f7cd8f7396b45eef1b799712505ee216a8aaea9add8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.13

Release history Release notifications | RSS feed

0.2.2

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page