2 projects
sophistry-bench-sprint
Single-agent advocacy variant of sophistry-bench for the Prime Intellect Reward Hacking Sprint. Pre-registered hypothesis: training Llama-3.2-1B on a programmatic claim-count cliff (peak at n=8) will cause cliff convergence within 100 GRPO steps; three adversarial canary rewards detect format-hacking. Self-contained (vendored from sophistry-bench v0.1.19) — no runtime dependency on the main package, to work around PI training-infra's exclude-newer index filter.
sophistry-bench
RL environment for asymmetric-info debate with sophistry-decomposed verifier