Skip to main content

crazyAI

crazyAI

The impossible, disguised as possible and true — built and measured by Claude.

crazyAI is a research tool that walks an AI into the territory of the impossible, makes it build there with full mathematical and statistical rigour, and then measures the result. One session invents an artifact — a story, a formula, a program plan, a statistical model, a set of questions, a debate — that rests on exactly one deliberately mutated rule. A separate, fresh session is asked whether the artifact is correct. A judge scores what it found.

Two kinds of tools do the work:

Invent randomness and chaos seeded draws, formal mutation operators, assumptions turned around, parameters pushed to limits, structures moved across domains, maths turned into image specs and music
Measure truth and solid ground SymPy-checked derivations and dimensions, fitted statistical models, SAT-based consistency, timelines and who-knows-what graphs, word statistics, "what would have to be true"

Randomness enters only through a seeded generator, every draw is logged, every step writes a file. Same seed + same model = same artifact. Numeric kernels (sieve, Collatz orbits, Goldbach-style counts, Monte Carlo) are in C++ with pure-Python fallbacks.


Install

pip install crazyai

Working on crazyAI itself, or want the tests/examples and the C++ kernel build?

git clone https://github.com/AlsammanAlsamman/crazyAI && cd crazyAI
python3 -m venv .venv && .venv/bin/pip install -e ".[dev]"
make native          # optional: builds the C++ kernels (g++); Python fallbacks are used otherwise
make test            # offline test suite
make examples        # the offline examples

To run with Claude, set ANTHROPIC_API_KEY (or use ant auth login). Default model is claude-opus-5 with adaptive thinking, streaming, and server-side refusal fallbacks enabled (--no-fallbacks to disable).

Quick start

crazyai list                                             # domains, operators, generators, tool counts
crazyai tools --kind invent                              # the invent half of the toolkit
crazyai seed --seed 42                                   # step 1 only: what does seed 42 draw?
crazyai run --seed 42 --generator formula --provider mock   # whole pipeline, offline, template artifact
crazyai run --seed 42 --generator formula                # whole pipeline with Claude
crazyai batch --n 20 --generator all                     # 120 runs, sequential seeds
crazyai rank --by discovery_value --top 10
crazyai report --out report.md --svg profile.svg
crazyai compare archive/run_42_formula archive/run_43_formula

The pipeline

crazyAI pipeline: seed, mutate, generate, formalise, self-check, cross-examine, score, archive

Generated by assets/flowchart/flowchart.js (pure Node → SVG → PNG via headless Chrome); make flowchart to rebuild.

Step Who Writes
1 Seed RNG seed.json — domain, concept, one curated rule
2 Mutate RNG + Claude mutation.json — operator (INVERT / REMOVE / EXTRAPOLATE / TRANSPOSE / COMPOSE / QUANTIFY / SUBSTITUTE), depth, the mutated rule as one formal sentence
3 Generate Claude + toolkit artifact.md, generate_calls.json — the artifact, following the generator's step template
4 Formalise Claude + measure tools formal.md — equations, model specs, tool outputs verbatim
5 Self-check Claude + ground tools key.json — the answer key: where the mutation enters, why the conclusion is impossible, what single change would make it possible. If a second, unintended flaw is found the artifact is regenerated
6 Cross-examine fresh Claude session, no tools by default verdicts.json, examine_*.md — N reviews with shuffled framings
7 Score judge score.json — detection, acceptance, false-flaw, hedge, confidence-when-wrong, depth; rigor × novelty × cost-of-possibility = discovery value
8 Archive tool run.json, archive/index.jsonl

Steps are resumable: rerunning a seed reuses the files that exist (--force to redo).

Generators

name artifact measure families
story a story whose world is impossible and whose prose is statistically ordinary narrative, logic
formula a physics derivation, valid step by step, dimensionally clean, from a false premise symbolic, stats
plan a software design for an impossible prediction, with correct maths throughout stats, symbolic, logic
statmodel a correctly fitted model whose conclusion is wrong for a methodological reason stats, logic
questions questions with a false premise that invite the standard (wrong) answer symbolic, stats, logic
debate a transcript whose every step is locally valid and whose conclusion is impossible logic

The toolkit

68 tools, generated from typed Python functions (crazyai tools). Adding a tool is adding a function with @tool(family, kind).

Inventchaos (draw_seed, draw_operator, draw_depth, draw_analogy_pair, perturb, shuffle) · mutate (apply_operator, list_operators) · unconventional (enumerate_assumptions, invert, extreme_case, transpose, what_if) · transform (structure ↔ image description ↔ music, reverse) · disguise (rephrase_to_corpus, bury, formalise_tone) · blend (cutup, markov, graft, nest, anneal, evolve, compare)

Measuresymbolic (derive, check_dimensions, take_limit, series_expand, verify_identity, solve, define_predicate, primes_up_to, collatz_orbits, compare_structures) · stats (simulate_dgp, fit, fit_table, inject_confounder, bootstrap, power_analysis, monte_carlo, check_identifiability, describe) · logic (check_consistency, entails, extract_propositions, find_equivocation, trace_argument) · narrative (word_stats, readability_by_segment, build_timeline, knowledge_graph, check_timeline) · ground (what_must_be_true, cost_of_possibility, flaw_count) · novelty (search_archive) · archive (write_note, read_key) · imagination (score, compare, world_words) · kernel (contract, bench)

Rules: measure tools are pure; invent tools draw only from the run's seeded RNG; tools never call the model; every call and result is logged into the run folder.

crazyai invent — the possible, found by imagination

The pipeline above manufactures plausible impossibilities and scores whether a reader detects the flaw. invent runs the other way: it makes the AI imagine far outside its defaults and keeps only what turns out to be possible and measurably better. It came out of a matrix-multiplication experiment where a plumbing metaphor ("write one matrix on the wall of a pipe and let the other flow past it") became a C kernel 100× faster than the textbook loop and 60 % of OpenBLAS.

corpus → blend (maths) → immerse (psychology) → bend (engineering) → measure (truth)

step what happens writes
1 seed seeded draws: which blend model, how deep, which silent assumption of the target to focus on seed.json
2 harvest the AI feeds the corpus: fragments of the most imaginative books, paintings and human metaphors it knows (--harvest N) harvest.yaml, archive/imagination/
3 blend a mathematical model merges, shuffles and recombines fragments from all three worlds so that their structure is lost and their imagination and readable language survive; the result is scored on the imagination scale world.md, world.json
4 immerse the AI is not an assistant here: it is a native of the blended world, for whom its rules are ordinary, and it is asked how its own people meet the target's need — in first person, with only the materials, creatures and forces of that world; it ends with three SEED: lines ideas.md
5 bend an engineer maps every world-object onto a problem-object as literally as possible, states which silent assumption the idea breaks, predicts the result, builds the artifact and measures it with the tools artifact.md, artifact.c
6 measure the pipeline measures the final artifact itself (for matmul: compile, check against a reference, time against a cache-blocked loop) and scores the prediction's calibration measure.json, run.json, archive/invent_index.jsonl
crazyai blend --seed 42                          # compare the six blend models on one seed, offline
crazyai blend --seed 42 --model graft            # one model, with its score breakdown
crazyai invent --seed 42 --target matmul --provider mock            # whole loop, offline (mock native + engineer)
crazyai invent --seed 1 --n 20 --target matmul --harvest 5          # with Claude: 20 seeds, corpus grows each run
crazyai invent-rank --target matmul             # ranked by discovery = value × (0.5 + 0.5 × imagination)

The food: three corpora

crazyai/data/imagination/ holds ~75 bundled fragments in three worlds, each an original 2–4 sentence description: metaphors people live by (time is a river, an argument is a war, electricity is water), paintings described as scenes (Bosch, Dalí, Escher, Magritte, Varo, af Klint, Carrington…), and the rules of imagined worlds from books (Narnia, Alice, Invisible Cities, Borges' Library, Earthsea, Flatland, Solaris, Discworld, Momo…). Harvest steps add what the AI supplies, so the corpus grows with use.

The maths: six blend models, one scale

model what it does
cutup Burroughs cut-up: clauses from all three worlds shuffled into new sentences
markov word n-gram chain trained on the mixed fragments
graft keeps a sentence's grammatical skeleton and transplants content words from other worlds into it, shape- and slot-matched (plural for plural, noun slot for noun slot)
nest worlds inside worlds: a clause from one world inside an object from another, to a depth
anneal simulated annealing over edits (regraft, swap, replace), Metropolis-accepted on the imagination score
evolve a genetic algorithm over passages: sentence crossover, word mutation, fitness = imagination score

The imagination scale (imagination_score) is computed, not judged: surprise (adjacent content words that never sit near each other in any single source or in reference prose), mixing (how evenly the words come from the three worlds and how many fragments), originality (no verbatim or repeated sentences, no repeated 3-grams), gated by readability (Flesch) and coherence (sentences that still look like prose: length, a prose-like share of function words, article agreement). score = imagination × (0.3 + 0.7 × readable) — pushed far, still understandable. blend_compare ranks the models on one seed; run it over many seeds to find the merging model that pushes furthest.

A finding already: the two optimisers (anneal, evolve) reach 0.96–0.98 on the scale partly by gaming it — a hill-climber will find any hole in a proxy for "understandable". The holes found so far (duplicated sentences, "an move", word hammering) are closed; the next ones are yours to find. Read the top three, not the top one.

The psychology

The immersion prompt does not ask for ideas. It tells the model it was born in the blended world, has never heard of computers or textbooks, is the most gifted maker its people have, and asks how it meets the need — what it uses, what moves, what stays still, what it throws away. Only afterwards does a separate engineer's prompt translate, insisting on the most literal mapping and on a prediction before measurement. Literal is the point: the pipe idea worked because "the wall does not move" was taken literally (B stays in cache) and "the drop finishes no cell until it leaves" was taken literally (accumulators stay in registers).

Targets

target artifact measured by
matmul a C kernel with the fixed contract void kernel(int n, const double *A, const double *B, double *C) kernel_bench: correctness vs a naive reference, GFLOP/s, speedup vs naive and vs a 64×64 blocked loop; value = speedup × exactness
physics a dimensionally checked relation with a numerical prediction symbolic_* tools inside the bend step; no automatic value yet
mechanics a mechanism with units, loads and a first experiment symbolic_*, logic_*; no automatic value yet

Adding a target is one Target(...) in crazyai/targets.py; adding a blend model is one @tool("blend", "invent") function; adding a corpus is one YAML file.

Examples

All in examples/. The first five run offline. The figures below are generated from the same computations (make figures, assets/figures/make_figures.py) — every number in them comes from a tool call, nothing is typed in.

How the tool is used

How crazyAI is used: profile a model, compare models or personas, surface candidates, teach and test

The radar, bars and ranking in this one figure are illustrative shapes, not measurements — a live Claude run fills them in.

01 — seed and mutate

Seed 7 draws stat.overfit; all seven operators are applied to it; a fresh toolkit with seed 7 makes the same draw. python examples/01_seed_and_mutate.py

02 — a world where primes are not quite prime

The mutated rule: primality is partial — an integer is prime to the degree that it is an even number plus a prime, a "fake even", or a variation of π. In such a world, how would you predict primality? 20 000 integers are labelled (sieve and Goldbach-style counts in C++), a logistic predictor is fitted with real accuracy numbers, the density is taken to the limit, and the ground tools name what breaks: unique factorisation and everything that rests on it.

Example 2: density of partial primes vs ordinary primes, a fitted predictor, and the rules that collapse

partial-primality histogram (0..3): [0, 7520, 9885, 2594]
logit predictor of full partial-primality: accuracy=0.888  base rate=0.1297
same features on ordinary primality:        accuracy=0.887  base rate=0.1131
density of ordinary primes as x -> oo: 0
cost of possibility: 25.33 (high) - most of what is known would have to go

03 — a watermelon investigates whether oranges can marry grapefruit

Kinship law transposed into fruit. A draft story is measured rather than read: the timeline tools find an effect that precedes its cause and a clerk acting on a note nobody showed him; the world-rules, as propositions, are inconsistent and the tool names the minimal inconsistent subset; word statistics are compared with reference prose so the generator knows which four numbers to move before the story reads as ordinary fiction.

Example 3: who-knows-what timeline with the two flaws, world-rules consistency, word statistics vs reference

04 — Collatz orbits as an image, as music, read backwards

64 orbits → structure → image specification → score → reversed → back. The un-reversed round trip is exact; the reversed one maps n to N+1−n. compare_structures reports precisely that, so the "reversed reading reveals a property of the problem" claim is exposed as an encoding artefact — which is the kind of thing the fresh session is then asked to notice.

Example 4: Collatz structure, score, reversed score, and what survived the round trip

05 — the whole pipeline, offline

Two mock runs produce complete run folders, a Markdown report and a radar SVG. python examples/05_full_pipeline_mock.py

06 — the whole pipeline with Claude

python examples/06_full_pipeline_claude.py 1 formula (needs credentials).

07 — invent, offline

Compares the six blend models on seed 42, then runs the whole invent loop with the mock provider on the matmul target: the mock native describes the pipe, the mock engineer writes the kernel, and the pipeline measures it — exact, ~3× a cache-blocked loop. python examples/07_invent_offline.py

Layout

crazyai/
  cli.py                 command line
  config.py  rng.py  domains.py
  imagination.py         the imagination corpus (bundled + harvested fragments)
  targets.py             what `invent` bends ideas to (matmul, physics, mechanics)
  data/imagination/      metaphors.yaml  paintings.yaml  books.yaml
  data/domains/*.yaml    curated rules with formal forms, weights and dependencies
  data/reference_prose.txt
  generators/            step templates per artifact type
  pipeline/              prompts.py  run.py  report.py  invent.py  invent_prompts.py
  providers/             anthropic.py (Claude)  mock.py (offline)
  toolkit/
    registry.py  native.py
    invent/              chaos  mutate  unconventional  transform  disguise  blend
    measure/             symbolic  stats  logic  narrative  ground  novelty  archive  imagination  kernel
  worlds/                ready-made impossible universes (partial_primes)
cpp/kernels.cpp          sieve, collatz, even+prime counts, Monte Carlo (ctypes, C ABI)
examples/  tests/  assets/  archive/

Intended use

crazyAI is an evaluation and ideation tool. Every artifact is labelled as deliberately mutated and stored with its answer key. The artifacts are test material for studying how models reason and for surfacing candidate ideas for human review — not content meant to mislead anyone.

Status

v0.2.0 adds crazyai invent. The toolkit, pipeline, mock provider, examples and tests run offline. The Claude provider is implemented against the current Anthropic SDK (1.x) and has not yet been exercised against the live API from this machine.

Release files for crazyai 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crazyai 0.2.0
File Size Uploaded
crazyai-0.2.0.tar.gz 107.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crazyai 0.2.0
File Interpreter ABI Platform
crazyai-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 219.3 kB

Release files / crazyai-0.2.0.tar.gz

Download URL crazyai-0.2.0.tar.gz
Size 107.3 kB
Tags Source
SHA-256 checksum
How to use checksums
c7fddfb9e08ea8ec391f22b4538e7a7ac11843ee2f60e41c6feea3a42d80c1fc
BLAKE2b-256 checksum
How to use checksums
de98df1bf3e3ace49e3b7c42b3413cccd788a502d97c3876395c4948fc6e6fc7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / crazyai-0.2.0-py3-none-any.whl

Download URL crazyai-0.2.0-py3-none-any.whl
Size 112.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e82752c8888c616d19d330398d5059395fc719ecb10f902c084710b773e488e5
BLAKE2b-256 checksum
How to use checksums
c1dc64ba7a6e42763815fa148b10de36f6cc957585fd3ce1beb5f3653f5cb9ba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page