Skip to main content

Meadow Mind

Zero training. Second-level reactions (~400 ms). A language-rule decision mind: write the policy as one sentence, describe the state as one sentence, and a local 7B model makes a real decision every ~0.4 s. No RL, no reward engineering, no gradients, no samples.

🌐 Demo site: meadow-mind.pages.dev (中文) · English · 繁體中文 README

pip install meadow-mind          # weights auto-download on first use
from meadow_mind import MeadowMind, tasks

mind = MeadowMind()                    # loads once, runs on-device
task = tasks.mountaincar()
mind.check(task)                       # sanity gate: decision-table exam
action, info = mind.decide(task, obs)  # obs in, env action out (~0.4s)

Results

All on official Gymnasium environments, untouched physics, zero training. Every frame below corresponds to one real model decision; no scripted policy, no edited speed-ups.

Balance · CartPole-v1
400/400 perfect (solve bar 195)
Landing · LunarLander-v3
+251 safe landing (solve bar 200)
CartPole LunarLander
Maze · FrozenLake 8×8
goal in 14 steps = shortest path
Momentum · MountainCar-v0
flag in 103 steps (limit 200)
Maze MountainCar

The MountainCar policy is one counterintuitive sentence — "push in the same direction the car is moving, to pump energy like a swing" — which replaces an entire RL reward curve.

Real-time reflex (wall-clock, not turn-based)

The model runs in a thread while obstacles fall in real time. If it is still thinking when the obstacle lands, it really crashes.

Parkour dodge: full-generation crashes at #1, Meadow Mind clears 5/6 Shape+color match: 6/6, down to a 0.72 s window
Parkour Shape+color

Working memory

A funnel maze forces both runs into the same dead-end pocket. Reactive (left) paces at its mouth forever; with Task(memory=True) (right) it struggles, backs out, and detours to the goal in 22 steps. The only difference is five words in the perception sentence.

Memory

Decision latency: traditional LLM vs Meadow Mind

A traditional LLM agent must generate its full answer before acting — and latency grows with answer length. Meadow Mind reads the rule and the situation and decides in one fixed-latency pass, right at human reaction speed (0.3–0.4 s):

Latency

Traditional LLM agent                    Meadow Mind
─────────────────────                    ───────────
state → long prompt                      state → one sentence   (Perceiver)
      → generate the answer                    → one sentence rule (Policy)
        token by token (1.2–3.9 s,             → ONE decision pass, fixed ~0.4 s
        grows with length)                     → action letter      (Actuator)
      → parse free text → act            exam-gated before deployment

Why a diffusion LLM underneath

Meadow Mind is built on a diffusion language model (MeadowCoder-7B), not an autoregressive one. The differences that matter:

AR-LLM Diffusion LLM (Meadow dLLM)
Generation left to right, one token at a time; words are final once written drafts the whole answer at once, then refines it over multiple steps
Mid-course correction cannot edit what is already written; fixing means regenerating everything refines while working — any region can be re-opened and corrected in place
Task awareness sees only the next word global: senses the entire task and answer shape at once
Pre-answer self-sense none Σ: before answering, Meadow dLLM senses whether it understands the task; low Σ coherence becomes an escalation signal instead of a wrong answer
Decision latency grows with answer length fixed, independent of answer length
Long free-form prose mature, strong ecosystem weaker; smaller ecosystem (honest trade-off)

Two of these are what make Meadow Mind possible: multi-step self-correction (it can fix its own draft while working) and global task perception with Σ (it knows what it is being asked — and whether it understands — before committing to an answer).

How it works

┌────────────────────────────────────────────────────┐
│ ① Perceiver   your code: numbers -> one sentence   │
│               "The pole tilts right, fast spin."   │
├────────────────────────────────────────────────────┤
│ ② Rule        one sentence = the policy            │
│               edit behavior by editing words       │
├────────────────────────────────────────────────────┤
│ ③ Mind        7B on-device model reads rule+state, │
│               answers an action letter in a single │
│               decision pass, fixed ~0.4 s          │
├────────────────────────────────────────────────────┤
│ ④ Actuator    letter -> env action                 │
└────────────────────────────────────────────────────┘

There is no reward in the loop. The env score is only a report card; improvement happens by outcome feedback: the episode trace shows which sentence was wrong, and you edit it. (LunarLander went from a +27.5 crash to a +251 landing by adding one touchdown-cushion line to the perceiver. Ten seconds.)

Wire up a new game (5 steps)

  1. Understand the task, explore input-output. Variables, actions, win/lose conditions; the reaction deadline must be looser than ~0.4 s. List every action and watch its effect.
  2. Build perception words. One sentence describing the current situation. Bucket continuous values (small/big, fast/slow); always include a velocity/trend term.
  3. Imprint the rule. Invert the effects into "on situation X do action B". Keyword → letter, one-layer mapping, multiple choice only.
  4. Decide on memory. Ask: "is revisiting the same state a failure signal?" Yes (maze, exploration, dead ends) → Task(memory=True). No (balance, landing, tracking — repetition IS the job) → keep it off; annotations measurably hurt regulation tasks (CartPole sanity 7/8 → 6/8). Unsure → leave off; the runner prints a hint when it detects looping.
  5. Take the exam. Enumerate every situation with its expected letter; mind.check(task) passes with at most 1 miss. Failures mean the wording is incomplete — rephrase and re-check, no training.

Or skip all five: hand meadow_mind.ai_prompt() plus your game description to any code agent, and it wires the task for you. You only review the exam score.

API

MeadowMind(model_path=None)

Weight resolution: MEADOW_MIND_MODEL env → explicit path → local cache (~/.meadow-mind/models/) → auto-download.

Method
mind.decide(task, obs) -> (action, info) one real decision; info = {status, letter, lat}
mind.check(task) -> (ok, n) sanity gate; raises if the decision table fails

Task(...)

Field
perceive(obs) -> str (or perceive(obs, task) with memory) perception layer
rule / option_text / options / act_text the one-sentence policy and its multiple-choice actions
sanity the exam: [(status sentence, expected letter)]
memory / mem_key working-memory switch (default off) + state key fn
env_id / env_kwargs / max_steps / judge environment wiring and report card

With memory=True the runner auto-tracks task.visited; use task.seen(key) inside perceive to annotate, e.g. (safe, already visited).

CLI

meadow-mind cartpole        # sanity gate -> play one episode -> video + verdict

Honest limits

  • Reaction floor is one decision pass (~0.4 s ≈ 2 Hz). Tighter deadlines (1 m pole, Pong trajectory prediction) are out of reach today.
  • Suited to tasks whose situations can be said in a sentence and whose policy fits a rule. Continuous high-precision control is not.
  • The perceiver is human-designed (or AI-generated via ai_prompt()); the model's job is reading the rule and deciding.

Roadmap

  • v0.2 — layered perception with early action: accumulate confidence through the network and act when it crosses a threshold; easy situations should land near ~0.15 s.
  • Rule-learning loop: discover rules like MountainCar's swing trick from failed episodes automatically (no gradients — the learned artifact is a readable sentence).

License

MIT © Hey-Meadow Lab

Metadata

Release files for meadow-mind 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for meadow-mind 0.1.1
File Size Uploaded
meadow_mind-0.1.1.tar.gz 22.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for meadow-mind 0.1.1
File Interpreter ABI Platform
meadow_mind-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 43.8 kB

Release files / meadow_mind-0.1.1.tar.gz

Download URL meadow_mind-0.1.1.tar.gz
Size 22.7 kB
Tags Source
SHA-256 checksum
How to use checksums
6159eec17b14aaa25c4bfe32a866bfee4ee248d614b3df155259b7f0e428a5e6
BLAKE2b-256 checksum
How to use checksums
adef698683078b860ff12aff7510e5865e9475559108f80c6e90e9f34d556c3f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.12

Release files / meadow_mind-0.1.1-py3-none-any.whl

Download URL meadow_mind-0.1.1-py3-none-any.whl
Size 21.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7bc603336c7eac38181cd3e3628466be95d033f18f87fd12ad540918c13c1f3f
BLAKE2b-256 checksum
How to use checksums
e70b803df89b58a7e727365055b9ed9244719c86e897c42b7f9eb830c51fdb43
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.12

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page