Skip to main content

RawIntent 4.0 — Generalization-First English → RawLang → Python

RawIntent turns ordinary English into a small validated language called RawLang, then executes only registered Python capabilities.

Version 4.0 changes the training system so the neural compiler is explicitly trained and measured on phrases it did not see during training.

messy / unseen English
        ↓
RawIntent compiler model
        ↓
RawLang
        ↓
parser + schema validation + permission gates
        ↓
registered Python functions
        ↓
result or plain-English error

The neural model does not execute arbitrary Python. It only emits RawLang.

What changed in 4.0

RawIntent 4.0 adds:

  • phrase-family train/validation/test splits performed before augmentation
  • exact leakage checks so one wording family cannot appear in multiple splits
  • held-out wording benchmarks for every action that has multiple example phrases
  • separate compositional generalization families
  • typo, slang, filler, synonym, casing, spacing, punctuation, and informal-English augmentation
  • explicit UNKNOWN_CAPABILITY output for requests the runtime does not support
  • unknown-capability precision / recall / F1 evaluation
  • syntax-validity, exact-match, semantic-match, and action-accuracy metrics
  • per-category metrics for single commands, compositions, and unsupported requests
  • optional external validation data during training, preventing random near-duplicate leakage

This is the key idea: the model should learn the meaning of known capabilities rather than memorize exact sentences.

Install

Runtime only:

pip install .

With the from-scratch model trainer:

pip install ".[ml]"

With HTTP + ML extras:

pip install ".[all]"

RawLang

Single command:

random.integer low=1 high=100

Composition:

LET n = random.integer low=1 high=100
LET root = math.sqrt value=$n
RETURN $root

Unsupported capability:

UNKNOWN_CAPABILITY

That last form is valid RawLang. It means the compiler understood that the request is outside the current capability schema instead of inventing a command.

Use immediately without a trained model

from rawintent import RawIntentLanguage

lang = RawIntentLanguage()

print(lang.compile("pick a random number between 1 and 10"))
# random.integer low=1 high=10

result = lang.run('sha256 hash "hello"')
print(result.value)

Without a checkpoint, RawIntent uses the deterministic parser. This remains useful as a runtime, fallback, and data-bootstrap system.

Generate a true generalization dataset

Recommended:

rawintent-lm generate-generalization-data \
  --out-dir rawintent-generalization \
  --train-variants 8 \
  --eval-variants 6 \
  --samples 3 \
  --seed 7

It writes:

rawintent-generalization/
├── train.jsonl
├── validation.jsonl
├── test.jsonl
└── manifest.json

The split occurs before augmentation.

If an action has phrases such as:

TRAIN FAMILY:
pick a random number

VALIDATION FAMILY:
choose an integer in a range

TEST FAMILY:
grab a number between two values

then every noisy variation of the test family remains outside training.

This prevents a misleading setup where:

TRAIN: pick a random number
TEST:  please pick a random number

would be counted as generalization.

Dataset records

Records contain family and category metadata:

{
  "source": "yo can u grab an integer minimum 10 maximum 100",
  "target": "random.integer low=10 high=100",
  "action": "random.integer",
  "family": "random.integer:phrase:2:...",
  "kind": "single",
  "split": "test"
}

Unsupported requests look like:

{
  "source": "train a tensorflow image classifier",
  "target": "UNKNOWN_CAPABILITY",
  "action": "unknown",
  "family": "unknown:tensorflow",
  "kind": "unknown",
  "split": "test"
}

Train from scratch with held-out validation phrases

rawintent-lm train \
  --data rawintent-generalization/train.jsonl \
  --validation-data rawintent-generalization/validation.jsonl \
  --out rawintent-compiler.pt \
  --epochs 20 \
  --d-model 192 \
  --heads 6 \
  --layers 4 \
  --ff 768

When --validation-data is supplied, RawIntent checks that no phrase family appears in both training and validation.

The model initializes with new random weights. No pretrained model or external tokenizer is downloaded.

Evaluate phrases the model never saw

rawintent-lm evaluate \
  --checkpoint rawintent-compiler.pt \
  --data rawintent-generalization/test.jsonl \
  --out generalization-results.json

Metrics include:

{
  "syntax_validity": 0.0,
  "exact_match": 0.0,
  "semantic_match": 0.0,
  "action_accuracy": 0.0,
  "unknown": {
    "precision": 0.0,
    "recall": 0.0,
    "f1": 0.0
  },
  "by_kind": {
    "single": {},
    "composition": {},
    "unknown": {}
  }
}

The zeros above are only an example of the JSON shape, not expected trained-model performance.

What the metrics mean

  • syntax validity — did the model output valid RawLang?
  • exact match — exact target text match
  • semantic match — same parsed program even if argument order differs
  • action accuracy — did it choose the correct RawLang action sequence?
  • unknown precision/recall/F1 — can it reject truly unsupported capabilities without over-rejecting supported ones?

Failures are included in the report with the original English, target RawLang, generated RawLang, phrase family, and syntax error if applicable.

Benchmark the deterministic parser too

rawintent-lm evaluate \
  --deterministic \
  --data rawintent-generalization/test.jsonl

This gives you a baseline against which to compare the neural compiler.

Why unseen phrases can work

The model is not trained to memorize one sentence per command. It sees many actions expressed through multiple wording families and noisy forms.

For example it may train on:

choose an integer from 1 to 20
pick a random value between 1 and 20

while the test set contains only:

grab me some number anywhere in the one-to-twenty range

All map to:

random.integer low=1 high=20

The held-out test split is designed to tell you whether that generalization is actually happening.

Compositional generalization

Training also includes multi-step programs:

LET password = secrets.password length=24
LET lower = text.lower text=$password
RETURN $lower

Entire composition families are held out. This tests whether the model can learn concepts such as:

produce value A
feed A into operation B
return B

rather than memorizing one fixed multi-step example.

Unknown concepts instead of hallucinated commands

The training data contains unsupported concepts such as external account operations, image/video generation, databases, deployment tools, live external data, and other capabilities not present in the registered runtime.

Expected output:

UNKNOWN_CAPABILITY

Executing that RawLang returns a structured English error:

from rawintent import RawLangExecutor

result = RawLangExecutor().execute("UNKNOWN_CAPABILITY")
print(result.error.code)
# UNKNOWN_CAPABILITY

print(result.error.message)
# I understand the request, but this RawIntent runtime does not have a registered capability for it yet.

The custom model

RawIntent includes a small encoder-decoder Transformer implemented in PyTorch:

from rawintent.compiler.model import ModelConfig, RawCompilerModel

config = ModelConfig(
    d_model=192,
    nhead=6,
    num_encoder_layers=4,
    num_decoder_layers=4,
    dim_feedforward=768,
)

model = RawCompilerModel(config)

The tokenizer is a fixed UTF-8 byte tokenizer with 260 token IDs, so arbitrary Unicode text does not require a downloaded vocabulary.

Compile with a trained checkpoint

rawintent-lm compile \
  --checkpoint rawintent-compiler.pt \
  "yo grab me some num anywhere from 20 up to 80"

Python:

from rawintent import RawIntentLanguage

lang = RawIntentLanguage("rawintent-compiler.pt")
print(lang.compile("make me a strong 24 character password"))

Normal application mode can fall back to deterministic parsing if neural output is invalid. Evaluation bypasses that fallback so the benchmark measures the neural model itself.

Execute RawLang directly

from rawintent import RawLangExecutor

program = """
LET n = math.power base=3 exponent=2
LET root = math.sqrt value=$n
RETURN $root
"""

result = RawLangExecutor().execute(program)
print(result.value)

Safety boundary

RawLang does not provide arbitrary eval(), exec(), os.system(), or unrestricted subprocess execution.

Unknown actions are rejected against the registered capability schema.

File changes are opt-in:

from rawintent import EnglishPython

engine = EnglishPython(allow_file_changes=True)

Network writes are separately opt-in:

engine = EnglishPython(
    allow_network=True,
    allow_network_writes=True,
)

Python capability coverage

The runtime retains the common capability packs from RawIntent 2/3, including operations backed by:

  • random, secrets
  • math, statistics, decimal, fractions
  • datetime, time, calendar, uuid
  • hashlib, hmac, Base64, hex, gzip, zlib
  • json, csv, configparser
  • re, text, HTML, URL/query helpers
  • pathlib, selected os / shutil
  • collections, itertools, heapq, bisect, fnmatch, difflib
  • HTTP with requests when installed, otherwise urllib

Inspect the exact language with:

rawintent-lm schema

Third-party packages

Explicit module exposure remains available:

from rawintent import EnglishPython

engine = EnglishPython()
engine.expose_module("numpy", ["mean", "median", "std"])

For neural understanding, add training families for those new RawLang actions, regenerate the split, and retrain.

CLI summary

rawintent-lm schema
rawintent-lm generate-data --out bootstrap.jsonl
rawintent-lm generate-generalization-data --out-dir rawintent-generalization
rawintent-lm train --data rawintent-generalization/train.jsonl --validation-data rawintent-generalization/validation.jsonl --out model.pt
rawintent-lm evaluate --checkpoint model.pt --data rawintent-generalization/test.jsonl
rawintent-lm evaluate --deterministic --data rawintent-generalization/test.jsonl
rawintent-lm compile --checkpoint model.pt "pick any number one to ten"
rawintent-lm run --checkpoint model.pt 'sha256 hash "hello"'
rawintent-lm run-rawlang 'math.sqrt value=81'

Important reality check

This architecture makes unseen phrasing possible and measurable; it does not make a small from-scratch model magically understand every sentence. Quality still depends on training diversity, model size, optimization, and the amount of real human language in the dataset.

Version 4.0's purpose is to make that distinction measurable: a model only scores well if it succeeds on wording and composition families that were genuinely held out from training.

Development

python -m pytest -q

MIT licensed.

Metadata

Release files for rawintent 4.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rawintent 4.0.0
File Size Uploaded
rawintent-4.0.0.tar.gz 64.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rawintent 4.0.0
File Interpreter ABI Platform
rawintent-4.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 126.6 kB

Release files / rawintent-4.0.0.tar.gz

Download URL rawintent-4.0.0.tar.gz
Size 64.1 kB
Tags Source
SHA-256 checksum
How to use checksums
efacb3bb12224b551007c5817dcb33de9eb3bd977532029cbfbfca129b1f9035
BLAKE2b-256 checksum
How to use checksums
acbf46bf0a46135694a8cf9e7a111bb12a171e615052e6ff4e2693f66e27ad07
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / rawintent-4.0.0-py3-none-any.whl

Download URL rawintent-4.0.0-py3-none-any.whl
Size 62.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a36a2b95fc2a2f606c9f38da78ad1b91ea557bbc6d709909239a439efc551cfd
BLAKE2b-256 checksum
How to use checksums
ca7f0cda3fa54fda415df1ad2d8f9d171b7a13d02b69716d36f5475b3699df7b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

5.2.2

2 release files

5.2.1

2 release files

5.2.0

2 release files

5.1.1

2 release files

5.1.0

2 release files

5.0.0

2 release files

4.1.0

2 release files

This release

4.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page