RawIntent 4.0 — Generalization-First English → RawLang → Python
RawIntent turns ordinary English into a small validated language called RawLang, then executes only registered Python capabilities.
Version 4.0 changes the training system so the neural compiler is explicitly trained and measured on phrases it did not see during training.
messy / unseen English
↓
RawIntent compiler model
↓
RawLang
↓
parser + schema validation + permission gates
↓
registered Python functions
↓
result or plain-English error
The neural model does not execute arbitrary Python. It only emits RawLang.
What changed in 4.0
RawIntent 4.0 adds:
- phrase-family train/validation/test splits performed before augmentation
- exact leakage checks so one wording family cannot appear in multiple splits
- held-out wording benchmarks for every action that has multiple example phrases
- separate compositional generalization families
- typo, slang, filler, synonym, casing, spacing, punctuation, and informal-English augmentation
- explicit
UNKNOWN_CAPABILITYoutput for requests the runtime does not support - unknown-capability precision / recall / F1 evaluation
- syntax-validity, exact-match, semantic-match, and action-accuracy metrics
- per-category metrics for single commands, compositions, and unsupported requests
- optional external validation data during training, preventing random near-duplicate leakage
This is the key idea: the model should learn the meaning of known capabilities rather than memorize exact sentences.
Install
Runtime only:
pip install .
With the from-scratch model trainer:
pip install ".[ml]"
With HTTP + ML extras:
pip install ".[all]"
RawLang
Single command:
random.integer low=1 high=100
Composition:
LET n = random.integer low=1 high=100
LET root = math.sqrt value=$n
RETURN $root
Unsupported capability:
UNKNOWN_CAPABILITY
That last form is valid RawLang. It means the compiler understood that the request is outside the current capability schema instead of inventing a command.
Use immediately without a trained model
from rawintent import RawIntentLanguage
lang = RawIntentLanguage()
print(lang.compile("pick a random number between 1 and 10"))
# random.integer low=1 high=10
result = lang.run('sha256 hash "hello"')
print(result.value)
Without a checkpoint, RawIntent uses the deterministic parser. This remains useful as a runtime, fallback, and data-bootstrap system.
Generate a true generalization dataset
Recommended:
rawintent-lm generate-generalization-data \
--out-dir rawintent-generalization \
--train-variants 8 \
--eval-variants 6 \
--samples 3 \
--seed 7
It writes:
rawintent-generalization/
├── train.jsonl
├── validation.jsonl
├── test.jsonl
└── manifest.json
The split occurs before augmentation.
If an action has phrases such as:
TRAIN FAMILY:
pick a random number
VALIDATION FAMILY:
choose an integer in a range
TEST FAMILY:
grab a number between two values
then every noisy variation of the test family remains outside training.
This prevents a misleading setup where:
TRAIN: pick a random number
TEST: please pick a random number
would be counted as generalization.
Dataset records
Records contain family and category metadata:
{
"source": "yo can u grab an integer minimum 10 maximum 100",
"target": "random.integer low=10 high=100",
"action": "random.integer",
"family": "random.integer:phrase:2:...",
"kind": "single",
"split": "test"
}
Unsupported requests look like:
{
"source": "train a tensorflow image classifier",
"target": "UNKNOWN_CAPABILITY",
"action": "unknown",
"family": "unknown:tensorflow",
"kind": "unknown",
"split": "test"
}
Train from scratch with held-out validation phrases
rawintent-lm train \
--data rawintent-generalization/train.jsonl \
--validation-data rawintent-generalization/validation.jsonl \
--out rawintent-compiler.pt \
--epochs 20 \
--d-model 192 \
--heads 6 \
--layers 4 \
--ff 768
When --validation-data is supplied, RawIntent checks that no phrase family appears in both training and validation.
The model initializes with new random weights. No pretrained model or external tokenizer is downloaded.
Evaluate phrases the model never saw
rawintent-lm evaluate \
--checkpoint rawintent-compiler.pt \
--data rawintent-generalization/test.jsonl \
--out generalization-results.json
Metrics include:
{
"syntax_validity": 0.0,
"exact_match": 0.0,
"semantic_match": 0.0,
"action_accuracy": 0.0,
"unknown": {
"precision": 0.0,
"recall": 0.0,
"f1": 0.0
},
"by_kind": {
"single": {},
"composition": {},
"unknown": {}
}
}
The zeros above are only an example of the JSON shape, not expected trained-model performance.
What the metrics mean
- syntax validity — did the model output valid RawLang?
- exact match — exact target text match
- semantic match — same parsed program even if argument order differs
- action accuracy — did it choose the correct RawLang action sequence?
- unknown precision/recall/F1 — can it reject truly unsupported capabilities without over-rejecting supported ones?
Failures are included in the report with the original English, target RawLang, generated RawLang, phrase family, and syntax error if applicable.
Benchmark the deterministic parser too
rawintent-lm evaluate \
--deterministic \
--data rawintent-generalization/test.jsonl
This gives you a baseline against which to compare the neural compiler.
Why unseen phrases can work
The model is not trained to memorize one sentence per command. It sees many actions expressed through multiple wording families and noisy forms.
For example it may train on:
choose an integer from 1 to 20
pick a random value between 1 and 20
while the test set contains only:
grab me some number anywhere in the one-to-twenty range
All map to:
random.integer low=1 high=20
The held-out test split is designed to tell you whether that generalization is actually happening.
Compositional generalization
Training also includes multi-step programs:
LET password = secrets.password length=24
LET lower = text.lower text=$password
RETURN $lower
Entire composition families are held out. This tests whether the model can learn concepts such as:
produce value A
feed A into operation B
return B
rather than memorizing one fixed multi-step example.
Unknown concepts instead of hallucinated commands
The training data contains unsupported concepts such as external account operations, image/video generation, databases, deployment tools, live external data, and other capabilities not present in the registered runtime.
Expected output:
UNKNOWN_CAPABILITY
Executing that RawLang returns a structured English error:
from rawintent import RawLangExecutor
result = RawLangExecutor().execute("UNKNOWN_CAPABILITY")
print(result.error.code)
# UNKNOWN_CAPABILITY
print(result.error.message)
# I understand the request, but this RawIntent runtime does not have a registered capability for it yet.
The custom model
RawIntent includes a small encoder-decoder Transformer implemented in PyTorch:
from rawintent.compiler.model import ModelConfig, RawCompilerModel
config = ModelConfig(
d_model=192,
nhead=6,
num_encoder_layers=4,
num_decoder_layers=4,
dim_feedforward=768,
)
model = RawCompilerModel(config)
The tokenizer is a fixed UTF-8 byte tokenizer with 260 token IDs, so arbitrary Unicode text does not require a downloaded vocabulary.
Compile with a trained checkpoint
rawintent-lm compile \
--checkpoint rawintent-compiler.pt \
"yo grab me some num anywhere from 20 up to 80"
Python:
from rawintent import RawIntentLanguage
lang = RawIntentLanguage("rawintent-compiler.pt")
print(lang.compile("make me a strong 24 character password"))
Normal application mode can fall back to deterministic parsing if neural output is invalid. Evaluation bypasses that fallback so the benchmark measures the neural model itself.
Execute RawLang directly
from rawintent import RawLangExecutor
program = """
LET n = math.power base=3 exponent=2
LET root = math.sqrt value=$n
RETURN $root
"""
result = RawLangExecutor().execute(program)
print(result.value)
Safety boundary
RawLang does not provide arbitrary eval(), exec(), os.system(), or unrestricted subprocess execution.
Unknown actions are rejected against the registered capability schema.
File changes are opt-in:
from rawintent import EnglishPython
engine = EnglishPython(allow_file_changes=True)
Network writes are separately opt-in:
engine = EnglishPython(
allow_network=True,
allow_network_writes=True,
)
Python capability coverage
The runtime retains the common capability packs from RawIntent 2/3, including operations backed by:
random,secretsmath,statistics,decimal,fractionsdatetime,time,calendar,uuidhashlib,hmac, Base64, hex, gzip, zlibjson,csv,configparserre, text, HTML, URL/query helperspathlib, selectedos/shutilcollections,itertools,heapq,bisect,fnmatch,difflib- HTTP with
requestswhen installed, otherwiseurllib
Inspect the exact language with:
rawintent-lm schema
Third-party packages
Explicit module exposure remains available:
from rawintent import EnglishPython
engine = EnglishPython()
engine.expose_module("numpy", ["mean", "median", "std"])
For neural understanding, add training families for those new RawLang actions, regenerate the split, and retrain.
CLI summary
rawintent-lm schema
rawintent-lm generate-data --out bootstrap.jsonl
rawintent-lm generate-generalization-data --out-dir rawintent-generalization
rawintent-lm train --data rawintent-generalization/train.jsonl --validation-data rawintent-generalization/validation.jsonl --out model.pt
rawintent-lm evaluate --checkpoint model.pt --data rawintent-generalization/test.jsonl
rawintent-lm evaluate --deterministic --data rawintent-generalization/test.jsonl
rawintent-lm compile --checkpoint model.pt "pick any number one to ten"
rawintent-lm run --checkpoint model.pt 'sha256 hash "hello"'
rawintent-lm run-rawlang 'math.sqrt value=81'
Important reality check
This architecture makes unseen phrasing possible and measurable; it does not make a small from-scratch model magically understand every sentence. Quality still depends on training diversity, model size, optimization, and the amount of real human language in the dataset.
Version 4.0's purpose is to make that distinction measurable: a model only scores well if it succeeds on wording and composition families that were genuinely held out from training.
Development
python -m pytest -q
MIT licensed.
Metadata
Release files for rawintent 4.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rawintent-4.0.0.tar.gz | 64.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rawintent-4.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 126.6 kB
Release files / rawintent-4.0.0.tar.gz
| Download URL | rawintent-4.0.0.tar.gz |
|---|---|
| Size | 64.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
efacb3bb12224b551007c5817dcb33de9eb3bd977532029cbfbfca129b1f9035
|
|
BLAKE2b-256 checksum How to use checksums |
acbf46bf0a46135694a8cf9e7a111bb12a171e615052e6ff4e2693f66e27ad07
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|
Release files / rawintent-4.0.0-py3-none-any.whl
| Download URL | rawintent-4.0.0-py3-none-any.whl |
|---|---|
| Size | 62.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a36a2b95fc2a2f606c9f38da78ad1b91ea557bbc6d709909239a439efc551cfd
|
|
BLAKE2b-256 checksum How to use checksums |
ca7f0cda3fa54fda415df1ad2d8f9d171b7a13d02b69716d36f5475b3699df7b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|