NeurInferno
Field-boundary inference for unlabeled binary protocol messages
Live demo · Model · Dataset · Report a bug
NeurInferno analyzes a batch of binary messages that share a format and predicts the byte gaps most likely to be field boundaries. It works without field labels at inference time and returns boundaries, confidence scores, and the corresponding byte segments.
Install
python -m pip install neurinferno
Python 3.10 or newer is required. Inference works on CPU; training requires a CUDA-capable environment.
Quick start
from neurinferno import FieldBoundaryModel
messages = [
"0001080006040001900c2d9bfa4649e7160700000000000043f03612",
"0001080006040002fbccad5c9fb1d014735252376fd2446375217d01",
"0001080006040002a388be2d1d684df0f31908ae1acad25b4dca7aa8",
"0001080006040001d41cd6668e87ac1e6d7b0000000000008593a2d9",
]
model = FieldBoundaryModel.from_pretrained()
for message_index, result in enumerate(model.infer(messages), start=1):
print(f"message {message_index}")
for segment in result.segments:
print(f" [{segment.start}:{segment.end}] {segment.hex}")
The first call to from_pretrained() downloads and caches the model artifact.
For reliable cross-message statistics, provide at least four messages from the
same format.
Command line
Place one hexadecimal message on each line of a file:
neurinferno infer messages.hex
neurinferno infer messages.hex --threshold 0.85
neurinferno infer messages.hex --format json
Standard input is supported when the path is omitted:
printf '0001aaff\n0002bbff\n' | neurinferno infer
Use a local checkpoint with --ckpt path/to/model.ckpt, or select another
Hugging Face model repository with --repo namespace/name.
Input and output
Inputs may use plain hex or common separators:
0001080006040001
0x00 0x01 0x08 0x00 0x06 0x04 0x00 0x02
00:01:08:00:06:04:00:03
Unsupported characters, empty messages, invalid thresholds, and oversized
batches produce explicit errors. Messages longer than the model context are
marked as truncated in MessageResult.
Each result contains:
| Attribute | Meaning |
|---|---|
hex |
Hexadecimal bytes processed by the model. |
n_bytes |
Number of processed bytes. |
original_n_bytes |
Input length before context truncation. |
truncated |
Whether the input exceeded the model context. |
scores |
Boundary probability for every gap between adjacent bytes. |
cuts |
Boolean decisions after applying the threshold. |
segments |
Predicted fields with start offset, end offset, and hex. |
How it works
flowchart LR
A[Same-format messages] --> B[Byte encoder]
B --> C[Cross-message statistics]
D[Frozen byte language model] --> E[Entropy features]
C --> F[Boundary head]
E --> F
F --> G[Gap scores and segments]
A byte-level transformer processes the messages together. Cross-message mean, maximum, and variance features capture how bytes behave at the same offset, while a frozen byte language model supplies entropy features. The boundary head assigns a probability to each gap between consecutive bytes.
Data
The full dataset is hosted separately from the Python package:
hf download sachithabey/neurinferno \
--repo-type dataset \
--local-dir data
This creates data/protocols/ with labeled protocol traces and data/grammar/
with synthetic formats.
Each messages.jsonl line has this structure:
{
"bytes_hex": "0001080006040001...",
"field_type_per_byte": [2, 2, 2, 2, 1, 1],
"boundary_per_gap": [0, 1, 0, 1, 1],
"format_id": "example_00000",
"endianness": "big"
}
| Field | Required | Meaning |
|---|---|---|
bytes_hex |
Yes | Lowercase hexadecimal bytes with an even length. |
boundary_per_gap |
Yes | Length n_bytes - 1; 1 marks a boundary after that byte. |
field_type_per_byte |
Yes | Length n_bytes; auxiliary training type identifiers. |
format_id |
Yes | Format identifier; identifiers ending in _corrupted are excluded from evaluation. |
endianness |
No | big or little. |
Field type identifiers are defined in
label_format.py.
Training and evaluation
Clone the repository and install the training dependencies:
git clone https://github.com/Sachithx/NeurInferno.git
cd NeurInferno
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[train]"
hf download sachithabey/neurinferno --repo-type dataset --local-dir data
Run training or evaluate existing checkpoints:
CUDA_VISIBLE_DEVICES=0 bash train.sh
bash download_checkpoints.sh
CUDA_VISIBLE_DEVICES=0 bash eval.sh
The training scripts skip folds that already have checkpoints. The configured LOPO and L4PO splits exclude held-out protocols from language-model training, main-model training, and fine-tuning.
Adding a protocol
- Create
data/protocols/<name>/messages.jsonl. - Add the name to
LOPO_PROTOCOLSindataset.pyandtrain.sh. - Add it to the required L4PO split definitions in
train.sh. - Train the affected folds and run the evaluation suite.
Directory names must be unique and contain only lowercase letters, numbers, and underscores.
Repository layout
src/neurinferno/ Installable package and model implementation
tests/ Unit and interface-helper tests
examples/ Small runnable examples
hf_space/ Hugging Face Space interface and deployment helper
data/ Training and evaluation data layout
train.sh Reproducible training entry point
eval.sh Evaluation entry point
Development
python -m pip install -e ".[dev,demo]"
ruff check .
ruff format --check .
pytest
python -m build
twine check dist/*
See CONTRIBUTING.md for contribution expectations and SECURITY.md for private vulnerability reporting.
Project status
NeurInferno is an alpha-stage research software project. Interfaces may evolve between minor releases. Pin a version when integrating it into another system.
License
Licensed under Apache-2.0. See LICENSE.
Metadata
Release files for neurinferno 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| neurinferno-0.2.0.tar.gz | 63.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| neurinferno-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 124.0 kB
Release files / neurinferno-0.2.0.tar.gz
| Download URL | neurinferno-0.2.0.tar.gz |
|---|---|
| Size | 63.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
82743cbc5fda43d875a00011e637eaeb6127a4a86089c13e021468fb755af12d
|
|
BLAKE2b-256 checksum How to use checksums |
99d88cd1d8295bc8323a11e72cd3a9046654252cf3d8ff9907645d8e0fba52bd
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.
Transparency logRelease files / neurinferno-0.2.0-py3-none-any.whl
| Download URL | neurinferno-0.2.0-py3-none-any.whl |
|---|---|
| Size | 60.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
08a2b4ecc27b248699447ff2ee827e806b3836da8435c40bcc7f0fb4b141b27d
|
|
BLAKE2b-256 checksum How to use checksums |
2eb1580a7928db261c4f17c48f1de58c14b3be17f2ba1a03a1119e7b86d91e52
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.
Transparency log