Skip to main content

NeurInferno

Field-boundary inference for unlabeled binary protocol messages

CI PyPI Python License Hugging Face Space

Live demo · Model · Dataset · Report a bug

NeurInferno analyzes a batch of binary messages that share a format and predicts the byte gaps most likely to be field boundaries. It works without field labels at inference time and returns boundaries, confidence scores, and the corresponding byte segments.

Install

python -m pip install neurinferno

Python 3.10 or newer is required. Inference works on CPU; training requires a CUDA-capable environment.

Quick start

from neurinferno import FieldBoundaryModel

messages = [
    "0001080006040001900c2d9bfa4649e7160700000000000043f03612",
    "0001080006040002fbccad5c9fb1d014735252376fd2446375217d01",
    "0001080006040002a388be2d1d684df0f31908ae1acad25b4dca7aa8",
    "0001080006040001d41cd6668e87ac1e6d7b0000000000008593a2d9",
]

model = FieldBoundaryModel.from_pretrained()
for message_index, result in enumerate(model.infer(messages), start=1):
    print(f"message {message_index}")
    for segment in result.segments:
        print(f"  [{segment.start}:{segment.end}] {segment.hex}")

The first call to from_pretrained() downloads and caches the model artifact. For reliable cross-message statistics, provide at least four messages from the same format.

Command line

Place one hexadecimal message on each line of a file:

neurinferno infer messages.hex
neurinferno infer messages.hex --threshold 0.85
neurinferno infer messages.hex --format json

Standard input is supported when the path is omitted:

printf '0001aaff\n0002bbff\n' | neurinferno infer

Use a local checkpoint with --ckpt path/to/model.ckpt, or select another Hugging Face model repository with --repo namespace/name.

Input and output

Inputs may use plain hex or common separators:

0001080006040001
0x00 0x01 0x08 0x00 0x06 0x04 0x00 0x02
00:01:08:00:06:04:00:03

Unsupported characters, empty messages, invalid thresholds, and oversized batches produce explicit errors. Messages longer than the model context are marked as truncated in MessageResult.

Each result contains:

Attribute Meaning
hex Hexadecimal bytes processed by the model.
n_bytes Number of processed bytes.
original_n_bytes Input length before context truncation.
truncated Whether the input exceeded the model context.
scores Boundary probability for every gap between adjacent bytes.
cuts Boolean decisions after applying the threshold.
segments Predicted fields with start offset, end offset, and hex.

How it works

flowchart LR
    A[Same-format messages] --> B[Byte encoder]
    B --> C[Cross-message statistics]
    D[Frozen byte language model] --> E[Entropy features]
    C --> F[Boundary head]
    E --> F
    F --> G[Gap scores and segments]

A byte-level transformer processes the messages together. Cross-message mean, maximum, and variance features capture how bytes behave at the same offset, while a frozen byte language model supplies entropy features. The boundary head assigns a probability to each gap between consecutive bytes.

Data

The full dataset is hosted separately from the Python package:

hf download sachithabey/neurinferno \
  --repo-type dataset \
  --local-dir data

This creates data/protocols/ with labeled protocol traces and data/grammar/ with synthetic formats.

Each messages.jsonl line has this structure:

{
  "bytes_hex": "0001080006040001...",
  "field_type_per_byte": [2, 2, 2, 2, 1, 1],
  "boundary_per_gap": [0, 1, 0, 1, 1],
  "format_id": "example_00000",
  "endianness": "big"
}
Field Required Meaning
bytes_hex Yes Lowercase hexadecimal bytes with an even length.
boundary_per_gap Yes Length n_bytes - 1; 1 marks a boundary after that byte.
field_type_per_byte Yes Length n_bytes; auxiliary training type identifiers.
format_id Yes Format identifier; identifiers ending in _corrupted are excluded from evaluation.
endianness No big or little.

Field type identifiers are defined in label_format.py.

Training and evaluation

Clone the repository and install the training dependencies:

git clone https://github.com/Sachithx/NeurInferno.git
cd NeurInferno
python -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[train]"
hf download sachithabey/neurinferno --repo-type dataset --local-dir data

Run training or evaluate existing checkpoints:

CUDA_VISIBLE_DEVICES=0 bash train.sh
bash download_checkpoints.sh
CUDA_VISIBLE_DEVICES=0 bash eval.sh

The training scripts skip folds that already have checkpoints. The configured LOPO and L4PO splits exclude held-out protocols from language-model training, main-model training, and fine-tuning.

Adding a protocol

  1. Create data/protocols/<name>/messages.jsonl.
  2. Add the name to LOPO_PROTOCOLS in dataset.py and train.sh.
  3. Add it to the required L4PO split definitions in train.sh.
  4. Train the affected folds and run the evaluation suite.

Directory names must be unique and contain only lowercase letters, numbers, and underscores.

Repository layout

src/neurinferno/   Installable package and model implementation
tests/             Unit and interface-helper tests
examples/          Small runnable examples
hf_space/          Hugging Face Space interface and deployment helper
data/              Training and evaluation data layout
train.sh           Reproducible training entry point
eval.sh            Evaluation entry point

Development

python -m pip install -e ".[dev,demo]"
ruff check .
ruff format --check .
pytest
python -m build
twine check dist/*

See CONTRIBUTING.md for contribution expectations and SECURITY.md for private vulnerability reporting.

Project status

NeurInferno is an alpha-stage research software project. Interfaces may evolve between minor releases. Pin a version when integrating it into another system.

License

Licensed under Apache-2.0. See LICENSE.

Metadata

Release files for neurinferno 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for neurinferno 0.2.0
File Size Uploaded
neurinferno-0.2.0.tar.gz 63.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for neurinferno 0.2.0
File Interpreter ABI Platform
neurinferno-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 124.0 kB

Release files / neurinferno-0.2.0.tar.gz

Download URL neurinferno-0.2.0.tar.gz
Size 63.8 kB
Tags Source
SHA-256 checksum
How to use checksums
82743cbc5fda43d875a00011e637eaeb6127a4a86089c13e021468fb755af12d
BLAKE2b-256 checksum
How to use checksums
99d88cd1d8295bc8323a11e72cd3a9046654252cf3d8ff9907645d8e0fba52bd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release files / neurinferno-0.2.0-py3-none-any.whl

Download URL neurinferno-0.2.0-py3-none-any.whl
Size 60.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
08a2b4ecc27b248699447ff2ee827e806b3836da8435c40bcc7f0fb4b141b27d
BLAKE2b-256 checksum
How to use checksums
2eb1580a7928db261c4f17c48f1de58c14b3be17f2ba1a03a1119e7b86d91e52
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 21, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page