Bytewise MIME detector
Bytewise is a standalone neural MIME detector trained on raw file bytes. It does not require Java, a Tika server, a filename, or a file extension. The repository preserves its complete research lineage: BFA/BFC baselines, neural experiments, strict-host validation, deduplication audits, and D3 reports.
The Bytewise 0.10.0 default is the hash-pinned, source-disjointly validated Bytewise-175 repaired model. It learns 175 MIME classes and exposes 177 supported outputs, including octet-stream fallback and M4V refinement. The prior 151-, 150-, 125-, 100-, 58-, and original 57-class models remain bundled as explicit rollback and reproducibility options.
Usage
Homebrew Python is an externally managed environment and must not be modified
with pip --break-system-packages. For the command-line application, install
Bytewise into an isolated Python 3.12 environment with uv:
uv tool install --python 3.12 "bytewise[metal]"
uv tool update-shell
Open a new terminal (or add $HOME/.local/bin to PATH) and verify it:
bytewise --version
bytewise model-info
bytewise doctor
Choose the runtime extra for the target platform:
# Apple Silicon GPU
python -m pip install "bytewise[metal]"
# Linux with an NVIDIA GPU and a current NVIDIA driver
python -m pip install "bytewise[cuda]"
# CPU inference on Linux or Windows
python -m pip install "bytewise[inference]"
The metal extra pins the validated TensorFlow 2.18 and tensorflow-metal
1.2 runtime. The cuda extra installs TensorFlow's pip-managed CUDA and cuDNN
libraries; the host still needs a compatible NVIDIA driver. Confirm CUDA is
visible with:
python -c 'import tensorflow as tf; print(tf.config.list_physical_devices("GPU"))'
For development or library use, keep the dependency in a project environment:
cd "$HOME/git/bytewise"
uv sync --python 3.12 --extra metal --group tests
uv run bytewise model-info
uv run python
Inside that uv run python session:
from bytewise import Detector
detector = Detector.load_default()
result = detector.detect_file("document.bin")
print(result.mime_type)
print(result.confidence)
print(result.alternatives)
The same model is available from the command line:
bytewise detect document.bin image.dat
bytewise detect document.bin --top-k 5 --threshold 0.80 --json
cat unknown.bin | bytewise detect -
bytewise supported
bytewise model-info
Select an earlier immutable model when reproducibility requires it:
legacy = Detector.load_default(model="legacy")
previous = Detector.load_default(model="previous")
production58 = Detector.load_default(model="production58")
production125_v2 = Detector.load_default(model="production125_v2")
java_router = Detector.load_default(model="java_router")
production175 = Detector.load_default(model="production175")
bytewise detect --model-version legacy document.bin
bytewise detect --model-version previous document.bin
bytewise detect --model-version production58 document.bin
bytewise detect --model-version production125_v2 document.bin
bytewise detect --model-version java_router Example.java
bytewise detect --model-version production175 document.bin
bytewise model-info --model-version legacy
Routine TensorFlow startup diagnostics are suppressed so CLI output remains script-friendly. To diagnose device selection or CUDA loading, enable them for one invocation:
bytewise doctor --require-gpu
bytewise detect --tensorflow-logs document.bin
Allow uncertain inputs to abstain so another detector can handle them:
detector = Detector.load_default(confidence_threshold=0.80)
result = detector.detect_bytes(payload)
if result.abstained:
# Fall back to tika-python, libmagic, or another detector.
pass
Production 175-class repaired model
The Bytewise 0.10.0 default is bytewise-175-repair-v1-candidate, the exact
passing two-epoch GPU checkpoint from the six-class repair cycle:
- Validation accuracy: 91.38%, up 1.17 points over the frozen 175 candidate
- Validation macro F1: 82.18%, up 2.68 points
- Source-disjoint repair accuracy: 58.78%, up 28.48 points
- Source-disjoint repair macro F1: 66.98%, up 28.33 points
- Learned labels: 175; supported outputs: 177
- Model SHA-256:
1cb9d0f958d6fb5c103897be2107aff4aeaf13c0e1cf29736a0e0ad9aa7d4ec9
The equivalent explicit alias is --model-version production175.
Preserved 151-class Java router
The Bytewise 0.9.0 model is bytewise-151-java-router-v1. It preserves the
0.8.0 result unless Java neural confidence and structural Java syntax both pass
the locked routing contract. Independent evaluation achieved 98.0% Java recall
and 98.99% Java F1 with zero changed predictions on the established Wikimedia
and complementary regression suites.
The equivalent explicit alias is --model-version java_router. The Bytewise
0.8.0 model remains selectable with --model-version production150.
Preserved 150-class fusion model
The Bytewise 0.8.0 default is bytewise-150-fusion-v1. It combines the frozen
150-class v3 model with the Bytewise 0.7.1 125-class parent using the externally
validated 0.4/0.6 probability weights and a 4x expansion-class bias:
- Internal accuracy: 92.35%; macro F1: 0.826
- Internal expansion accuracy: 86.75%; macro F1: 0.892
- Fresh mixed holdout: +1.14 accuracy points overall and +1.01 expansion points
- Expanded legacy confirmation: 85.41% accuracy, +0.54 points over the parent
- Learned labels: 150; supported outputs: 152
- Fusion bundle SHA-256:
e2f85dcaf8416b318fc6e01e311e8872d8be7d8b65fcb8b0cfd3a7084fd05e99
The exact contract, model components, and external validation are documented in the Bytewise-150 report.
The Bytewise 0.7.1 125-class model remains selectable with
--model-version production125_v2.
The Bytewise 0.6.0 125-class model remains selectable with
--model-version production125_v1.
The Bytewise 0.5.0 100-class model remains selectable with
--model-version previous. The Bytewise 0.3.0 58-class model remains
selectable with --model-version production58.
Legacy v1 model
The preserved legacy model is polar-byte-transformer-seed550-v1:
- Validation accuracy: 96.77%
- Independent-reference accuracy: 94.59%
- Input: first 4,096 raw bytes
- Labels: 57 MIME types
- Model SHA-256:
72a2f5f2dd0fbb4ffaf88488618bc8e034c03876c7cea94d11da23439ed5b849
See the model card and Full-v3 report for the complete evidence.
Repository organization
src/bytewise/: production API and migrated byte-frequency research codeartifacts/: immutable model release bundleconfigs/,scripts/: reproducible experiments and evaluationreports/: curated D3 reports from early pilots through Full-v3docs/: research and repository-extraction documentationtests/: standalone production and research regression tests
Dataset bytes, feature caches, databases, raw predictions, and transient logs are intentionally kept outside Git.
Release files for bytewise 0.10.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| bytewise-0.10.0.tar.gz | 16.6 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| bytewise-0.10.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 33.2 MB
Release files / bytewise-0.10.0.tar.gz
| Download URL | bytewise-0.10.0.tar.gz |
|---|---|
| Size | 16.6 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4b536b16761fca8d016ff704094960ab1f84d47a7baf79cea2ba28972b6affda
|
|
BLAKE2b-256 checksum How to use checksums |
c063dc474e01f19239a8bc956d1cc05a45537dd497515235d06e8432ba70b9e7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency logRelease files / bytewise-0.10.0-py3-none-any.whl
| Download URL | bytewise-0.10.0-py3-none-any.whl |
|---|---|
| Size | 16.6 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
91519a832f339482d9486b32b42c273ecad61b6a30a9631414a5b9bdf1e3fde7
|
|
BLAKE2b-256 checksum How to use checksums |
ddc13b128a105fd257a3339da3b2e90bde9f63a4e76bd29be8bb2c66dced7b62
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.
Transparency log