What it is
attr-eomt is a standalone EoMT (Encoder-only Mask Transformer) for instance segmentation and object detection, with one feature that sets it apart: independent per-instance attribute heads. Alongside the mask/box + class output, it predicts one or several orthogonal attributes for every detected instance — read straight off the same per-query embedding the detector already computes. No second model, no second pass, and the primary detection metric is untouched (the figure above tells the whole story).
It's a clean-room, Apache-2.0 reimplementation: the weights you train are yours to release.
from eomt import EoMT
model = EoMT("l") # fresh large model (DINOv2 backbone)
model.train(data="coco", epochs=50) # COCO 2017 auto-downloads if missing
model = EoMT("runs/train/eomt-l") # reload a run — size/classes/heads auto-detected
model.predict("images/", plot=True) # render masks/boxes + per-instance attributes
Architecture
EoMT is a DINOv2-with-registers ViT whose last few transformer blocks are augmented
with a fixed set of learnable queries (the Mask2Former idea) — each query is one
"slot" that latches onto one object instance. After the encoder runs, every query emits
a single vector, the per-query embedding of shape [B, Q, hidden]. The whole model
is then just "turn that embedding into predictions": a class head for the primary
label and a mask/box head for geometry. It is NMS-free, so two overlapping
garments stay two distinct queries instead of being merged — the property that lets
attributes stay attached to the right instance.
The attribute heads add nothing to this picture except themselves: they tap the exact same embedding (captured non-invasively with a forward hook), each a small classifier on top.
This collapses what is classically a two-stage pipeline — detect, crop each box, run a second classifier per crop — into a single pass. Attributes therefore cost only a thin head each, see full-image context (not just a cropped box), and never inherit a second model's cropping errors — the modern, single-stage formulation of the DETR / Mask2Former lineage (see the figure at the top).
Two model families: segmentation & detection
Both families share the same DINOv2 encoder, query mechanism, NMS-free matching and auxiliary heads — they differ only in the head on top and what they output:
| family | --task |
output | metric driving best.pt |
|---|---|---|---|
| instance (default) | instance |
per-instance masks + boxes + class | segm/mAP |
| detect | detect |
per-instance boxes + class (DETR-style box head, no masks) | bbox/mAP |
EoMT("l").train(data="coco", family="instance") # masks (default)
EoMT("l").train(data="coco", family="detect") # boxes only
The family is recorded in the checkpoint, so val / predict pick the right
post-processing automatically. Everything below applies identically to both.
Models & sizes
| size | backbone | hidden | layers | heads | queries |
|---|---|---|---|---|---|
s |
DINOv2-small | 384 | 12 | 6 | 100 |
b |
DINOv2-base | 768 | 12 | 12 | 200 |
l |
DINOv2-large | 1024 | 24 | 16 | 200 |
Default input is a patch-14-aligned square (644 = 14 × 46) so DINOv2 weights load 1:1.
Compute & inference speed
Measured on a single NVIDIA GeForce RTX 5090, 644 × 644 input, batch size 1.
GFLOPs are multiply-accumulates at that resolution (attention included); latency /
throughput are the median over 50 runs after warm-up, under torch.amp.autocast
(fp16) — the package's own inference path.
instance family (masks + boxes + class):
| size | params | GFLOPs | latency (fp16) | throughput (fp16) | throughput (fp32) |
|---|---|---|---|---|---|
s |
24.0 M | 128 | 8.4 ms | 119 img/s | 70 img/s |
b |
93.9 M | 430 | 17.4 ms | 58 img/s | 32 img/s |
l |
317 M | 1144 | 30.2 ms | 33 img/s | 15 img/s |
detect family (boxes + class, no mask head):
| size | params | GFLOPs | latency (fp16) | throughput (fp16) | throughput (fp32) |
|---|---|---|---|---|---|
s |
22.7 M | 89 | 2.9 ms | 348 img/s | 120 img/s |
b |
88.6 M | 276 | 5.3 ms | 190 img/s | 60 img/s |
l |
308 M | 881 | 13.6 ms | 74 img/s | 21 img/s |
Dropping the mask-upsampling head makes detect substantially lighter and ~1.3–3×
faster. Figures are for the detector itself (backbone + queries + heads); the
attribute heads add a thin linear/MLP per head and are negligible by design.
Factorizing the label space
This is the contribution. Conventional detectors fold every distinction into one flat
label space: an object's type × viewpoint × occlusion × … becomes a Cartesian product of
leaf classes that explodes combinatorially, starves each leaf of examples, and multiplies
the Hungarian matcher's targets. attr-eomt factorizes instead — a small, general primary
head plus independent attribute heads that add, not multiply.
Because the heads are independent, the primary taxonomy stays compact and every class keeps
its full sample count; attributes ride along for near-zero compute; and the model composes
attribute × class combinations that never appear in the training data — combinations a
flat label space cannot even represent.
Example: clothing with per-instance attributes
One model segments each garment (primary classes like vest_dress / short_sleeve_top
/ long_sleeve_dress / skirt / trousers …) and, for every detection, reads off
four independent attribute heads — scale (small / modest / large),
occlusion (no / slight / medium), zoom_in (no / medium / large) and
viewpoint (frontal / side / back). The renderer prints the primary class + score
on the first row and each attribute + its confidence on the rows beneath it.
The four attributes are orthogonal to the garment class — they vary independently — which is exactly the case that's awkward to fold into the primary class space. The same pattern fits any "class plus per-instance sub-labels" task: retail shelves → product + facing, documents → element + role, cells → type + health.
Trained on the public DeepFashion2 dataset (13 garment classes + 4 attribute heads) and rendered with the package's own renderer (
eomt.visualize.draw_instances).
Training — it rides on the detector's own match
Attributes never run their own matcher. Detection already solves "which query is responsible for which ground-truth object" via the Hungarian matcher; attributes simply reuse that same query→GT assignment and read the answer off the matched queries.
- Embedding source. Each head reads the per-query embedding — the input to EoMT's
class_predictor, captured with a forward hook ([B, Q, hidden]). - Matching. Supervision reuses EoMT's own Hungarian matcher
(
model.eomt.criterion.matcher), so every attribute is trained on the same query→GT assignment the detection loss used; the attribute is read after matching. - Gate. An optional IoU gate drops barely-overlapping matched pairs (common early in training) so attributes only learn from queries that actually localize the object.
- Loss. Cross-entropy per head over matched queries, summed across heads and scaled
by
aux_w(default1.0), added to the detector loss. Empty-match batches contribute a graph-preserving zero, and missing labels useignore_indexand contribute nothing. - Checkpoint selection is unchanged. The attribute "rides along": its per-head
matched-query accuracy is shown live and written to
metrics.csv, but never drivesbest.pt(stillsegm/mAPorbbox/mAP). - Inference. Each result attaches
aux = {head: {"ids", "probs"}}for the kept detections, andpredict(plot=True)renders each attribute next to the class label using names stored in the checkpoint.
Data format (auto-discovered from the COCO JSON)
Attributes live inside the COCO annotations — each annotation is already a
per-instance object, so alignment is automatic and pycocotools still parses it. Just
two additions to a standard COCO file; no YAML changes — heads (count, classes,
names) are discovered from the JSON, the same as nc.
1. A top-level attributes list — one entry per head, defining its vocabulary:
"attributes": [
{
"name": "scale",
"categories": [
{"id": 1, "name": "small"},
{"id": 2, "name": "modest"},
{"id": 3, "name": "large"}
]
},
{
"name": "viewpoint",
"categories": [
{"id": 0, "name": "frontal"},
{"id": 1, "name": "side"},
{"id": 2, "name": "back"}
]
}
]
2. A per-annotation attributes map — {head: raw_id} on each instance:
{
"id": 1, "image_id": 42, "category_id": 1,
"segmentation": [...], "bbox": [...], "area": 1234, "iscrowd": 0,
"attributes": {"scale": 3, "viewpoint": 0}
}
Notes:
- Raw ids are remapped to a contiguous
0..n-1per head (soscale's1/2/3become0/1/2);categoriesmay be omitted, in which case the id set is inferred. - A missing or out-of-vocab per-annotation value is ignored (
-100), not trained as class0— so a partially tagged dataset is valid: each head learns only from the instances that actually carry its value. A JSON with noattributes⇒ detection-only, exactly as before.
Class-conditional heads. Give an attribute definition an optional applies_to list of
primary-class names or ids, and that head is only trained on — and only emitted for —
instances of those classes (hard routing on the primary class). Omit it and the head applies
to every class. So different attributes can attach to different classes, each with its own
label set, in one model:
"attributes": [
{"name": "posture", "categories": [...], "applies_to": ["cat", "dog"]}
]
At inference a scoped head reports ids = -1 ("not applicable") for detections whose class it
does not cover. The scope is stored in the checkpoint, so it survives reload.
Sidecar format (optional). You can keep the COCO JSON as plain, standard COCO and put the
attributes beside it instead of inside it: an attributes.yaml schema in the dataset root plus
attributes/<split>.json values keyed by annotation id ({ann_id: {head: value}}). If present
(and the JSON has no embedded attributes), it is merged in memory at load — so a plain COCO
dataset always works and the sidecar is picked up automatically when you add it. Embedded
attributes in the JSON take precedence.
A tiny, self-contained example (two heads, including a non-contiguous id set) lives in sample_data/.
Install
pip install attr-eomt # from PyPI
pip install "attr-eomt[logging]" # + tensorboard/wandb
pip install -e ".[dev]" # from source (editable; [dev] adds pytest/build/twine)
Usage
Everything goes through one class. Initialize from a size (fresh model, pretrained
DINOv2 backbone) or from a checkpoint / run folder (family, size, classes, image
size, normalization and any auxiliary heads are auto-detected from the .pt):
from eomt import EoMT
# Train on COCO 2017 (auto-downloaded on first run):
EoMT("l").train(data="coco", epochs=50, batch=4)
# ...or any COCO-format dataset (point at its data.yaml):
EoMT("s").train(data="sample_data/data.yaml", epochs=1, batch=1)
# Validate and predict from a trained run:
EoMT("runs/train/eomt-l").val(data="coco")
EoMT("runs/train/eomt-l").predict("images/", plot=True) # writes annotated images
For the full training recipe, every train() knob, and int8 compression, see the
annotated explainer → — it's the deep dive.
Roadmap / future work
- Model export. ONNX / TensorRT (and friends) for deployment — currently out of scope; the inference path is being kept export-friendly.
- Keypoints. A keypoint/pose head family alongside
instanceanddetect(the code already carries afamilyparameter so new heads slot in without API churn). - Pretrained COCO checkpoints. None are published yet. COCO-trained
s/b/lweights will be released on the Hugging Face Hub (thefrom_pretrained/hf://loading plumbing is already in place and waiting for them). - Multi-image re-ID via contrastive learning. Train the auxiliary head with a contrastive objective so each instance's query embedding becomes a re-identification vector — matching the same object across images, frames and cameras for tracking and retrieval. The aux head already produces a per-instance embedding from the detector's own matched queries; re-ID reuses that signal instead of bolting on a separate model.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file attr_eomt-1.1.0.tar.gz.
File metadata
- Download URL: attr_eomt-1.1.0.tar.gz
- Upload date:
- Size: 1.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a5c448630f24650cecc681158fc40ce66e36f285d110281078a4d3f7f0998aa9
|
|
| MD5 |
71e5e9bf401906f1979554262d52d183
|
|
| BLAKE2b-256 |
fb62ec2160aca2015e8f6c12c753bc8d5600e24e33b998bfe5aa0f4050899549
|
File details
Details for the file attr_eomt-1.1.0-py3-none-any.whl.
File metadata
- Download URL: attr_eomt-1.1.0-py3-none-any.whl
- Upload date:
- Size: 104.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e28a82befde0a0e4eecdf2244c074ec525fc07fc806c01187a506294d5a62290
|
|
| MD5 |
56cb073e3dec0e5689ba43c71a322ec3
|
|
| BLAKE2b-256 |
5e41355e46082e415adee2a772de3825f062dcaded720549ad4787313528f476
|