Skip to main content

locate-anything

A tiny Python package that wraps the full nvidia/LocateAnything-3B workflow — dependency setup, model loading, preprocessing, inference, and output parsing — behind a single class, so you don't have to re-copy notebook cells every time.

Install

Because torch needs a CUDA-specific build, install it first from the PyTorch index, then install this package:

pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu124
pip install locate-anything

(Or, from a local checkout: pip install . after the torch step above.)

LocateAnything-3B is a gated model, so you'll also need to authenticate once:

huggingface-cli login

or pass a token directly when creating the client (see below).

Usage

Python API

from locate_anything import LocateAnything

# Loads tokenizer, processor, and model once (put this outside any loop)
la = LocateAnything()  # optionally: LocateAnything(hf_token="hf_...")

result = la.detect("screenshot.png", save_path="detected.png")

print(result["total"])         # number of detections
print(result["counts"])        # {"button": 12, "icon": 30, ...}
print(result["detections"])    # [{"label": "button", "bbox_pixels": [x1,y1,x2,y2]}, ...]
result["annotated_image"].show()  # PIL.Image with boxes drawn

Custom categories:

result = la.detect(
    "screenshot.png",
    categories=["play button", "volume slider", "progress bar"],
)

Batch of images:

results = la.detect_batch(["a.png", "b.png", "c.png"])

Command line

locate-anything screenshot.png -o detected.png --json-output detections.json

What it handles for you

  • Dependency management — pinned versions declared in pyproject.toml (torch must be installed separately due to the CUDA index URL requirement).
  • Model initialization — tokenizer, processor, and model loaded once per LocateAnything instance, in bfloat16 on the best available device.
  • Preprocessing — builds the chat-template prompt and processes image/video inputs for you.
  • Inference — calls model.generate(...) with sane defaults (generation_mode="hybrid", use_cache=True).
  • Output processing — parses the model's <box> tags back into pixel coordinates, tallies counts per label, and optionally draws + saves an annotated image.

Package layout

locate_anything/
├── __init__.py        # public API: LocateAnything, DEFAULT_CATEGORIES
├── config.py           # default model name, categories, regex pattern
├── core.py              # LocateAnything class (load + detect)
├── postprocessing.py    # parse_detections(), draw_detections()
└── cli.py                # `locate-anything` command line entry point

Release files for locate-anything 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for locate-anything 0.1.0
File Size Uploaded
locate_anything-0.1.0.tar.gz 6.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for locate-anything 0.1.0
File Interpreter ABI Platform
locate_anything-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 14.5 kB

Release files / locate_anything-0.1.0.tar.gz

Download URL locate_anything-0.1.0.tar.gz
Size 6.6 kB
Tags Source
SHA-256 checksum
How to use checksums
0a5ff6820b2b5133456da28aa020690b94a8a5df0b393aa82886c44405dd8944
BLAKE2b-256 checksum
How to use checksums
e26d703e9177bf2f7e3cfb2e13a97ed4411bfeab0575cd2fb3ca5b20af8e2899
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.8

Release files / locate_anything-0.1.0-py3-none-any.whl

Download URL locate_anything-0.1.0-py3-none-any.whl
Size 7.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
783ab9ed46fafb3baba6bdcbbee5a1dd25bad2377b657c2bd24e0508c34dbb67
BLAKE2b-256 checksum
How to use checksums
58570201af068df00a837f13faad4294e8df0f7efb8f07938d7716d283916e1f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.8

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page