HuggingFace Plugin for Vision Agents
HuggingFace integration for Vision Agents. Supports cloud-based inference via HuggingFace's Inference Providers API and local on-device inference via Transformers.
Installation
# Cloud inference (HuggingFace Inference API)
uv add "vision-agents[huggingface]"
# or directly
uv add vision-agents-plugins-huggingface
# Local inference (Transformers - LLM, VLM, object detection)
uv add "vision-agents-plugins-huggingface[transformers]"
# Local inference with quantization (4-bit / 8-bit)
uv add "vision-agents-plugins-huggingface[transformers-quantized]"
Cloud Inference (API-based)
Configuration
export HF_TOKEN=your_huggingface_token
Text-only LLM
from vision_agents.plugins import huggingface
llm = huggingface.LLM(
model="meta-llama/Meta-Llama-3-8B-Instruct",
provider="together", # or "groq", "cerebras", etc.
)
response = await llm.simple_response("Hello, how are you?")
print(response.text)
Vision Language Model (VLM)
from vision_agents.plugins import huggingface
vlm = huggingface.VLM(
model="Qwen/Qwen2-VL-7B-Instruct",
fps=1,
frame_buffer_seconds=10,
)
response = await vlm.simple_response("What do you see?")
print(response.text)
Local Inference (Transformers)
Runs models directly on your hardware (GPU/CPU/MPS). Requires the [transformers] extra.
Local LLM
from vision_agents.plugins import huggingface
llm = huggingface.TransformersLLM(
model="meta-llama/Llama-3.2-3B-Instruct",
)
@llm.register_function()
async def get_weather(city: str) -> str:
"""Get the current weather for a city."""
return f"The weather in {city} is sunny."
response = await llm.simple_response("What's the weather in Paris?")
Supported Providers
With 4-bit quantization (~4x memory reduction)
llm = huggingface.TransformersLLM( model="meta-llama/Llama-3.2-3B-Instruct", quantization="4bit", )
**Parameters:**
- `model` (str): HuggingFace model ID
- `device`: `"auto"`, `"cuda"`, `"mps"`, or `"cpu"`
- `quantization`: `"none"`, `"4bit"`, or `"8bit"`
- `torch_dtype`: `"auto"`, `"float16"`, `"bfloat16"`, or `"float32"`
- `max_new_tokens` (int): Max tokens per response (default: 512)
### Local VLM
```python
from vision_agents.plugins import huggingface
vlm = huggingface.TransformersVLM(
model="Qwen/Qwen2-VL-2B-Instruct",
)
Parameters:
model(str): HuggingFace model IDdevice:"auto","cuda","mps", or"cpu"quantization:"none","4bit", or"8bit"fps(int): Frames per second to capture (default: 1)frame_buffer_seconds(int): Seconds of video to buffer (default: 10)max_frames(int): Max frames per inference (default: 4)
Local Object Detection
Runs detection models like RT-DETRv2 on video frames and emits DetectionCompletedEvent with bounding boxes.
from vision_agents.core import Agent
from vision_agents.plugins import huggingface
processor = huggingface.TransformersDetectionProcessor(
model="PekingU/rtdetr_v2_r101vd",
conf_threshold=0.5,
fps=5,
)
agent = Agent(processors=[processor], ...)
@agent.events.subscribe
async def on_detection(event: huggingface.DetectionCompletedEvent):
for obj in event.objects:
print(f"{obj['label']} ({obj['confidence']:.0%})")
Parameters:
model(str): HuggingFace model ID (default:"PekingU/rtdetr_v2_r101vd")conf_threshold(float): Confidence threshold 0-1 (default: 0.5)fps(int): Frame processing rate (default: 10)classes(list[str], optional): Filter to specific class namesdevice:"auto","cuda","mps", or"cpu"annotate(bool): Draw bounding boxes on output video (default: True)
Metadata
Release files for vision-agents-plugins-huggingface 0.6.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| vision_agents_plugins_huggingface-0.6.9.tar.gz | 21.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| vision_agents_plugins_huggingface-0.6.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 52.8 kB
Release files / vision_agents_plugins_huggingface-0.6.9.tar.gz
| Download URL | vision_agents_plugins_huggingface-0.6.9.tar.gz |
|---|---|
| Size | 21.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f3ea38a60412d2b23143621b6038302c3a329f5a89fd65a72c35572effa460b3
|
|
BLAKE2b-256 checksum How to use checksums |
7af69d4b5b800416593f9f40907e5891446a8149f8bd7cb3f7aa7783372ab2d7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.10.10 {"installer":{"name":"uv","version":"0.10.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / vision_agents_plugins_huggingface-0.6.9-py3-none-any.whl
| Download URL | vision_agents_plugins_huggingface-0.6.9-py3-none-any.whl |
|---|---|
| Size | 31.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
db809301d216bfde6d22273c006216af957e307ae5b7d4215b5563c7e9b3e74e
|
|
BLAKE2b-256 checksum How to use checksums |
9e8de4136ef1a7b2d17906887886be55701ee95023a9482706c1229e6c086a91
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.10.10 {"installer":{"name":"uv","version":"0.10.10","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|