NobodyWho
Run LLMs locally and efficiently on any device
NobodyWho is a lightweight, open-source inference engine that makes it simple to run open-weights language models directly inside your Python applications. No API keys, no cloud infrastructure, no complexity—just fast, easy local AI.
Free to use in commercial projects under the EUPL-1.2 license, no API key required. Supports text, vision, hearing, speech-to-text, text-to-speech, voice activity detection, embeddings, RAG & tool calling.
- Documentation — Python & other frameworks documentation
- Discord — Get help, share ideas, and connect with other developers
- GitHub Issues — Report bugs
- GitHub Discussions — Ask questions and request features
Key Features
- Run locally, offline, for free - No API keys or cloud services required
- Fast, simple tool calling - Just pass normal Python functions
- Reliable tool execution - Automatically derives grammar from function signatures
- Speech-to-text - Transcribe spoken audio into text with Whisper models
- Text-to-speech - Generate natural-sounding speech from text
- Voice activity detection - Detect speech in an audio stream to know when to start and stop listening
- Vision & embeddings - Multimodal image and audio input, plus embeddings and reranking for semantic search and RAG
- Infinite conversations - Conversation-aware preemptive context shifting prevents mid-conversation crashes
- GPU accelerated - Vulkan-powered inference for maximum performance
- Thousands of compatible models - Works with any LLM in GGUF format
- Powered by llama.cpp - Built on the proven llama.cpp engine
Installation
pip install nobodywho
Supported Model Format
NobodyWho uses the GGUF format, a binary format optimized for fast loading and efficient LLM inference. A wide selection of GGUF models is available on Hugging Face.
You can also download a model without any extra dependencies by passing huggingface:owner/repo/filename.gguf where you'd normally pass the model path:
from nobodywho import Chat
chat = Chat("huggingface:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf")
Chat
Every interaction with your LLM starts by instantiating a Chat object. Call .ask() to send a message, and .completed() to block until the whole response is ready:
from nobodywho import Chat
chat = Chat("./model.gguf", system_prompt="You are a helpful assistant.")
response = chat.ask("Is water wet?").completed()
print(response) # Yes, indeed, water is wet!
Your messages and the model's responses are remembered inside the Chat object, so follow-up questions keep their context. You can pass "auto" as the model path to pick a chat model based on available memory.
See the Chat documentation for details.
Tool Calling
Give your LLM the ability to interact with the outside world by turning any Python function into a tool with the @tool decorator. NobodyWho inspects the function signature to derive the parameters and configures the sampler for you:
import math
from nobodywho import Chat, tool
@tool(description="Calculates the area of a circle given its radius")
def circle_area(radius: float) -> str:
area = math.pi * radius ** 2
return f"Circle with radius {radius} has area {area:.2f}"
chat = Chat("./model.gguf", tools=[circle_area])
response = chat.ask("What is the area of a circle with a radius of 2?").completed()
print(response)
See the Tool Calling documentation for more.
Sampling
A sampler decides how the next token is picked from the model's probability distribution. Use a preset to tune creativity, or constrain the output to a specific format such as JSON:
from nobodywho import Chat, SamplerPresets
# Lower temperature = more deterministic output
chat = Chat("./model.gguf", sampler=SamplerPresets.temperature(0.2))
You can also force the output to match a JSON schema, a regex, or a custom grammar:
import json
from nobodywho import Chat, SamplerPresets
chat = Chat("./model.gguf", sampler=SamplerPresets.constrain_with_json_schema({
"type": "object",
"properties": {
"name": {"type": "string", "maxLength": 50},
"age": {"type": "integer"},
},
"required": ["name", "age"],
"additionalProperties": False,
}))
person = json.loads(chat.ask("Give me a person as JSON with name and age.").completed())
See the Sampling documentation for more.
Embeddings & RAG
For semantic search, document similarity, or retrieval-augmented generation (RAG), NobodyWho supports embeddings and cross-encoders.
Turn text into vectors with an Encoder and compare them with cosine_similarity:
from nobodywho import Encoder, cosine_similarity
encoder = Encoder("./embedding-model.gguf")
query = encoder.encode("How do I reset my password?")
doc = encoder.encode("You can reset your password in the account settings")
print(cosine_similarity(query, doc))
For more accurate ranking, use a CrossEncoder to build a knowledge-base search tool:
from nobodywho import Chat, CrossEncoder, tool
crossencoder = CrossEncoder("./reranker-model.gguf")
knowledge = [
"Our company offers a 30-day return policy for all products",
"Free shipping is available on orders over $50",
"Customer support is available via email and phone",
]
@tool(description="Search the knowledge base for relevant information")
def search_knowledge(query: str) -> str:
ranked = crossencoder.rank_and_sort(query, knowledge)
return "\n".join(doc for doc, score in ranked[:3])
chat = Chat(
"./model.gguf",
system_prompt="Use the search_knowledge tool to answer customer questions.",
tools=[search_knowledge],
)
print(chat.ask("What is your return policy?").completed())
See the Embeddings & RAG documentation for more.
Vision and Hearing
Include images and audio in your prompts, so the model can see and hear content alongside text. You need a multimodal LLM plus a matching projection model (usually named mmproj), which have to be trained together.
from nobodywho import Model, Chat, Prompt, Text, Image, Audio
model = Model("./multimodal-model.gguf", projection_model_path="./mmproj.gguf")
chat = Chat(model, system_prompt="You are a helpful assistant that can hear and see!")
prompt = Prompt([
Text("Tell me what you see in the image and what you hear in the audio."),
Image("./dog.png"),
Audio("./sound.mp3"),
])
print(chat.ask(prompt).completed())
See the Multimodal documentation for model recommendations and advanced tips.
Speech to Text
Transcribe spoken audio into text using Whisper models in ONNX format:
from nobodywho import SpeechToText
stt = SpeechToText(source="hf://onnx-community/whisper-base")
text = stt.transcribe_file("recording.mp3").completed()
print(text)
You can also transcribe raw mono i16 PCM buffers with transcribe_pcm, and stream the transcription token by token.
See the Speech to Text documentation for more.
Text to Speech
Generate natural-sounding speech from text, ready to save as a WAV file or play back in your app:
from pathlib import Path
from nobodywho import TextToSpeech
tts = TextToSpeech(
source="hf://NobodyWho/Kokoro-82M", # Hugging Face repo or local folder.
voice="bf_emma", # Voice to use from the model.
language="en-gb", # Language code for the input text.
)
wav = tts.synthesize("Hello from NobodyWho!")
Path("out.wav").write_bytes(wav)
NobodyWho supports the Kokoro, Pocket TTS, and Supertonic speech synthesis architectures.
See the Text to Speech documentation for more.
Voice Activity Detection
Detect speech automatically in an audio stream, so you know when to stop listening to the microphone and start transcribing:
from nobodywho import VoiceActivityDetection, VoiceActivityDetectionEvent, SpeechToText
vad = VoiceActivityDetection(source="hf://onnx-community/silero-vad", sample_rate=16000)
stt = SpeechToText(source="hf://onnx-community/whisper-base")
while chunk := read_mic(): # however you read from the microphone
if vad.push(chunk) == VoiceActivityDetectionEvent.SpeechEnded:
break
speech = vad.finish() # buffered speech, ready to pass to SpeechToText
print(stt.transcribe_pcm(speech, sample_rate=16000).completed())
You can also segment speech out of an existing recording with segment().
See the Voice Activity Detection documentation for more.
Streaming & Async API
To stream tokens as soon as they arrive, iterate over the response instead of calling .completed():
from nobodywho import Chat
chat = Chat("./model.gguf")
for token in chat.ask("How are you?"):
print(token, end="", flush=True)
For non-blocking inference, swap Chat for ChatAsync — the API is identical, and you can await a full response or stream tokens with async for:
import asyncio
from nobodywho import ChatAsync
async def main():
chat = ChatAsync("./model.gguf")
async for token in chat.ask("How are you?"):
print(token, end="", flush=True)
asyncio.run(main())
The other model types also have async variants: EncoderAsync, CrossEncoderAsync, and SpeechToTextAsync.
See the Streaming & Async documentation for more.
Documentation
Full documentation available at: https://docs.nobodywho.ooo/python/
License
EUPL-1.2 - Free for commercial and proprietary use. Modified versions of the library itself must remain open source.
Metadata
Release files for nobodywho 3.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Built distributions (wheels)
| File | Reset | |||
|---|---|---|---|---|
| nobodywho-3.0.0-cp39-abi3-win_amd64.whl | CPython 3.9 | abi3 | Windows x86-64 | Details |
| nobodywho-3.0.0-cp39-abi3-manylinux_2_34_x86_64.whl | CPython 3.9 | abi3 | Linux glibc 2.34+ x86-64 | Details |
| nobodywho-3.0.0-cp39-abi3-manylinux_2_34_aarch64.whl | CPython 3.9 | abi3 | Linux glibc 2.34+ ARM64 | Details |
| nobodywho-3.0.0-cp39-abi3-macosx_11_0_arm64.whl | CPython 3.9 | abi3 | macOS 11.0+ ARM64 | Details |
Total release size: 156.7 MB
Release files / nobodywho-3.0.0-cp39-abi3-win_amd64.whl
| Download URL | nobodywho-3.0.0-cp39-abi3-win_amd64.whl |
|---|---|
| Size | 43.1 MB |
| Tags | CPython 3.9 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
e58ccef9e2cd574f4b3b883fb5842afbefe6e79bbbf6db8f4341896f5032cd80
|
|
BLAKE2b-256 checksum How to use checksums |
cb0f19b0fb8614192bfa9bf8842135e874fb59814d426d323d6590e6ed31be65
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|
Release files / nobodywho-3.0.0-cp39-abi3-manylinux_2_34_x86_64.whl
| Download URL | nobodywho-3.0.0-cp39-abi3-manylinux_2_34_x86_64.whl |
|---|---|
| Size | 43.4 MB |
| Tags | CPython 3.9 Linux glibc 2.34+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
1755cfc11fd6a9a0161301c96ce9989fc8cd3c765f6409530571d94913e2c336
|
|
BLAKE2b-256 checksum How to use checksums |
5786ee70b6325dd9232faca6ec112a17b9ce61fbec3bc8a98ef6050a226d9228
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|
Release files / nobodywho-3.0.0-cp39-abi3-manylinux_2_34_aarch64.whl
| Download URL | nobodywho-3.0.0-cp39-abi3-manylinux_2_34_aarch64.whl |
|---|---|
| Size | 41.5 MB |
| Tags | CPython 3.9 Linux glibc 2.34+ ARM64 abi3 |
|
SHA-256 checksum How to use checksums |
1138ecbaa2d21365c0a6171c2fbcd2ad849c0dc5e3b1c65bce076722c0fcf3a0
|
|
BLAKE2b-256 checksum How to use checksums |
645b48810990b77f7e3d84fbc3545329ea8de5766bfb10611f12ff15b4ada502
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|
Release files / nobodywho-3.0.0-cp39-abi3-macosx_11_0_arm64.whl
| Download URL | nobodywho-3.0.0-cp39-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 28.8 MB |
| Tags | CPython 3.9 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
dd42944219ddd1393cd4d7b6f7754168d5e0ccacb5e94998b225cabdb45b28c1
|
|
BLAKE2b-256 checksum How to use checksums |
b39022c760a4b93ceb6c4cfff1d7e9a270a58abf5e85e565d02a2f21e7e12f4f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.15
|