Skip to main content

NobodyWho

Run LLMs locally and efficiently on any device

NobodyWho is a lightweight, open-source inference engine that makes it simple to run open-weights language models directly inside your Python applications. No API keys, no cloud infrastructure, no complexity—just fast, easy local AI.

Free to use in commercial projects under the EUPL-1.2 license, no API key required. Supports text, vision, hearing, speech-to-text, text-to-speech, voice activity detection, embeddings, RAG & tool calling.

Key Features

  • Run locally, offline, for free - No API keys or cloud services required
  • Fast, simple tool calling - Just pass normal Python functions
  • Reliable tool execution - Automatically derives grammar from function signatures
  • Speech-to-text - Transcribe spoken audio into text with Whisper models
  • Text-to-speech - Generate natural-sounding speech from text
  • Voice activity detection - Detect speech in an audio stream to know when to start and stop listening
  • Vision & embeddings - Multimodal image and audio input, plus embeddings and reranking for semantic search and RAG
  • Infinite conversations - Conversation-aware preemptive context shifting prevents mid-conversation crashes
  • GPU accelerated - Vulkan-powered inference for maximum performance
  • Thousands of compatible models - Works with any LLM in GGUF format
  • Powered by llama.cpp - Built on the proven llama.cpp engine

Installation

pip install nobodywho

Supported Model Format

NobodyWho uses the GGUF format, a binary format optimized for fast loading and efficient LLM inference. A wide selection of GGUF models is available on Hugging Face.

You can also download a model without any extra dependencies by passing huggingface:owner/repo/filename.gguf where you'd normally pass the model path:

from nobodywho import Chat

chat = Chat("huggingface:NobodyWho/Qwen_Qwen3-0.6B-GGUF/Qwen_Qwen3-0.6B-Q4_K_M.gguf")

Chat

Every interaction with your LLM starts by instantiating a Chat object. Call .ask() to send a message, and .completed() to block until the whole response is ready:

from nobodywho import Chat

chat = Chat("./model.gguf", system_prompt="You are a helpful assistant.")
response = chat.ask("Is water wet?").completed()
print(response)  # Yes, indeed, water is wet!

Your messages and the model's responses are remembered inside the Chat object, so follow-up questions keep their context. You can pass "auto" as the model path to pick a chat model based on available memory.

See the Chat documentation for details.

Tool Calling

Give your LLM the ability to interact with the outside world by turning any Python function into a tool with the @tool decorator. NobodyWho inspects the function signature to derive the parameters and configures the sampler for you:

import math
from nobodywho import Chat, tool

@tool(description="Calculates the area of a circle given its radius")
def circle_area(radius: float) -> str:
    area = math.pi * radius ** 2
    return f"Circle with radius {radius} has area {area:.2f}"

chat = Chat("./model.gguf", tools=[circle_area])
response = chat.ask("What is the area of a circle with a radius of 2?").completed()
print(response)

See the Tool Calling documentation for more.

Sampling

A sampler decides how the next token is picked from the model's probability distribution. Use a preset to tune creativity, or constrain the output to a specific format such as JSON:

from nobodywho import Chat, SamplerPresets

# Lower temperature = more deterministic output
chat = Chat("./model.gguf", sampler=SamplerPresets.temperature(0.2))

You can also force the output to match a JSON schema, a regex, or a custom grammar:

import json
from nobodywho import Chat, SamplerPresets

chat = Chat("./model.gguf", sampler=SamplerPresets.constrain_with_json_schema({
    "type": "object",
    "properties": {
        "name": {"type": "string", "maxLength": 50},
        "age":  {"type": "integer"},
    },
    "required": ["name", "age"],
    "additionalProperties": False,
}))
person = json.loads(chat.ask("Give me a person as JSON with name and age.").completed())

See the Sampling documentation for more.

Embeddings & RAG

For semantic search, document similarity, or retrieval-augmented generation (RAG), NobodyWho supports embeddings and cross-encoders.

Turn text into vectors with an Encoder and compare them with cosine_similarity:

from nobodywho import Encoder, cosine_similarity

encoder = Encoder("./embedding-model.gguf")
query = encoder.encode("How do I reset my password?")
doc = encoder.encode("You can reset your password in the account settings")
print(cosine_similarity(query, doc))

For more accurate ranking, use a CrossEncoder to build a knowledge-base search tool:

from nobodywho import Chat, CrossEncoder, tool

crossencoder = CrossEncoder("./reranker-model.gguf")
knowledge = [
    "Our company offers a 30-day return policy for all products",
    "Free shipping is available on orders over $50",
    "Customer support is available via email and phone",
]

@tool(description="Search the knowledge base for relevant information")
def search_knowledge(query: str) -> str:
    ranked = crossencoder.rank_and_sort(query, knowledge)
    return "\n".join(doc for doc, score in ranked[:3])

chat = Chat(
    "./model.gguf",
    system_prompt="Use the search_knowledge tool to answer customer questions.",
    tools=[search_knowledge],
)
print(chat.ask("What is your return policy?").completed())

See the Embeddings & RAG documentation for more.

Vision and Hearing

Include images and audio in your prompts, so the model can see and hear content alongside text. You need a multimodal LLM plus a matching projection model (usually named mmproj), which have to be trained together.

from nobodywho import Model, Chat, Prompt, Text, Image, Audio

model = Model("./multimodal-model.gguf", projection_model_path="./mmproj.gguf")
chat = Chat(model, system_prompt="You are a helpful assistant that can hear and see!")

prompt = Prompt([
    Text("Tell me what you see in the image and what you hear in the audio."),
    Image("./dog.png"),
    Audio("./sound.mp3"),
])
print(chat.ask(prompt).completed())

See the Multimodal documentation for model recommendations and advanced tips.

Speech to Text

Transcribe spoken audio into text using Whisper models in ONNX format:

from nobodywho import SpeechToText

stt = SpeechToText(source="hf://onnx-community/whisper-base")
text = stt.transcribe_file("recording.mp3").completed()
print(text)

You can also transcribe raw mono i16 PCM buffers with transcribe_pcm, and stream the transcription token by token.

See the Speech to Text documentation for more.

Text to Speech

Generate natural-sounding speech from text, ready to save as a WAV file or play back in your app:

from pathlib import Path
from nobodywho import TextToSpeech

tts = TextToSpeech(
    source="hf://NobodyWho/Kokoro-82M",  # Hugging Face repo or local folder.
    voice="bf_emma",                     # Voice to use from the model.
    language="en-gb",                    # Language code for the input text.
)

wav = tts.synthesize("Hello from NobodyWho!")
Path("out.wav").write_bytes(wav)

NobodyWho supports the Kokoro, Pocket TTS, and Supertonic speech synthesis architectures.

See the Text to Speech documentation for more.

Voice Activity Detection

Detect speech automatically in an audio stream, so you know when to stop listening to the microphone and start transcribing:

from nobodywho import VoiceActivityDetection, VoiceActivityDetectionEvent, SpeechToText

vad = VoiceActivityDetection(source="hf://onnx-community/silero-vad", sample_rate=16000)
stt = SpeechToText(source="hf://onnx-community/whisper-base")

while chunk := read_mic():  # however you read from the microphone
    if vad.push(chunk) == VoiceActivityDetectionEvent.SpeechEnded:
        break

speech = vad.finish()  # buffered speech, ready to pass to SpeechToText
print(stt.transcribe_pcm(speech, sample_rate=16000).completed())

You can also segment speech out of an existing recording with segment().

See the Voice Activity Detection documentation for more.

Streaming & Async API

To stream tokens as soon as they arrive, iterate over the response instead of calling .completed():

from nobodywho import Chat

chat = Chat("./model.gguf")
for token in chat.ask("How are you?"):
    print(token, end="", flush=True)

For non-blocking inference, swap Chat for ChatAsync — the API is identical, and you can await a full response or stream tokens with async for:

import asyncio
from nobodywho import ChatAsync

async def main():
    chat = ChatAsync("./model.gguf")
    async for token in chat.ask("How are you?"):
        print(token, end="", flush=True)

asyncio.run(main())

The other model types also have async variants: EncoderAsync, CrossEncoderAsync, and SpeechToTextAsync.

See the Streaming & Async documentation for more.

Documentation

Full documentation available at: https://docs.nobodywho.ooo/python/

License

EUPL-1.2 - Free for commercial and proprietary use. Modified versions of the library itself must remain open source.

Metadata

Release files for nobodywho 3.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for nobodywho 3.0.0
File
nobodywho-3.0.0-cp39-abi3-win_amd64.whl CPython 3.9 abi3 Windows x86-64 Details
nobodywho-3.0.0-cp39-abi3-manylinux_2_34_x86_64.whl CPython 3.9 abi3 Linux glibc 2.34+ x86-64 Details
nobodywho-3.0.0-cp39-abi3-manylinux_2_34_aarch64.whl CPython 3.9 abi3 Linux glibc 2.34+ ARM64 Details
nobodywho-3.0.0-cp39-abi3-macosx_11_0_arm64.whl CPython 3.9 abi3 macOS 11.0+ ARM64 Details

Total release size: 156.7 MB

Release files / nobodywho-3.0.0-cp39-abi3-win_amd64.whl

Download URL nobodywho-3.0.0-cp39-abi3-win_amd64.whl
Size 43.1 MB
Tags CPython 3.9 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
e58ccef9e2cd574f4b3b883fb5842afbefe6e79bbbf6db8f4341896f5032cd80
BLAKE2b-256 checksum
How to use checksums
cb0f19b0fb8614192bfa9bf8842135e874fb59814d426d323d6590e6ed31be65
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / nobodywho-3.0.0-cp39-abi3-manylinux_2_34_x86_64.whl

Download URL nobodywho-3.0.0-cp39-abi3-manylinux_2_34_x86_64.whl
Size 43.4 MB
Tags CPython 3.9 Linux glibc 2.34+ x86-64 abi3
SHA-256 checksum
How to use checksums
1755cfc11fd6a9a0161301c96ce9989fc8cd3c765f6409530571d94913e2c336
BLAKE2b-256 checksum
How to use checksums
5786ee70b6325dd9232faca6ec112a17b9ce61fbec3bc8a98ef6050a226d9228
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / nobodywho-3.0.0-cp39-abi3-manylinux_2_34_aarch64.whl

Download URL nobodywho-3.0.0-cp39-abi3-manylinux_2_34_aarch64.whl
Size 41.5 MB
Tags CPython 3.9 Linux glibc 2.34+ ARM64 abi3
SHA-256 checksum
How to use checksums
1138ecbaa2d21365c0a6171c2fbcd2ad849c0dc5e3b1c65bce076722c0fcf3a0
BLAKE2b-256 checksum
How to use checksums
645b48810990b77f7e3d84fbc3545329ea8de5766bfb10611f12ff15b4ada502
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / nobodywho-3.0.0-cp39-abi3-macosx_11_0_arm64.whl

Download URL nobodywho-3.0.0-cp39-abi3-macosx_11_0_arm64.whl
Size 28.8 MB
Tags CPython 3.9 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
dd42944219ddd1393cd4d7b6f7754168d5e0ccacb5e94998b225cabdb45b28c1
BLAKE2b-256 checksum
How to use checksums
b39022c760a4b93ceb6c4cfff1d7e9a270a58abf5e85e565d02a2f21e7e12f4f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

This release

3.0.0 This release

4 release files

2.0.0

4 release files

1.7.0

4 release files

1.6.0

4 release files

1.5.0

4 release files

1.4.0

4 release files

1.3.0

4 release files

1.2.0

4 release files

1.1.0

4 release files

1.0.0

4 release files

0.9.1

4 release files

0.9.0

4 release files

0.8.0

4 release files

0.7.0

4 release files

0.6.2

4 release files

0.6.1

4 release files

0.6.0

4 release files

0.5.0

4 release files

0.4.2

4 release files

0.4.1

4 release files

0.4.0

3 release files

0.3.0

3 release files

0.2.0

3 release files

0.1.0

3 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page