fermion-research
The fermion command runs the Phonon speech models and the Neutrino language models from one install. It transcribes audio files, transcribes the microphone live on Apple silicon, and serves an OpenAI-compatible endpoint. Phonon-2 is the current speech model. It runs on Apple silicon through MLX, on CPUs under Linux (x86-64 and Arm), Windows and macOS (Apple silicon and Intel Macs), and on NVIDIA GPUs through the CUDA container.
Models
| Model | Download | Weights |
|---|---|---|
Phonon-2 (phonon-2) |
164 MB | FermionResearch/Phonon-2 |
Phonon-1 (phonon-1) |
415 MB | FermionResearch/Phonon-1 |
Phonon-1 Micro (phonon-1-micro) |
285 MB | FermionResearch/Phonon-1-Micro |
Phonon-1 Big (phonon-1-big) |
581 MB | FermionResearch/Phonon-1-Big |
Accuracy and speed for each model are on its model page (fermionresearch.com/models/phonon-2 and the model cards above).
Install
The package installs from PyPI.
pip install fermion-research
On Apple silicon, one more line adds the MLX speech runtime.
pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
On Linux, Windows and Intel Mac CPUs, the package runs with torch. On Linux, install torch from its CPU wheel index first, which skips the GPU build.
pip install --no-deps torch --index-url https://download.pytorch.org/whl/cpu # Linux only
pip install fermion-research torch safetensors soundfile scipy zstandard
Intel Macs: Phonon-2 runs on the CPU engine (Python 3.10 to 3.12; the install line above resolves the last Intel torch wheel, 2.2.2, with transformers 5.0).
Supported CPUs: x86-64 with SSE4.1 or newer (AVX2, AVX-512 VNNI and AMX processors run faster tiers of the same kernels) and 64-bit Arm with NEON (dotprod and i8mm processors run faster tiers), on Linux, Windows and Intel Macs; Apple silicon Macs run the MLX engine. fermion describe shows the features found on your machine and the tier it runs.
The CPU container runs on amd64 and arm64.
docker run --rm -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cpu:2.0.3 transcribe phonon-2 /audio/recording.wav
The CUDA container runs Phonon-2 on NVIDIA GPUs.
docker run --rm --gpus all -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cuda:1.0.4 transcribe phonon-2 /audio/recording.wav
Run
Transcribe a file, transcribe the microphone live, or start a local OpenAI-compatible server. phonon runs Phonon-2, phonon-1 runs Phonon-1, and fermion <command> <model> runs any model by name. Name the model. Phonon never guesses.
phonon transcribe meeting.wav # transcribe a file with Phonon-2
phonon transcribe meeting.wav --hotwords "Kushal, Sigil" # favour names and terms
phonon listen # live microphone transcription (Apple silicon)
phonon serve # OpenAI-compatible HTTP server on 127.0.0.1:8000
fermion transcribe phonon-2 meeting.wav # the same, naming the model
Any OpenAI-compatible client can then send audio to the server.
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
-F "file=@meeting.wav" \
-F "model=phonon-2"
Dictation tools
A dictation tool keeps one Phonon server running and sends each recording to it, over a local port or an owner-only Unix socket.
fermion serve phonon-2 --port 8010 --threads 4 # or: phonon serve --port 8010
fermion serve phonon-2 --unix-socket ~/.cache/fermion/phonon.sock # owner-only socket in place of an API key
engine = "whisper"
[whisper]
backend = "remote"
remote_endpoint = "http://127.0.0.1:8010"
remote_model = "phonon-2"
Keeping the server running at login is covered in docs/server.md.
fermion models lists every model with its aliases and marks the ones already on the machine.
Neutrino
The same install runs the Neutrino language models. fermion chat opens a conversation, fermion generate completes a prompt, and fermion serve starts an OpenAI-compatible chat endpoint, each with the model named.
fermion chat neutrino-8b # interactive
fermion generate neutrino-8b 'Write a haiku'
fermion serve neutrino-8b # OpenAI-compatible chat completions
Documentation
- The command line
- The HTTP server and its API
- Installation on every platform
- Running on CPUs, no GPU required
- The NVIDIA CUDA image
- Fixes for common problems
Licence
The Phonon-2 weights are released under CC-BY-4.0. They are a derivative of NVIDIA's parakeet-tdt-0.6b-v3, with the changes listed in the NOTICE file of the weights repository. The Phonon-1 family weights and the command line are released under Apache-2.0, and Phonon-1 is built on Qwen/Qwen3-ASR-0.6B, also under Apache-2.0.
Metadata
Release files for fermion-research 0.2.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| fermion_research-0.2.9.tar.gz | 1.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| fermion_research-0.2.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.5 MB
Release files / fermion_research-0.2.9.tar.gz
| Download URL | fermion_research-0.2.9.tar.gz |
|---|---|
| Size | 1.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0e19ed285dd87be804e29e617f7f62d7d6a7ecddb79ef28a46fa407fe3949cf7
|
|
BLAKE2b-256 checksum How to use checksums |
9f0488421fe19d2de0ccb987a69de618e1387ffa647c463b2e3617c39911edaa
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|
Release files / fermion_research-0.2.9-py3-none-any.whl
| Download URL | fermion_research-0.2.9-py3-none-any.whl |
|---|---|
| Size | 1.3 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
68bd35b5f2ad397ab0202752ef3ea0e95dad64dc019af0f3baebdf5f95fe308e
|
|
BLAKE2b-256 checksum How to use checksums |
0a32bc70c7911d1ff1719e3e4a0db95d3bea5228e8249edd2b02dad6bc97fbb4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.13
|