Skip to main content

fermion-research

The fermion command runs the Phonon speech models and the Neutrino language models from one install. It transcribes audio files, transcribes the microphone live on Apple silicon, and serves an OpenAI-compatible endpoint. Phonon-2 is the current speech model. It runs on Apple silicon through MLX, on CPUs under Linux (x86-64 and Arm), Windows and macOS (Apple silicon and Intel Macs), and on NVIDIA GPUs through the CUDA container.

Models

Model Download Weights
Phonon-2 (phonon-2) 164 MB FermionResearch/Phonon-2
Phonon-1 (phonon-1) 415 MB FermionResearch/Phonon-1
Phonon-1 Micro (phonon-1-micro) 285 MB FermionResearch/Phonon-1-Micro
Phonon-1 Big (phonon-1-big) 581 MB FermionResearch/Phonon-1-Big

Accuracy and speed for each model are on its model page (fermionresearch.com/models/phonon-2 and the model cards above).

Install

Python 3.10 or newer (macOS's built-in python3 is 3.9). On a Mac, create a virtual environment with a newer Python first.

python3.12 -m venv .venv && source .venv/bin/activate
pip install fermion-research

On Linux, Windows and Intel Mac CPUs, that line is the whole setup: Phonon-2 runs on the CPU engine. On Linux, install torch from its CPU wheel index first, which skips the GPU build.

pip install --no-deps torch --index-url https://download.pytorch.org/whl/cpu   # Linux only
pip install fermion-research

On Apple silicon, add the MLX speech runtime.

pip install "fermion-research[mlx]"

Intel Macs: Phonon-2 runs on the CPU engine (Python 3.10 to 3.12; the install resolves the last Intel torch wheel, 2.2.2, with transformers 5.0).

Supported CPUs: x86-64 with SSE4.1 or newer (AVX2, AVX-512 VNNI and AMX processors run faster tiers of the same kernels) and 64-bit Arm with NEON (dotprod and i8mm processors run faster tiers), on Linux, Windows and Intel Macs; Apple silicon Macs run the MLX engine. fermion describe shows the features found on your machine and the tier it runs.

The CPU container runs on amd64 and arm64.

docker run --rm -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cpu:2.0.8 transcribe phonon-2 /audio/recording.wav

The CUDA container runs Phonon-2 on NVIDIA GPUs.

docker run --rm --gpus all -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cuda:1.0.7 transcribe phonon-2 /audio/recording.wav

Run

Transcribe a file, transcribe the microphone live, or start a local OpenAI-compatible server. phonon runs Phonon-2, phonon-1 runs Phonon-1, and fermion <command> <model> runs any model by name. Name the model. Phonon never guesses.

phonon transcribe meeting.wav               # transcribe a file with Phonon-2
phonon transcribe meeting.wav --hotwords "Kushal, Sigil"   # favour names and terms
phonon listen                               # live microphone transcription (Apple silicon)
phonon serve                                # OpenAI-compatible HTTP server on 127.0.0.1:8000
fermion transcribe phonon-2 meeting.wav     # the same, naming the model

Any OpenAI-compatible client can then send audio to the server.

curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
  -F "file=@meeting.wav" \
  -F "model=phonon-2"

Streaming

phonon serve also transcribes live audio over a WebSocket at /v1/audio/stream. The protocol and a reference client are in docs/server.md.

Dictation tools

A dictation tool keeps one Phonon server running and sends each recording to it, over a local port or an owner-only Unix socket.

fermion serve phonon-2 --port 8010 --threads 4         # or: phonon serve --port 8010
fermion serve phonon-2 --unix-socket ~/.cache/fermion/phonon.sock   # owner-only socket in place of an API key
engine = "whisper"
[whisper]
backend = "remote"
remote_endpoint = "http://127.0.0.1:8010"
remote_model = "phonon-2"

Keeping the server running at login is covered in docs/server.md.

fermion models lists every model with its aliases and marks the ones already on the machine.

Neutrino

The same install runs the Neutrino language models. fermion chat opens a conversation, fermion generate completes a prompt, and fermion serve starts an OpenAI-compatible chat endpoint, each with the model named.

fermion chat neutrino-8b                     # interactive
fermion generate neutrino-8b 'Write a haiku'
fermion serve neutrino-8b                    # OpenAI-compatible chat completions

Documentation

Licence

The Phonon-2 weights are released under CC-BY-4.0. They are a derivative of NVIDIA's parakeet-tdt-0.6b-v3, with the changes listed in the NOTICE file of the weights repository. The Phonon-1 family weights and the command line are released under Apache-2.0, and Phonon-1 is built on Qwen/Qwen3-ASR-0.6B, also under Apache-2.0.

Metadata

Release files for fermion-research 0.2.10

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fermion-research 0.2.10
File Size Uploaded
fermion_research-0.2.10.tar.gz 1.3 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for fermion-research 0.2.10
File Interpreter ABI Platform
fermion_research-0.2.10-py3-none-any.whl Python 3 none any Details

Total release size: 2.7 MB

Release files / fermion_research-0.2.10.tar.gz

Download URL fermion_research-0.2.10.tar.gz
Size 1.3 MB
Tags Source
SHA-256 checksum
How to use checksums
6225bed009880091ddecf545fe17c17eb7358fc85acb2bf2b1b1077563e4cb3c
BLAKE2b-256 checksum
How to use checksums
f75031f7c992e4e582940adf8f642eefb71e554b81bb5a619f2e012d4380e2dd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / fermion_research-0.2.10-py3-none-any.whl

Download URL fermion_research-0.2.10-py3-none-any.whl
Size 1.4 MB
Tags Python 3
SHA-256 checksum
How to use checksums
112c3b99851962782c7306f5343a5e764a2e985d91d19c4ce3ddbabf4ca6bc69
BLAKE2b-256 checksum
How to use checksums
87ae2bf70a140ea6ce9a124fc5c2332f7c079a1410ec5ecfb94a87c923853a02
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.2.10 This release

2 release files

0.2.9

2 release files

0.2.7

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.22

2 release files

0.1.21

2 release files

0.1.20

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

0.1.16

2 release files

0.1.15

2 release files

0.1.14

2 release files

0.1.13

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

1 release file

0.1.6

1 release file

0.1.5

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page