Skip to main content

speech2caret

speech2caret logo

Use your speech to write to the current caret position!

Goals

  • ✅ Simple: A minimalist tool that does one thing well.
  • ✅ Local: Runs entirely on your machine (uses Hugging Face models for speech recognition).
  • ✅ Efficient: Optimised for low CPU and memory usage, thanks to an event-driven architecture that responds instantly to key presses without wasting resources.

Note: Tested only on Linux (Ubuntu). Other operating systems are currently unsupported.

Demo (turn volume on):

demo video

Installation

1. System Dependencies

First, install the required system libraries:

sudo apt update
sudo apt install libportaudio2 ffmpeg

2. Grant Permissions

To read keyboard events and simulate key presses, evdev needs access to your keyboard input device. Add your user to the input group to grant the necessary permissions:

sudo usermod -aG input $USER
newgrp input  # or log out and back in 

3. Install and Run

You can install and run speech2caret using pip or uv:

# Install the package
uv add speech2caret  # or pip install speech2caret

# Run the application
speech2caret

Alternatively, you can run it directly without installation using uvx(the --index pytorch-cpu=... flag ensures only CPU packages are downloaded, avoiding GPU-related dependencies):

uvx --index pytorch-cpu=https://download.pytorch.org/whl/cpu --from speech2caret speech2caret

Configuration

The first time you run speech2caret, it creates a config file at ~/.config/speech2caret/config.ini.

You’ll need to manually edit it with the following values:

keyboard_device_path

This is the path to your keyboard input device. You can find the path either following this, or by running the command below and looking for an entry that ends with -event-kbd.

ls /dev/input/by-path/

start_stop_key and resume_pause_key

These are the keys you'll use to control the app.

To find the correct name for a key, you can use the provided Python script below. First, ensure you have your keyboard_device_path from the step above, then run this command:

uvx --from evdev python -c '
keyboard_device_path = "PASTE_YOUR_KEYBOARD_DEVICE_PATH_HERE"

from evdev import InputDevice, categorize, ecodes, KeyEvent
dev = InputDevice(keyboard_device_path)
print(f"Listening for key presses on {dev.name}...")
for event in dev.read_loop():
    if event.type == ecodes.EV_KEY:
        key_event = categorize(event)
        if key_event.keystate == KeyEvent.key_down:
            print(f" {key_event.keycode}")
'

Press the keys you wish to use, and their names will be printed to the terminal. For a full list of available key names, see here.

Additional (Optional) Configuration

You can configure audio cues to notify when recording has started, stopped, paused, or resumed. To do this, update the start_recording_audio_path, stop_recording_audio_path, resume_recording_audio_path, and pause_recording_audio_path config variables in ~/.config/speech2caret/config.ini with the absolute paths to your choice of audio files.

Word Replacement

You can define custom word or phrase replacements in the [word_replacement] section of ~/.config/speech2caret/config.ini file. This allows you to automatically substitute specific spoken words with desired text.

For example, to replace "new line" with a newline character or " underscore " with _, you can configure it as follows:

[word_replacement]
"new line" = "\n"
" underscore " = "_"

How to Use

  1. Run the speech2caret command in your terminal.
  2. Press your configured start_stop_key to begin recording.
  3. Press the resume_pause_key to toggle between pausing and resuming.
  4. When you are finished, press the start_stop_key again.
  5. The recorded audio will be transcribed and typed at your current caret position.

Metadata

Release files for speech2caret 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for speech2caret 0.3.0
File Size Uploaded
speech2caret-0.3.0.tar.gz 7.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for speech2caret 0.3.0
File Interpreter ABI Platform
speech2caret-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 17.1 kB

Release files / speech2caret-0.3.0.tar.gz

Download URL speech2caret-0.3.0.tar.gz
Size 7.3 kB
Tags Source
SHA-256 checksum
How to use checksums
92457870e7d26c4e99fc38ae097a380df32902a658333d84e1dc6d8a754fa021
BLAKE2b-256 checksum
How to use checksums
4626f365493f6177587880d1ffe7d5c74fd2c6f1cf917acc5ccb10bfd51c2a83
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Nov 14, 2025.

Transparency log

Release files / speech2caret-0.3.0-py3-none-any.whl

Download URL speech2caret-0.3.0-py3-none-any.whl
Size 9.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
e797f364bb04a491a27bae32f26efa2f54fa986d44416d7c694f033058768224
BLAKE2b-256 checksum
How to use checksums
fef8f285e31edcbf0365337da4cffdcffe7f729133c701fe17637af09bbaa810
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Nov 14, 2025.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page