Skip to main content

pipecat-hecttor

Hecttor speech enhancement for Pipecat.

Two integrations are provided:

  • HecttorFilter — an input audio filter that removes background noise from the user's audio in real time using the Hecttor SDK's ASR-optimized speech enhancer. It runs before audio reaches the STT service, improving transcription accuracy in noisy environments. You can choose among several enhancement models and blend the enhanced output with the original audio.
  • HecttorAudioProcessor — a frame processor that enhances the input audio with two different blend factors: one for the STT/agent path and one for VAD and turn-taking (TT) models, since the optimal blend for transcription is usually not the optimal blend for endpointing.

This integration is maintained by Saima AI, the company behind Hecttor.

Installation

Install the package:

uv add pipecat-hecttor

The filter requires the hecttor_sdk Python package, which is not published to PyPI. Contact Hecttor for SDK access and an API key — you'll receive a wheel for your platform and Python version:

pip install hecttor_sdk-<version>-<python>-<platform>.whl

Set your API key:

export HECTTOR_API_KEY=your_api_key_here

Usage

Add the filter to any Pipecat transport via the audio_in_filter parameter:

from pipecat_hecttor import HecttorFilter
from pipecat.transports.base_transport import TransportParams

params = TransportParams(
    audio_in_enabled=True,
    audio_in_filter=HecttorFilter(),
    audio_out_enabled=True,
)

Configuration

Parameter Default Description
api_key None Hecttor API key. Falls back to the HECTTOR_API_KEY environment variable.
model_name "coda-vi-1.0" ASR enhancement model: crest-1.0, crest-2.0, mist-1.0, coda-1.0, or coda-vi-1.0.
chunk_size_ms 20 Chunk size in milliseconds, 16 or 20. crest-2.0, coda-1.0, and coda-vi-1.0 require 20.
enhancer_weight None Blend factor [0.0, 1.0] between original (0.0) and enhanced (1.0) audio. None uses the model's default.
smart_blending False Gated "smart" wet/dry blend (requires hecttor_sdk >= 3.3.0): the blend is applied in the time domain and gated by a noise gate driven by the enhanced signal, so lower weights restore the original voice without the background noise in pauses. All models except coda-1.0.

Per-consumer blends: HecttorAudioProcessor

When you want STT and the VAD/turn-taking models to hear differently blended audio, use HecttorAudioProcessor instead of the transport filter (never both at once). It exploits Pipecat's pipeline ordering: STT consumes audio before the user context aggregator, which hosts the VAD and turn analyzers. Stage 1 sits after the transport input and rewrites frames with the ASR blend; stage 2 (vad_tt_stage()) sits after STT and swaps in the VAD/TT blend:

from pipecat_hecttor import HecttorAudioProcessor

hecttor = HecttorAudioProcessor(asr_weight=1.0, vad_tt_weight=0.5)

pipeline = Pipeline(
    [
        transport.input(),       # no audio_in_filter
        hecttor,                 # stage 1: frames now carry the ASR blend
        stt,                     # hears the ASR blend
        hecttor.vad_tt_stage(),  # stage 2: swaps in the VAD/TT blend
        user_aggregator,         # VAD + turn analyzers hear the VAD/TT blend
        llm,
        tts,
        transport.output(),
        assistant_aggregator,
    ]
)

HecttorAudioProcessor accepts the same api_key, model_name, and chunk_size_ms parameters as HecttorFilter, plus asr_weight and vad_tt_weight blend factors in [0.0, 1.0] (None uses the model's default weight).

Notes:

  • The current implementation runs two enhancer sessions, one per weight, doubling enhancement compute and SDK usage accounting. A single-pass multi-weight SDK API may replace this later.
  • Anything placed downstream of vad_tt_stage() (e.g. audio recorders) sees the VAD/TT blend.

Runtime toggle

Disable and re-enable denoising at runtime with Pipecat's FilterEnableFrame:

from pipecat.frames.frames import FilterEnableFrame

await worker.queue_frame(FilterEnableFrame(False))  # disable
await worker.queue_frame(FilterEnableFrame(True))  # re-enable

Running the example

examples/voice-hecttor.py is a complete voice bot with Hecttor denoising, Deepgram STT, OpenAI LLM, and Cartesia TTS.

  1. Install the example's dependencies:

    uv add pipecat-hecttor "pipecat-ai[deepgram,cartesia,openai,silero,webrtc,runner]"
    
  2. Install the hecttor_sdk wheel (see Installation).

  3. Create a .env file with your keys:

    HECTTOR_API_KEY=...
    DEEPGRAM_API_KEY=...
    OPENAI_API_KEY=...
    CARTESIA_API_KEY=...
    
  4. Run the bot and open the printed URL in your browser:

    python examples/voice-hecttor.py
    

Testing the filter offline

scripts/test_hecttor_filter_audiofile.py runs the filter over a pre-recorded audio file so you can compare original and enhanced audio:

uv add soundfile
python scripts/test_hecttor_filter_audiofile.py input.wav output.wav

It reports the realtime factor and before/after audio statistics.

Testing the processor offline

scripts/test_hecttor_processor_audiofile.py runs the full two-stage HecttorAudioProcessor pipeline over a pre-recorded audio file — including a real Silero VAD on the swapped stream when pipecat-ai[silero] is installed — and verifies the processor's invariants, exiting non-zero on any violation:

uv add soundfile
python scripts/test_hecttor_processor_audiofile.py input.wav --asr-weight 1.0 --vad-tt-weight 0.3

It saves both blends (asr_blend.wav, vad_tt_blend.wav) so you can compare them by ear, and reports frame counts, VAD events, and the realtime factor.

Compatibility

Tested with Pipecat v1.7.0.

Pipecat evolves rapidly; if you hit a compatibility issue with a newer release, please open an issue.

License

BSD 2-Clause — see LICENSE.

Metadata

Release files for pipecat-hecttor 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pipecat-hecttor 0.3.0
File Size Uploaded
pipecat_hecttor-0.3.0.tar.gz 20.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pipecat-hecttor 0.3.0
File Interpreter ABI Platform
pipecat_hecttor-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 33.3 kB

Release files / pipecat_hecttor-0.3.0.tar.gz

Download URL pipecat_hecttor-0.3.0.tar.gz
Size 20.2 kB
Tags Source
SHA-256 checksum
How to use checksums
d9daa168ac9994fa7355acfbbb30c6741ca44e668e10c51af731b1a7bd4bdaaf
BLAKE2b-256 checksum
How to use checksums
7b29690a75412296bef63f38a2e57bc31927576cdc24ae896bc0840adfa9d997
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / pipecat_hecttor-0.3.0-py3-none-any.whl

Download URL pipecat_hecttor-0.3.0-py3-none-any.whl
Size 13.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f4feb31065c1a852e809ad8bb195243c5d4fc141deddefe50cab45194eec16b4
BLAKE2b-256 checksum
How to use checksums
67643e5c9dfba837d7f85f365a530a97fd39d2bfbe4527f2da387635d1ffd405
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page