Skip to main content

Pipecat service integration for Gnani speech AI — STT & TTS for Indian languages

Project description

pipecat-gnani

PyPI License

Pipecat service integration for Gnani — high-accuracy Speech-to-Text and low-latency Text-to-Speech for Indian languages.

Gnani is a production-ready speech AI platform supporting 10+ Indian languages, real-time streaming, and multilingual transcription.

This integration is maintained by Gnani.ai.

Installation

pip install pipecat-gnani

Or with uv:

uv add pipecat-gnani

This will also install the gnani-vachana (>= 0.7.3) core SDK as a dependency. The Python import package name remains gnani.

The WebRTC quickstart below needs the example extra (Pipecat runner, Silero VAD, WebRTC stack, and Groq LLM):

pip install "pipecat-gnani[example]"
# or: uv add "pipecat-gnani[example]"

From source:

git clone https://github.com/Gnani-AI-Mintlify/pipecat-gnani.git
cd pipecat-gnani
uv pip install -e ".[example]"

Prerequisites

You need a Gnani API key. Gnani APIs have this.

Set your credentials as environment variables:

export GNANI_API_KEY="your-api-key"

Quickstart example

The example lives in this repository's examples/foundational/ directory (not in the PyPI wheel). Clone the repo, then run a small WebRTC bot: Gnani WebSocket STT/TTS, OpenAI LLM, and the Pipecat runner CLI.

  1. Copy examples/foundational/env.example to examples/foundational/.env and set GNANI_API_KEY and GROQ_API_KEY.
  2. With the example extra installed (see Install), from the clone root:
cd examples/foundational
python agent.py -t webrtc

If you use uv in a git checkout, from the repository root:

uv sync --extra example
cd examples/foundational
uv run --extra example python agent.py -t webrtc
  1. Open the Pipecat playground at http://localhost:7860/client, connect, and speak.

Environment variables

Variable Purpose
GNANI_API_KEY API key for Gnani Vachana STT and TTS
GROQ_API_KEY API key for the Groq LLM in the foundational example

Quick Start — Pipeline snippet

The snippet below shows the core Pipeline([...]) wiring used in the foundational example. See examples/foundational/agent.py for the full runnable version.

import os

from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import (
    LLMContextAggregatorPair,
    LLMUserAggregatorParams,
)
from pipecat.services.groq.llm import GroqLLMService
from pipecat.transcriptions.language import Language
from pipecat_gnani import GnaniSTTService, GnaniTTSService

# transport = ...  # WebRTC — see examples/foundational/agent.py

stt = GnaniSTTService(
    api_key=os.environ["GNANI_API_KEY"],
    settings=GnaniSTTService.Settings(language=Language.HI_IN),
)

tts = GnaniTTSService(
    api_key=os.environ["GNANI_API_KEY"],
    settings=GnaniTTSService.Settings(voice="Pranav"),
)

llm = GroqLLMService(
    api_key=os.environ["GROQ_API_KEY"],
    settings=GroqLLMService.Settings(
        model="llama-3.1-8b-instant",
    ),
)

context = LLMContext()
aggregators = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer()),
)

pipeline = Pipeline(
    [
        transport.input(),
        stt,
        aggregators.user(),
        llm,
        tts,
        transport.output(),
        aggregators.assistant(),
    ]
)

task = PipelineTask(
    pipeline,
    params=PipelineParams(enable_metrics=True),
)

Swap service classes in agent.py for REST or SSE variants — see the file for options (WebSocket STT + TTS is the default for lowest latency and interruption support).

Service Construction

Speech-to-Text (REST)

from pipecat_gnani import GnaniHttpSTTService
from pipecat.transcriptions.language import Language

stt = GnaniHttpSTTService(
    api_key="your-api-key",
    aiohttp_session=session,
    settings=GnaniHttpSTTService.Settings(
        language=Language.HI_IN,
    ),
)

Speech-to-Text (Streaming WebSocket)

from pipecat_gnani import GnaniSTTService
from pipecat.transcriptions.language import Language

stt = GnaniSTTService(
    api_key="your-api-key",
    settings=GnaniSTTService.Settings(
        language=Language.HI_IN,
    ),
)

Text-to-Speech (REST)

from pipecat_gnani import GnaniHttpTTSService

tts = GnaniHttpTTSService(
    api_key="your-api-key",
    aiohttp_session=session,
    settings=GnaniHttpTTSService.Settings(
        voice="Pranav",
    ),
)

Text-to-Speech (SSE Streaming)

from pipecat_gnani import GnaniSSETTSService

tts = GnaniSSETTSService(
    api_key="your-api-key",
    aiohttp_session=session,
    settings=GnaniSSETTSService.Settings(
        voice="Pranav",
    ),
)

Text-to-Speech (WebSocket Streaming)

from pipecat_gnani import GnaniTTSService

tts = GnaniTTSService(
    api_key="your-api-key",
    settings=GnaniTTSService.Settings(
        voice="Pranav",
    ),
)

Services

STT Services

Service Transport Base Class Description
GnaniHttpSTTService REST POST SegmentedSTTService File-based transcription via POST /stt/v3. Requires VAD in pipeline.
GnaniSTTService WebSocket STTService Real-time streaming via wss://api.vachana.ai/stt/v3/stream with VAD events. Emits TranscriptionFrame (final) and InterimTranscriptionFrame when the API sets is_final: false (today Gnani sends final transcripts only).

Streaming PCM Specification

All streaming audio must be sent as raw PCM binary frames — no container format (WAV, MP3) mid-stream.

Property 16 kHz 8 kHz
Encoding PCM signed 16-bit little-endian PCM signed 16-bit little-endian
Sample Rate 16,000 Hz 8,000 Hz
Channels 1 (mono) 1 (mono)
Samples per chunk 512 512
Bytes per frame 1,024 bytes (512 samples × 2 bytes) 1,024 bytes (512 samples × 2 bytes)
Frame duration 32 ms 64 ms

Frames must be sent at real-time cadence. See STT Realtime — PCM Specification for full details.

TTS Services

Service Transport Base Class Description
GnaniHttpTTSService REST POST TTSService Single-request synthesis via POST /api/v1/tts/inference.
GnaniSSETTSService SSE TTSService Streaming synthesis via POST /api/v1/tts/sse. Lower latency than REST.
GnaniTTSService WebSocket InterruptibleTTSService Streaming via wss://api.vachana.ai/api/v1/tts. Lowest latency, interruption support. TTSTextFrames are emitted by the Pipecat base class after each synthesis request.

Supported Languages

STT Languages (Speech-to-Text)

STT uses BCP-47 locale codes (e.g. hi-IN, bn-IN). Note: STT uses the -IN suffix (unlike TTS).

For the full list of supported languages, see STT — Supported Languages.

TTS Languages (Text-to-Speech)

TTS uses ISO 639 language codes (e.g. hi, bn). Note: TTS does not use the -IN suffix.

For the full list of supported languages, see TTS — Supported Languages.

Available Voices

See the official voice list for the latest supported voices.

Voice Gender Description
Pranav Male Bold, Trustworthy
Kaveri Female Confident, Bright
Shubhra Female Gentle, Expressive
Deepak Male Grounded, Conversational

Architecture

gnani-vachana (>=0.7.3)  ← Core SDK on PyPI (import as `gnani`)
        ↑
pipecat-gnani            ← This package (Pipecat service adapters)
  ├── STT: REST + WebSocket
  └── TTS: REST + SSE + WebSocket

This package wraps the gnani SDK into Pipecat's SegmentedSTTService, STTService, TTSService, and InterruptibleTTSService base classes.

Documentation

Pipecat Compatibility

Tested with Pipecat v1.5.0.

License

BSD 2-Clause — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pipecat_gnani-0.5.9.tar.gz (18.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pipecat_gnani-0.5.9-py3-none-any.whl (21.4 kB view details)

Uploaded Python 3

File details

Details for the file pipecat_gnani-0.5.9.tar.gz.

File metadata

  • Download URL: pipecat_gnani-0.5.9.tar.gz
  • Upload date:
  • Size: 18.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pipecat_gnani-0.5.9.tar.gz
Algorithm Hash digest
SHA256 d5f1555da2f095e555ccfc4c1d4dd409e99ef247e8023dfa166468d1c7840b92
MD5 dfce0f29da8899c807d5c4a06cc837c4
BLAKE2b-256 a558fa9e620bd81d4fc1355ec5936a063476184afb554bcb5902b6b26a0e49d0

See more details on using hashes here.

Provenance

The following attestation bundles were made for pipecat_gnani-0.5.9.tar.gz:

Publisher: workflow.yml on Gnani-AI-Mintlify/pipecat-gnani

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pipecat_gnani-0.5.9-py3-none-any.whl.

File metadata

  • Download URL: pipecat_gnani-0.5.9-py3-none-any.whl
  • Upload date:
  • Size: 21.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pipecat_gnani-0.5.9-py3-none-any.whl
Algorithm Hash digest
SHA256 b05dd6f9f1b3b6b234159c62f9e9dcfc4b1bf3f600e7db16735ac6fbb4caef5e
MD5 1e59ebe276d20a3126f19d02b765707b
BLAKE2b-256 a5bf30a049eb75ed412ff87fa290bfb8303dc871612cc6ab1f22d95b864c8e8c

See more details on using hashes here.

Provenance

The following attestation bundles were made for pipecat_gnani-0.5.9-py3-none-any.whl:

Publisher: workflow.yml on Gnani-AI-Mintlify/pipecat-gnani

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page