Skip to main content

Pipecat service integration for Gnani speech AI — STT & TTS for Indian languages

Project description

pipecat-gnani

PyPI License

Pipecat service integration for Gnani — high-accuracy Speech-to-Text and low-latency Text-to-Speech for Indian languages.

Gnani is a production-ready speech AI platform supporting 10+ Indian languages, real-time streaming, and multilingual transcription.

This integration is maintained by Gnani.ai.

Installation

pip install pipecat-gnani

Or with uv:

uv add pipecat-gnani

This will also install the gnani-vachana (>= 0.7.9) core SDK as a dependency. The Python import package name remains gnani.

The WebRTC quickstart below needs the example extra (Pipecat runner, Silero VAD, WebRTC stack, and Groq LLM):

pip install "pipecat-gnani[example]"
# or: uv add "pipecat-gnani[example]"

From source:

git clone https://github.com/Gnani-AI-Mintlify/pipecat-gnani.git
cd pipecat-gnani
uv pip install -e ".[example]"

Prerequisites

You need a Gnani API key. Gnani APIs have this.

Set your credentials as environment variables:

export GNANI_API_KEY="your-api-key"

Quickstart example

The example lives in this repository's examples/foundational/ directory (not in the PyPI wheel). Clone the repo, then run a small WebRTC bot: Gnani WebSocket STT/TTS, OpenAI LLM, and the Pipecat runner CLI.

  1. Copy examples/foundational/env.example to examples/foundational/.env and set GNANI_API_KEY and GROQ_API_KEY.
  2. With the example extra installed (see Install), from the clone root:
cd examples/foundational
python agent.py -t webrtc

If you use uv in a git checkout, from the repository root:

uv sync --extra example
cd examples/foundational
uv run --extra example python agent.py -t webrtc
  1. Open the Pipecat playground at http://localhost:7860/client, connect, and speak.

Environment variables

Variable Purpose
GNANI_API_KEY API key for Gnani Vachana STT and TTS
GROQ_API_KEY API key for the Groq LLM in the foundational example

Quick Start — Pipeline snippet

The snippet below shows the core Pipeline([...]) wiring used in the foundational example. See examples/foundational/agent.py for the full runnable version.

import os

from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.task import PipelineParams, PipelineTask
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import (
    LLMContextAggregatorPair,
    LLMUserAggregatorParams,
)
from pipecat.services.groq.llm import GroqLLMService
from pipecat.transcriptions.language import Language
from pipecat_gnani import GnaniSTTService, GnaniTTSService

# transport = ...  # WebRTC — see examples/foundational/agent.py

stt = GnaniSTTService(
    api_key=os.environ["GNANI_API_KEY"],
    settings=GnaniSTTService.Settings(language=Language.HI_IN),
)

tts = GnaniTTSService(
    api_key=os.environ["GNANI_API_KEY"],
    settings=GnaniTTSService.Settings(voice="Pranav"),
)

llm = GroqLLMService(
    api_key=os.environ["GROQ_API_KEY"],
    settings=GroqLLMService.Settings(
        model="llama-3.1-8b-instant",
    ),
)

context = LLMContext()
aggregators = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(vad_analyzer=SileroVADAnalyzer()),
)

pipeline = Pipeline(
    [
        transport.input(),
        stt,
        aggregators.user(),
        llm,
        tts,
        transport.output(),
        aggregators.assistant(),
    ]
)

task = PipelineTask(
    pipeline,
    params=PipelineParams(enable_metrics=True),
)

Swap service classes in agent.py for REST or SSE variants — see the file for options (WebSocket STT + TTS is the default for lowest latency and interruption support).

Service Construction

Speech-to-Text (REST)

from pipecat_gnani import GnaniHttpSTTService
from pipecat.transcriptions.language import Language

stt = GnaniHttpSTTService(
    api_key="your-api-key",
    aiohttp_session=session,
    settings=GnaniHttpSTTService.Settings(
        language=Language.HI_IN,
    ),
)

Speech-to-Text (Streaming WebSocket)

from pipecat_gnani import GnaniSTTService
from pipecat.transcriptions.language import Language

stt = GnaniSTTService(
    api_key="your-api-key",
    settings=GnaniSTTService.Settings(
        language=Language.HI_IN,
    ),
)

Text-to-Speech (REST)

from pipecat_gnani import GnaniHttpTTSService

tts = GnaniHttpTTSService(
    api_key="your-api-key",
    aiohttp_session=session,
    settings=GnaniHttpTTSService.Settings(
        voice="Pranav",
    ),
)

Text-to-Speech (SSE Streaming)

from pipecat_gnani import GnaniSSETTSService

tts = GnaniSSETTSService(
    api_key="your-api-key",
    aiohttp_session=session,
    settings=GnaniSSETTSService.Settings(
        voice="Pranav",
    ),
)

Text-to-Speech (WebSocket Streaming)

from pipecat_gnani import GnaniTTSService

tts = GnaniTTSService(
    api_key="your-api-key",
    settings=GnaniTTSService.Settings(
        voice="Pranav",
    ),
)

Services

STT Services

Service Transport Base Class Description
GnaniHttpSTTService REST POST SegmentedSTTService File-based transcription via POST /stt/v3. Requires VAD in pipeline.
GnaniSTTService WebSocket STTService Real-time streaming via wss://api.vachana.ai/stt/v3/stream with VAD events. Emits TranscriptionFrame (final) and InterimTranscriptionFrame when the API sets is_final: false (today Gnani sends final transcripts only).

Streaming PCM Specification

All streaming audio must be sent as raw PCM binary frames — no container format (WAV, MP3) mid-stream.

Property 16 kHz 8 kHz
Encoding PCM signed 16-bit little-endian PCM signed 16-bit little-endian
Sample Rate 16,000 Hz 8,000 Hz
Channels 1 (mono) 1 (mono)
Samples per chunk 512 512
Bytes per frame 1,024 bytes (512 samples × 2 bytes) 1,024 bytes (512 samples × 2 bytes)
Frame duration 32 ms 64 ms

Frames must be sent at real-time cadence. See STT Realtime — PCM Specification for full details.

TTS Services

Service Transport Base Class Description
GnaniHttpTTSService REST POST TTSService Single-request synthesis via POST /api/v1/tts/inference.
GnaniSSETTSService SSE TTSService Streaming synthesis via POST /api/v1/tts/sse. Lower latency than REST.
GnaniTTSService WebSocket InterruptibleTTSService Streaming via wss://api.vachana.ai/api/v1/tts. Lowest latency, interruption support. TTSTextFrames are emitted by the Pipecat base class after each synthesis request.

Supported Languages

STT Languages (Speech-to-Text)

STT uses BCP-47 locale codes (e.g. hi-IN, bn-IN). Note: STT uses the -IN suffix (unlike TTS).

For the full list of supported languages, see STT — Supported Languages.

TTS Languages (Text-to-Speech)

TTS uses ISO 639 language codes (e.g. hi, bn). Note: TTS does not use the -IN suffix.

For the full list of supported languages, see TTS — Supported Languages.

Available Voices

See the official voice list for the latest supported voices.

Voice Gender Description
Pranav Male Bold, Trustworthy
Kaveri Female Confident, Bright
Shubhra Female Gentle, Expressive
Deepak Male Grounded, Conversational

Architecture

gnani-vachana (>=0.7.9)  ← Core SDK on PyPI (import as `gnani`)
        ↑
pipecat-gnani            ← This package (Pipecat service adapters)
  ├── STT: REST + WebSocket
  └── TTS: REST + SSE + WebSocket

This package wraps the gnani SDK into Pipecat's SegmentedSTTService, STTService, TTSService, and InterruptibleTTSService base classes.

Documentation

Pipecat Compatibility

Tested with Pipecat v1.5.0.

License

BSD 2-Clause — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pipecat_gnani-0.5.11.tar.gz (19.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pipecat_gnani-0.5.11-py3-none-any.whl (22.1 kB view details)

Uploaded Python 3

File details

Details for the file pipecat_gnani-0.5.11.tar.gz.

File metadata

  • Download URL: pipecat_gnani-0.5.11.tar.gz
  • Upload date:
  • Size: 19.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for pipecat_gnani-0.5.11.tar.gz
Algorithm Hash digest
SHA256 e7b9c1fd863a4da1e4882d3abf8bed7284a69ab94e9ece66f036fd2767b24d80
MD5 a2ea3f281a2eaa6f7c7fb1c0368bb7e7
BLAKE2b-256 712f2c6a3f8499dd1888bb4f71fdb252b4c3ee6fd5e3447027fc05222e00a6b1

See more details on using hashes here.

Provenance

The following attestation bundles were made for pipecat_gnani-0.5.11.tar.gz:

Publisher: workflow.yml on Gnani-AI-Mintlify/pipecat-gnani

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pipecat_gnani-0.5.11-py3-none-any.whl.

File metadata

  • Download URL: pipecat_gnani-0.5.11-py3-none-any.whl
  • Upload date:
  • Size: 22.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for pipecat_gnani-0.5.11-py3-none-any.whl
Algorithm Hash digest
SHA256 f260106854868697e98c858efb2e546c0d97d2e175c7fc7f44f2094fc15b1e5e
MD5 4792bb4f8baf5fa22053c2c984ba3979
BLAKE2b-256 865532c8addddc660ccd62fa72aa7b7ff1ea06a2c78e9af8d5d6caa7c52ad943

See more details on using hashes here.

Provenance

The following attestation bundles were made for pipecat_gnani-0.5.11-py3-none-any.whl:

Publisher: workflow.yml on Gnani-AI-Mintlify/pipecat-gnani

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page