aytgcalls
Play audio into Telegram group voice chats from Python.
- Kurigram for MTProto signaling (
phone.*) - aiortc for real media transport — ICE → DTLS-SRTP → RTP/Opus
- FFmpeg for decoding anything into 48 kHz stereo PCM
- No py-tgcalls, no tgcalls, no Telethon. Not wrapped, not vendored, not a dependency.
from aytgcalls import AyClient, AyCreds
client = AyClient(AyCreds.from_env())
await client.start()
await client.play(-1001234567890, "song.mp3") # auto-join + play
await client.skip(-1001234567890)
await client.end(-1001234567890)
await client.stop()
Every action takes chat_id first — one class, multiple voice chats.
Table of contents
- How it works
- A bot cannot join a voice chat
- Installation
- Generating a STRING_SESSION
- Quick start
- Three ways to use it
- API reference
- Queue
- Events
- Video / screen sharing
- Bot command interface
- Configuration
- Error handling
- Verifying it works on a VPS
- Troubleshooting
- What is and is not implemented
- Development
How it works
chat_id
│ channels.getFullChannel / messages.getFullChat
▼
InputGroupCall ──► phone.joinGroupCall(params = {ssrc, ufrag, pwd, fingerprints})
│
▼
UpdateGroupCallConnection
{transport:{…}, audio:{payload-types, rtp-hdrexts}}
│
┌──────────────────────────┴───────────────────────────────────────┐
│ ICE (controlling) → DTLS (client) → SRTP keys │
└──────────────────────────┬───────────────────────────────────────┘
▼
file/URL ─ FFmpeg ─► PCM s16le 48k/2 ─► ring buffer ─► gain ─► Opus 20 ms ─► RTP ─► SFU
The full protocol write-up is in PROTOCOL.md.
A bot cannot join a voice chat
This is a Telegram limitation, not a library one. phone.joinGroupCall is user-only.
So aytgcalls uses the standard two-account pattern:
| Account | Role |
|---|---|
User ("assistant"), via API_ID + API_HASH + STRING_SESSION |
Joins the call and streams audio. Required. |
Bot, via BOT_TOKEN |
Optional. Parses /play, /skip… and dispatches to the assistant. Never touches the call. |
GroupCall checks this at join time and raises BotClientNotAllowed if you hand it a bot session.
Installation
pip install aytgcalls
Removing conflicting Pyrogram forks
Kurigram is a maintained Pyrogram fork and still imports as pyrogram. If the real
pyrogram or pyrofork is installed alongside it, they overwrite each other's files and
you get bizarre import errors. Remove them first:
pip uninstall -y pyrogram pyrofork tgcalls py-tgcalls pytgcalls
pip install -U kurigram aytgcalls
Optional speedup (recommended):
pip install "aytgcalls[fast]" # adds tgcrypto
Verify:
python -c "import pyrogram; print(pyrogram.__version__)" # e.g. 2.2.24 (kurigram)
python -c "import importlib.util; assert importlib.util.find_spec('tgcalls') is None; print('no tgcalls ✓')"
FFmpeg and libopus
# Debian / Ubuntu
sudo apt update && sudo apt install -y ffmpeg libopus0
# Fedora / RHEL
sudo dnf install -y ffmpeg opus
# macOS
brew install ffmpeg opus
Check:
ffmpeg -version
python -c "from aytgcalls.media.opus import opus_available; print('opus ✓' if opus_available() else 'opus ✗')"
If ffmpeg is not on PATH, point at it explicitly:
export AYTGCALLS_FFMPEG=/usr/local/bin/ffmpeg
Generating a STRING_SESSION
Get API_ID / API_HASH from https://my.telegram.org → API development tools, then:
export API_ID=1234567
export API_HASH=0123456789abcdef0123456789abcdef
python examples/scripts/gen_session.py
Log in with the phone number of a normal account (not a bot). The script prints the session string once.
Treat
STRING_SESSIONlike a password. Anyone with it has full access to that account. Keep it in an environment variable or a secret manager.aytgcallsnever hardcodes credentials and never logs them.
Quick start
import asyncio
from aytgcalls import AyClient, AyCreds
async def main():
client = AyClient(AyCreds.from_env())
await client.start()
await client.play(-1001234567890, "song.mp3") # auto-join + play
await asyncio.sleep(30)
await client.stop()
asyncio.run(main())
That's it — one play() call handles join, queue, play, and auto-leave.
Three ways to use it
AyClient — one class, everything
The recommended entry point. One instance per bot, multiple voice chats:
from aytgcalls import AyClient, AyCreds
client = AyClient(AyCreds.from_env())
await client.start()
# every method takes chat_id first
await client.play(chat_id, "song.mp3") # auto-join + play or queue
await client.pause(chat_id)
await client.resume(chat_id)
await client.skip(chat_id)
await client.seek(chat_id, 45)
await client.forward(chat_id, 10)
await client.rewind(chat_id, 10)
await client.volume(chat_id, 80)
await client.mute(chat_id)
await client.unmute(chat_id)
await client.loop(chat_id, "queue")
await client.shuffle(chat_id)
await client.play_video(chat_id, "clip.mp4")
# introspection
print(client.position(chat_id))
print(client.now_playing(chat_id))
print(client.is_connected(chat_id))
# lifecycle
await client.end(chat_id) # stop + leave this chat
await client.stop() # leave all + shutdown
AyClient wraps AyFac internally and mirrors every per-chat control. It also accepts a
pre-built Pyrogram Client if you need more control over startup.
AyFac — factory for many chats
For when you want the factory directly:
from aytgcalls import AyFac
fac = AyFac(user_client)
await fac.play(chat_id, "song.mp3") # creates, joins, plays or queues
await fac.pause(chat_id)
await fac.skip(chat_id)
await fac.volume(chat_id, 80)
await fac.stop(chat_id) # stop + leave
await fac.leave_all() # shutdown
AyCall — one call, full control
For a single persistent voice chat:
from aytgcalls import AyCall
call = AyCall(user_client, chat_id)
await call.play("song.mp3") # joins + plays
await call.pause()
await call.resume()
await call.skip()
await call.end() # stop + leave
API reference
All methods are available on AyClient, AyFac (with chat_id first), and AyCall
(without chat_id). The table shows the AyClient signature.
Playback
await client.play(chat_id, source) # join / play / queue, all automatic
await client.play(chat_id, source, force=True) # jump the queue
await client.add(chat_id, source) # alias for play()
await client.pause(chat_id)
await client.resume(chat_id)
await client.stop_playback(chat_id) # stop audio, stay in the call
await client.stop(chat_id) # stop + leave (same as end())
await client.end(chat_id)
await client.previous(chat_id)
await client.replay(chat_id)
source can be a local path, an http(s) URL, a Pyrogram Message, or an AudioSource.
Seeking
await client.seek(chat_id, 90) # absolute seconds → where we landed
await client.forward(chat_id, 10) # skip forward (default 10 s)
await client.rewind(chat_id, 10) # skip backward (default 10 s)
Queue
await client.shuffle(chat_id)
await client.clear_queue(chat_id)
await client.loop(chat_id, "off") # "off" | "track" | "queue" | "shuffle"
await client.loop(chat_id, 3) # repeat current track 3 more times
await client.loop(chat_id) # read current mode
Volume
await client.volume(chat_id, 80) # local gain, 0..200 (%)
await client.mute(chat_id)
await client.unmute(chat_id)
Video
await client.play_video(chat_id, "clip.mp4") # stream video
await client.stop_video(chat_id) # stop video, audio continues
Introspection
client.now_playing(chat_id) # TrackInfo: title, state, position, duration …
client.position(chat_id) # current playback position (seconds)
client.duration(chat_id) # track duration (None for live)
client.volume(chat_id) # current volume setting
client.playback_state(chat_id) # PlaybackState: PLAYING | PAUSED | IDLE …
client.is_connected(chat_id) # whether we are in the voice chat
client.get_call(chat_id) # underlying GroupCall, or None
await client.get_stats(chat_id) # CallStats: packets, bytes, frames, ICE state …
client.active_calls # dict of all joined calls
len(client) # number of active calls
Join / leave
await client.join(chat_id) # join without playing
await client.leave(chat_id) # leave this chat
Queue
from aytgcalls import AyLoop
await call.queue.add("a.mp3")
await call.queue.add("b.mp3", position=0)
await call.queue.extend(["c.mp3", "https://example.com/d.mp3"])
await call.queue.remove(1)
await call.queue.clear()
await call.queue.shuffle()
await call.queue.move(2, 0)
call.queue.current # AudioSource | None
call.queue.items # tuple of upcoming tracks
call.queue.history # recently finished tracks
await call.queue.next() # advance (respects loop mode)
await call.queue.previous() # step back through history
await call.loop(3) # repeat the current track 3 more times
await call.loop("track") # forever
await call.loop("queue") # whole queue
await call.loop("shuffle") # shuffle + keep looping
await call.loop("off")
Loop accepts a count, a LoopMode, or any friendly word (one, song, all,
playlist, repeat, shuffle, off), so a chat command can be forwarded straight to it.
After loop(3) plays the track three more times the mode falls back to off by itself.
play() returns (track, started_now) so a bot can reply either "playing" or "queued at
#3" without tracking state itself. When a track ends the next one starts on its own.
Transitions are near-gapless: the next FFmpeg process starts the moment the previous one hits EOF, feeding the same ring buffer, so the 20 ms RTP cadence never breaks.
Events
@call.on_stream_end
async def _(call, source, reason):
# reason: FINISHED | FAILED | STOPPED | TIMEOUT
print("finished", source.display_name, reason.value)
@call.on_disconnect
async def _(call, reason):
# reason: REQUESTED | CALL_ENDED | KICKED | TRANSPORT_FAILED | SFU_TIMEOUT
print("disconnected", reason.value)
Video / screen sharing
await call.play_video("clip.mp4") # stream a video file
await call.stop_video() # stop video, audio continues
await call.play("song.mp3") # audio keeps going
The video track is encoded to H.264 by FFmpeg and sent on the dedicated presentation
SSRC (ssrc-groups). Requires Telegram's presentation join path.
Bot command interface
See examples/bot_plus_assistant.py for a complete
command surface using AyFac:
/play <file|url> or reply to a voice/audio message /now /queue
/pause /resume /skip /previous /replay /stop
/seek <secs> /forward [secs] /rewind [secs]
/volume <0-200> /mute /unmute
/loop <n|track|queue|shuffle|off>
There is no /join and no /add — /play covers both.
Configuration
from aytgcalls import AyConfig
config = AyConfig(
ffmpeg_path="ffmpeg",
opus_bitrate=96_000, # 64–128 kbps is Telegram's comfortable range
buffer_ms=400, # jitter buffer depth
prefetch_ms=200, # buffered before the first frame is released
volume=100,
ice_servers=(), # Telegram's SFU is ICE-lite on a public IP; STUN not needed
connect_timeout=20.0,
keepalive_interval=10.0, # phone.checkGroupCall
auto_reconnect=True,
reconnect_max_attempts=8,
)
call = AyCall(client, config=config)
Everything is also readable from the environment:
| Variable | Meaning |
|---|---|
API_ID, API_HASH, STRING_SESSION |
user session (required) |
BOT_TOKEN |
optional command bot |
AYTGCALLS_FFMPEG |
path to the ffmpeg binary |
AYTGCALLS_OPUS_BITRATE |
Opus bitrate in bits/s |
AYTGCALLS_BUFFER_MS, AYTGCALLS_PREFETCH_MS |
buffering |
AYTGCALLS_VOLUME |
initial volume percent |
AYTGCALLS_ICE_SERVERS |
comma-separated STUN/TURN URLs |
AYTGCALLS_CONNECT_TIMEOUT, AYTGCALLS_KEEPALIVE_INTERVAL |
timeouts |
AYTGCALLS_DEBUG=1 |
verbose logging with redacted signaling JSON |
from aytgcalls import enable_debug
enable_debug()
Automation you can turn off
AyConfig(
auto_join=True, # play() joins by itself
auto_leave=True, # leave when the queue runs out
auto_leave_delay=3.0, # grace period, so a quick next request keeps the call
)
The grace period matters: if a user queues another song within auto_leave_delay, the
pending leave is cancelled and the call stays up.
Error handling
AytgcallsError
├── BotClientNotAllowed a bot session was used
├── GroupCallNotFound no active voice chat in that chat
├── NotInGroup the account cannot access the chat
├── AlreadyJoined / NotJoined
├── AlreadyPlaying / NotPlaying
├── InvalidAudioSource / MediaSourceError
├── FFmpegError / FFmpegNotInstalled
├── OpusError
├── TransportError
│ ├── ICEFailed
│ └── DTLSHandshakeFailed
└── TelegramCallError wraps DATA_JSON_INVALID, GROUPCALL_INVALID,
GROUPCALL_FORBIDDEN, CHAT_ADMIN_REQUIRED,
GROUPCALL_SSRC_DUPLICATE_MUCH, JOIN_AS_PEER_INVALID
TelegramCallError attaches an explanation to every known RPC error id:
from aytgcalls.exceptions import AytgcallsError
try:
await client.play(chat_id, "song.mp3")
except AytgcallsError as exc:
await message.reply(f"❌ {exc}")
Verifying it works on a VPS
export API_ID=... API_HASH=... STRING_SESSION=...
export TEST_CHAT_ID=-1001234567890 # a chat with a RUNNING voice chat
python examples/scripts/live_check.py
It joins, streams a tone, asserts that RTP packets actually left the host at the expected 50 packets/second, prints the stats, and leaves cleanly. Exit code 0 = pass.
✅ preflight: ffmpeg, libopus, no py-tgcalls
✅ user session: id=… @…
✅ joined. ssrc=1735203981
✅ ICE=completed DTLS=connected
packets_sent=601 (+600) bytes=147840 frames=600 silence=3 underruns=3
✅ RTP is flowing at the expected 20 ms cadence
✅ real audio frames (not silence) reached the encoder
✅ left cleanly
Troubleshooting
Nobody can hear anything, but packets_sent keeps rising
The media path is fine; you are almost certainly server-side muted. In a group where
members join muted, an admin must unmute the assistant, or the account needs speaking
rights. Watch the logs for Server-side muted with can_self_unmute=False. You can also
try await call.mute(False).
ICEFailed: ICE did not connect
Outbound UDP is blocked. Telegram's media servers need arbitrary outbound UDP (ports vary,
commonly the 40000–65535 range) to 91.108.x.x / 149.154.x.x. Cloud firewalls that only
allow TCP will fail here. There is no TCP fallback in this package.
DTLSHandshakeFailed
ICE succeeded but the handshake did not. Usually a stale call: leave, wait a few seconds,
rediscover and rejoin (auto_reconnect=True does this for you). Check your clock is
correct — certificate validity is time-sensitive.
CHAT_ADMIN_REQUIRED
The account lacks rights for that action. Promote the assistant, or ask an admin to allow
members to speak.
GROUPCALL_SSRC_DUPLICATE_MUCH
A previous session never left cleanly. aytgcalls picks a fresh SSRC per join; wait a few
seconds and retry.
FFmpegNotInstalled
apt install ffmpeg, or set AYTGCALLS_FFMPEG=/path/to/ffmpeg.
GroupCallNotFound
The voice chat is not running (or is scheduled for later). This package joins existing
calls; it does not create them.
Choppy audio / lots of underruns in the stats
Increase buffer_ms / prefetch_ms, especially for remote URLs on a slow link.
ImportError mentioning pyrogram
Two forks are installed at once. See
Removing conflicting Pyrogram forks.
The call is an RTMP broadcast
TransportError: … RTMP/stream broadcast. Those calls do not accept WebRTC publishers at
all; you would have to push to the RTMP URL from phone.getGroupCallStreamRtmpUrl.
See PROTOCOL.md §7.
What is and is not implemented
Implemented:
- discovery, join, keepalive, leave, mute/unmute, volume, update handling, reconnect
- one-call automation:
play()auto-joins, auto-queues and the call auto-leaves when the queue empties - full playback control: play / pause / resume / skip / previous / replay / stop / end, seek / forward / rewind, loop (count·track·queue·shuffle), live position + duration
- Telegram voice notes and audio messages as sources, downloaded and cleaned up for you
- JSON ⇄ SDP bridge (both directions, round-trip tested)
- ICE (controlling) + DTLS-SRTP (client) + RTP with pinned SSRC and the SFU's payload type
- FFmpeg → PCM → ring buffer → gain → Opus 20 ms → paced RTP
- queue with loop/shuffle/history, near-gapless transitions, deterministic teardown
- video / screen sharing via presentation SSRC (H.264, FFmpeg-encoded)
Not implemented, deliberately:
- Receiving other participants' audio (we publish only)
- RTMP / stream-mode calls (detected and rejected with a clear error)
- E2E conference calls (
public_key/blockarguments ofphone.joinGroupCall)
Development
git clone … && cd aytgcalls
pip install -e ".[dev]"
pytest -q # ~250 tests, no network required
ruff check .
mypy aytgcalls
Two of the test modules go further than unit tests:
examples/tests/test_loopback.pystands up a second aiortc peer locally, serialises its parameters into exactly the JSON shape Telegram sends, and runs a real ICE + DTLS-SRTP + Opus RTP session against it.examples/tests/test_integration.pydrives the public API —play()→pause/resume/skip/volume→leave()— against a fake Kurigram client that returns real TL objects and hands off to that local peer. It asserts the far end decodes 48 kHz audio, thatphone.checkGroupCallkeepalives run, thatphone.leaveGroupCallgets our SSRC, and that no tasks or FFmpeg processes leak.
What neither can prove is that Telegram's production SFU accepts the join payload; that
is what examples/scripts/live_check.py is for.
License
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file aytgcalls-0.4.0.tar.gz.
File metadata
- Download URL: aytgcalls-0.4.0.tar.gz
- Upload date:
- Size: 87.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
412cacd1316ecba5398ac0e9595e62141bf5ebdbdee21e6c9484ba4825d722e0
|
|
| MD5 |
f8246844fc7fd956057bb0aa0f841e6f
|
|
| BLAKE2b-256 |
cbdcef4560846aa5c94ee128bd5106e38cdcd3aa3c87fc53f796df28eeccd81c
|
File details
Details for the file aytgcalls-0.4.0-py3-none-any.whl.
File metadata
- Download URL: aytgcalls-0.4.0-py3-none-any.whl
- Upload date:
- Size: 97.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6b5de9c28cbc062cb3788966233dc5c991fa71a43e77e26ecc0c654cabc48465
|
|
| MD5 |
8a025bbf8645144107db3f9ee60935af
|
|
| BLAKE2b-256 |
e2d599d5a63e5ee362dca5c8b46bbee51294a2d748f0e8f6c96686fe0e2cdf5c
|