ChatVoice / Speakr Voice Workspace
ChatVoice is the ChatArch repository and Python package shell for the Speakr voice workspace: a FastAPI + browser product with realtime meeting transcription, local audio capture, AI notes, TTS, and full-duplex voice conversation. The deployed product uses Speakr as its canonical public service, while ChatVoice is the package, CLI, and repository name.
Public site: https://speakr.public.wzhecnu.cn/
Repository: https://github.com/ChatArch/ChatVoice
PyPI package: https://pypi.org/project/ChatVoice/
Documentation: https://arch.gh.wzhecnu.cn/ChatVoice/
The former qwen-audio-demo.public.wzhecnu.cn entry is retired and returns HTTP 410.
Features
- 语音合成 (TTS): server-side proxy for
qwen-audio-3.0-tts-plus, returning playable MP3/WAV audio. - 声音克隆: server-side DashScope TTS v2 voice enrollment (
VoiceEnrollmentService) creates reusablevoice_idvalues; the browser never sees provider credentials. - 实时对话: 独立的豆包式语音对话页;browser WebSocket -> FastAPI proxy -> Qwen Realtime,支持服务端模型列表、VAD、流式文字、24 kHz PCM 播放、自然打断、对话历史和 Markdown 导出。
- 会议录音首页: mobile-first recording surface with live transcript, waveform, pause/resume, finish, and local audio download.
- 暂停整理: pausing a recording commits the current ASR window so pending live/rewrite text can be finalized before continuing.
- 标题刷新与新建快捷入口: the recorder header includes direct
刷新标题and新建buttons for full-session title regeneration and faster mobile topic creation. - 清空防误触: resetting a meeting with existing text/audio state requires confirmation.
- 音频留存默认关闭: 原始录音默认不保存在服务器,也不自动写入浏览器存储;需要音频文件时,用户先点
保存音频,录音结束后再从当前浏览器下载。 - Bounded local archive: when the user opts into
保存音频, MediaRecorder emits one-second chunks into browser-only IndexedDB; the full Blob is assembled only when the user requests a download. - 语音转写: the recorder streams microphone PCM16 to the ASR WebSocket and appends normalized final segments to the timeline.
- API-first ASR: production ASR is designed around
api-server, where the ChatVoice backend calls either a managed ASR API or a self-hosted GPU ASR server.stub-localremains available for contract smoke, andfunasr-gpu/funasr-cpuremain compatibility channels. - Realtime ASR WebSocket:
WS /ws/asr/streamaccepts continuous PCM16 microphone frames and returns cumulative revision events. Long recordings transparently roll a bounded context window while confirmed text continues to grow. - 会议纪要: final transcript segments can be sent to a server-side Qwen-compatible model for summary, action items, risks, and open questions.
- 双模式会议历史: guests keep meeting text and summaries only in browser IndexedDB; signed-in accounts sync records through authenticated server storage.
- 0.1 API 访问: signed-in users can generate one-time-visible API tokens from the web settings panel;
chatvoice data ...can then read meetings, summaries, and realtime conversations from a running service. - 受邀账号登录: public registration is disabled. Accounts are provisioned by
chatvoice accounts add; passwords use salted PBKDF2 hashes, sessions use HttpOnly cookies, and record writes require CSRF tokens.
Security model
- The browser never receives or stores provider credentials.
- Set provider credentials only in server-side environment/config storage; never expose them to the browser.
- Do not commit real env files, model caches, probe output, audio files, runtime logs, or generated API token values.
- Guest meeting and conversation records never enter the server database. Audio/transcript data still passes through ASR/summary or Realtime services while a request is processed.
- Raw recording blobs are not uploaded by the meeting-history feature; the current page keeps them only for local download.
- Realtime history stores text, model, and voice only; raw conversation audio is never written to history storage.
- API tokens are stored server-side as hashes. Token values are displayed only once on creation.
Quick start from the released package
python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "ChatVoice[web]==0.1.6"
chatvoice --tree
chatvoice service plan --ensure-dirs --json
chatvoice serve app --host 127.0.0.1 --port 18087
Open:
http://127.0.0.1:18087/
For a credential-free wiring smoke, use:
export CHATVOICE_ASR_CHANNEL=stub-local
chatvoice serve app --host 127.0.0.1 --port 18087
For a real ASR backend, keep credentials server-side and call an API provider:
export CHATVOICE_ASR_CHANNEL=api-server
export CHATVOICE_ASR_API_URL="https://<asr-service>/v1/transcribe"
# Configure the optional ASR bearer token in server-side config/env storage; do not put it in argv.
chatvoice serve app --host 127.0.0.1 --port 18087
Meeting summary generation is also a server-side model boundary: configure the notes model/provider in server-side environment or config storage, and let the browser/API read only the saved summary text.
Fresh account, browser, token, and data flow
Create one invited account in the same local runtime database used by the service:
read -r -s CHATVOICE_ACCOUNT_LOGIN
export CHATVOICE_ACCOUNT_LOGIN
chatvoice accounts add person@example.com --display-name "Person" --password-env CHATVOICE_ACCOUNT_LOGIN --json
chatvoice accounts list --json
Then:
- open the web service;
- log in with the invited account;
- create or open a meeting and generate its summary;
- open 识别设置 → API Token → 生成 Token;
- copy the token immediately; it is only shown once.
The same token lifecycle is available from CLI after the service is running:
chatvoice tokens create --url http://127.0.0.1:18087 --account person@example.com --password-env CHATVOICE_ACCOUNT_LOGIN --name automation --json
chatvoice tokens list --url http://127.0.0.1:18087 --account person@example.com --password-env CHATVOICE_ACCOUNT_LOGIN --json
Use the token with the data API/CLI:
read -r -s CHATVOICE_DATA_READ
export CHATVOICE_DATA_READ
chatvoice data meetings --url http://127.0.0.1:18087 --token-env CHATVOICE_DATA_READ --json
chatvoice data conversations --url http://127.0.0.1:18087 --token-env CHATVOICE_DATA_READ --json
Database and concurrency
The packaged v0.1.6 web app defaults to SQLite WAL at:
<chatarch-home>/chatvoice/data/meetings.sqlite3
Use one service process (--workers 1) with SQLite. For high-concurrency production, migrate the storage layer to Postgres/MySQL before scaling workers or nodes. An external database URL setting is detected by chatvoice doctor / chatvoice service plan, but the v0.1.6 packaged legacy storage layer still supports SQLite only.
运行目录与数据结构
pip install 后代码安装在当前 Python 环境的 site-packages/chatvoice/,CLI 在对应环境的 bin/chatvoice;生产建议使用独立 venv。运行数据不写入源码目录,默认 root 解析顺序是 CHATVOICE_RUNTIME_ROOT、CHATVOICE_HOME、CHATARCH_HOME/chatvoice、~/.chatarch/chatvoice。默认结构:
~/.chatarch/chatvoice/
├── data/meetings.sqlite3
├── logs/
├── run/
├── temp/asr/
└── model-cache/
后端 SQLite meetings.sqlite3 目前包含 accounts、auth_sessions、api_tokens、meeting_records、conversation_records。转写、summary、实时对话消息以 JSON 字符串保存;原始音频不进后端数据库。访客模式仍使用浏览器 IndexedDB 保存本地会议和录音分片。高并发 Postgres/MySQL 迁移列入 TODO,当前 0.1.6 仍只支持 SQLite WAL + 单服务进程。详见 运行目录与数据结构。
API surface
GET /api/heartbeat 是轻量服务心跳,用来判断 Web 服务、SQLite 只读探测和 ASR 状态是否正常;asr.status 会返回 ready、processing 或 degraded,并包含 FunASR 模型是否已热、最近一次识别成功/失败、耗时和输出长度。录音 WebSocket 也会发 asr.stream.processing / asr.stream.heartbeat,前端会显示“模型加载中/识别处理中/失败原因”,不再静默录音无文字。
GET /api/status: redacted backend status, models, ASR channels, and route shapes.GET /api/heartbeat: lightweight service/database/ASR heartbeat with model warm-up, processing, and recent error/success state.POST /api/tts: JSON{text, voice, format}->audio/mpegoraudio/wav.POST /api/voice-cloning/create: JSON{audio_url, prefix, target_model, language_hints}-> server-createdvoice_id.GET /api/voice-cloning/list: list server-side voice enrollment ids by prefix.GET /api/asr/channels: available ASR channels.POST /api/asr: programmatic/smoke multipart upload endpoint withchannel=api-server|funasr-gpu|funasr-cpu|stub-local.WS /ws/asr/stream: bounded PCM16 stream used by the recorder.GET /api/realtime/models: Realtime models currently exposed by the configured account.WS /ws/realtime?model=<id>: browser-to-backend Realtime proxy.POST /api/meeting-notes/polish: Qwen-compatible chat completion endpoint for transcript polish + realtime summary structure.POST /api/auth/register: intentionally returns403; self-registration is disabled.POST /api/auth/login,GET /api/auth/session,POST /api/auth/logout: invited-account session lifecycle.GET|PUT|DELETE /api/meetings[/<id>]: authenticated meeting record storage. Writes require the session CSRF token.GET|PUT|DELETE /api/conversations[/<id>]: authenticated text-only Realtime conversation storage. Writes require the session CSRF token.GET|POST /api/tokens,DELETE /api/tokens/<id>: signed-in session token management.GET /api/data/meetings[/<id>],GET /api/data/conversations[/<id>]: bearer-token data export for meetings, summaries, and realtime conversations.
Verification
python -m pytest -q
PYTHONPATH=src python -m chatvoice.cli --tree
python -m mkdocs build --strict
python -m build
python -m twine check dist/*
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file chatvoice-0.1.6.tar.gz.
File metadata
- Download URL: chatvoice-0.1.6.tar.gz
- Upload date:
- Size: 101.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cdb652bf889672d307f30d65c5364bff01eb7e27dc382dab99c2ad81d258859c
|
|
| MD5 |
d6040e97341c18a5707af9c679398704
|
|
| BLAKE2b-256 |
fa5eff36f0fc8ba4a9c09f22a33e34da3099719924f06ff673416e3fb81a039f
|
Provenance
The following attestation bundles were made for chatvoice-0.1.6.tar.gz:
Publisher:
publish.yml on ChatArch/ChatVoice
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
chatvoice-0.1.6.tar.gz -
Subject digest:
cdb652bf889672d307f30d65c5364bff01eb7e27dc382dab99c2ad81d258859c - Sigstore transparency entry: 2550239384
- Sigstore integration time:
-
Permalink:
ChatArch/ChatVoice@22b58b4626411cd44c0fb483fbb72e7843b8009a -
Branch / Tag:
refs/tags/v0.1.6 - Owner: https://github.com/ChatArch
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@22b58b4626411cd44c0fb483fbb72e7843b8009a -
Trigger Event:
push
-
Statement type:
File details
Details for the file chatvoice-0.1.6-py3-none-any.whl.
File metadata
- Download URL: chatvoice-0.1.6-py3-none-any.whl
- Upload date:
- Size: 98.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cfe3bad16441e81caaf681536bd5a5eecbc17ebee39c3b0660ae5fab31da5c2a
|
|
| MD5 |
2408536b7fb3d51a620715b108eea960
|
|
| BLAKE2b-256 |
b744a8ce3b819003b67382b99450e2f4b8e3f8842b6a25091f6a17d3b9c665dd
|
Provenance
The following attestation bundles were made for chatvoice-0.1.6-py3-none-any.whl:
Publisher:
publish.yml on ChatArch/ChatVoice
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
chatvoice-0.1.6-py3-none-any.whl -
Subject digest:
cfe3bad16441e81caaf681536bd5a5eecbc17ebee39c3b0660ae5fab31da5c2a - Sigstore transparency entry: 2550239406
- Sigstore integration time:
-
Permalink:
ChatArch/ChatVoice@22b58b4626411cd44c0fb483fbb72e7843b8009a -
Branch / Tag:
refs/tags/v0.1.6 - Owner: https://github.com/ChatArch
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@22b58b4626411cd44c0fb483fbb72e7843b8009a -
Trigger Event:
push
-
Statement type: