ChatVoice / Speakr Voice Workspace
ChatVoice is the ChatArch repository and Python package shell for the Speakr voice workspace: a FastAPI + browser product with realtime meeting transcription, AI notes, TTS, and full-duplex voice conversation. The deployed product uses Speakr as its canonical public service, while ChatVoice is the package, CLI, and repository name.
Public site: https://speakr.public.wzhecnu.cn/
Repository: https://github.com/ChatArch/ChatVoice
PyPI package: https://pypi.org/project/ChatVoice/
Documentation: https://arch.gh.wzhecnu.cn/ChatVoice/
The former qwen-audio-demo.public.wzhecnu.cn entry is retired and returns HTTP 410.
Features
-
Markdown Todo:摘要下方点击“转为 Todo”,在独立页面编辑、通过对话完善、撤销并导出
.md;不自动生成或执行任务,原摘要保持不变。详见 Markdown Todo。 -
语音合成 (TTS): server-side proxy for
qwen-audio-3.0-tts-plus, returning playable MP3/WAV audio when the model provider key is configured. -
本地声音复刻: Voice Studio is one unified panel: choose 系统音色 (built-in TTS voices) or 我的复刻声音 (upload/record an authorized reference audio sample) and share the same text box. The clone path runs a one-shot VoiceClone/IndexTTS-2.5 job through a local sidecar; the browser receives a playable/downloadable result for the current job only. Within a session the reference audio is kept so the cloned voice can be reused for new text until the page is left. No voice profile or generated-audio history is saved. See 声音复刻使用指南.
-
实时对话: 独立的豆包式语音对话页;browser WebSocket -> FastAPI proxy -> Qwen Realtime,支持服务端模型列表、VAD、流式文字、24 kHz PCM 播放、自然打断、对话历史和 Markdown 导出。
-
会议记录首页: mobile-first recording surface with live transcript, waveform, pause/resume, and finish flow. Audio is used for realtime transcription; the product saves text and summaries, not recording files.
-
暂停整理: pausing a recording commits the current ASR window so pending live/rewrite text can be finalized before continuing.
-
标题刷新与新建快捷入口: the recorder header includes direct
刷新标题and新建buttons for full-session title regeneration and faster mobile topic creation. -
会议标签: the recorder header includes a compact
#tag picker. Meetings default to no tags, support multi-selectthought/diarypresets plus repeated custom tags, and expose tags through web storage and API-token data reads. -
清空防误触: resetting a meeting with existing text/summary state requires confirmation.
-
原始录音不保存: 当前会议记录功能不提供“保存录音”或“下载录音”。服务器只保存文字、摘要和会议元数据;访客模式只在浏览器保存文字/摘要。详见 录音保存边界。
-
语音转写: the recorder streams microphone PCM16 to the ASR WebSocket and appends normalized final segments to the timeline.
-
API-first ASR: production ASR is designed around
api-server, where the ChatVoice backend calls either a managed ASR API or a self-hosted GPU ASR server.stub-localremains available for contract smoke, andfunasr-gpu/funasr-cpuremain compatibility channels. -
Realtime ASR WebSocket:
WS /ws/asr/streamaccepts continuous PCM16 microphone frames and returns cumulative revision events. Long recordings transparently roll a bounded context window while confirmed text continues to grow. -
会议纪要与标题独立模型: 纪要、画布修改和标题可通过 ChatEnv 分别配置独立的 OpenAI-compatible Base/Key/Model;未配置时保留旧路径,不影响语音 Token Plan 保护。见 独立文本配置 / English。
0.1.15.post1是源码构建 hotfix,不代表已经发布到 PyPI。 -
双模式会议历史: guests keep meeting text and summaries only in browser IndexedDB; signed-in accounts sync records through authenticated server storage.
-
0.1 API 访问: signed-in users can generate one-time-visible API tokens from the web settings panel;
chatvoice data ...can then read meetings, summaries, and realtime conversations from a running service. -
受邀账号登录: public registration is disabled. Accounts are provisioned by
chatvoice accounts add; passwords use salted PBKDF2 hashes, sessions use HttpOnly cookies, and record writes require CSRF tokens.
登录后端与前端边界
ChatVoice 依赖 ChatLogin>=0.1.1,<0.2.0 的认证与会话核心,但保留自己的 HTML、CSS、原生 JavaScript 和登录/访客弹窗,不注入 ChatLogin 默认模板。宿主适配层复用现有 accounts 与 auth_sessions 表、原账号 ID、PBKDF2 材料和 Cookie;不新建第二套用户库、不强制改密码。
/api/auth/* 和原 JSON 字段保持兼容。已有账号映射为普通用户,不引入 Web Admin;会议/对话 owner 检查与 API Token scope 仍由 ChatVoice 负责。访客记录继续保存在浏览器 IndexedDB,不因接入登录库而自动上传。部署与生产数据迁移仍是独立操作,发布包不等于重启服务。
Security model
- The browser never receives or stores provider credentials.
- Set provider credentials only in server-side environment/config storage; never expose them to the browser.
- Do not commit real env files, model caches, probe output, audio files, runtime logs, or generated API token values.
- Guest meeting and conversation records never enter the server database. Audio/transcript data still passes through ASR/summary or Realtime services while a request is processed.
- Raw meeting recordings are not saved by the meeting recorder: no backend raw-audio database/object store, no browser-local recording chunks, and no recording download endpoint in the current version.
- Realtime history stores text, model, and voice only; raw conversation audio is never written to history storage.
- API tokens are stored server-side as hashes. Token values are displayed only once on creation.
Quick start from the released package
本分支 0.1.15.post3 是本地源码 hotfix,尚未发布;下面锁定版本的 PyPI 命令仅在正式发布后可用,发布前应安装经过验证的本地 wheel。独立 TTS 配置 支持可配置协议、端点、模型、凭据和音色,配置不完整时拒绝而不回退;ASR、实时对话和 VoiceClone 不变。纪要和标题配置仍见独立文本模型。
python -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install "ChatVoice[web]==0.1.16"
chatvoice --tree
chatvoice --tree-brief
chatvoice service plan --ensure-dirs --json
chatvoice serve app --host 127.0.0.1 --port 18087
Open:
http://127.0.0.1:18087/
For a credential-free wiring smoke, use:
export CHATVOICE_ASR_CHANNEL=stub-local
chatvoice serve app --host 127.0.0.1 --port 18087
For a real ASR backend, keep credentials server-side and call an API provider:
export CHATVOICE_ASR_CHANNEL=api-server
export CHATVOICE_ASR_API_URL="https://<asr-service>/v1/transcribe"
# Configure the optional ASR bearer token in server-side config/env storage; do not put it in argv.
chatvoice serve app --host 127.0.0.1 --port 18087
Meeting summary generation is a separate server-side model boundary: summarize/polish/revise can use CRS via CHATVOICE_MEETING_NOTES_PROVIDER=crs-chat-completions and CHATVOICE_MEETING_NOTES_CRS_PROFILE=apple. Realtime and legacy TTS retain Token Plan CHATVOICE_OPENAI_API_*; independent TTS uses only CHATVOICE_TTS_*.
Fresh account, browser, token, and data flow
Create one invited account in the same local runtime database used by the service:
read -r -s CHATVOICE_ACCOUNT_LOGIN
export CHATVOICE_ACCOUNT_LOGIN
chatvoice accounts add person@example.com --display-name "Person" --password-env CHATVOICE_ACCOUNT_LOGIN --json
chatvoice accounts list --json
Then:
- open the web service;
- log in with the invited account;
- create or open a meeting and generate its summary;
- open 识别设置 → API Token → 生成 Token;
- copy the token immediately; it is only shown once.
The same token lifecycle is available from CLI after the service is running:
chatvoice tokens create --url http://127.0.0.1:18087 --account person@example.com --password-env CHATVOICE_ACCOUNT_LOGIN --name automation --json
chatvoice tokens list --url http://127.0.0.1:18087 --account person@example.com --password-env CHATVOICE_ACCOUNT_LOGIN --json
Use the token with the data API/CLI:
read -r -s CHATVOICE_DATA_READ
export CHATVOICE_DATA_READ
chatvoice data meetings --url http://127.0.0.1:18087 --token-env CHATVOICE_DATA_READ --json
chatvoice data conversations --url http://127.0.0.1:18087 --token-env CHATVOICE_DATA_READ --json
Database and concurrency
The packaged web app stores service data in one SQLite WAL file at:
<chatarch-home>/chatvoice/data/meetings.sqlite3
Use one service process (--workers 1) with SQLite. Back up and move data with the CLI single-file dump/restore commands. There is no DATABASE_URL ChatVoice setting in the packaged storage layer; the active database is just the resolved SQLite file. Future high-concurrency Postgres/MySQL support is a separate storage-layer migration.
运行目录与数据结构
pip install 后代码安装在当前 Python 环境的 site-packages/chatvoice/,CLI 在对应环境的 bin/chatvoice;生产建议使用独立 venv。运行数据不写入源码目录,默认 root 解析顺序是 CHATVOICE_RUNTIME_ROOT、CHATVOICE_HOME、CHATARCH_HOME/chatvoice、~/.chatarch/chatvoice。默认结构:
~/.chatarch/chatvoice/
├── data/meetings.sqlite3
├── logs/
├── run/
├── temp/asr/
└── model-cache/
后端 SQLite meetings.sqlite3 目前包含 accounts、auth_sessions、api_tokens、meeting_records、conversation_records。转写、summary、会议标签、实时对话消息以 JSON 字符串保存;原始音频不进后端数据库。访客模式仍使用浏览器 IndexedDB 保存本地会议文字、标签和摘要,不保存录音分片。当前版本支持 SQLite WAL + 单服务进程;数据备份/迁移使用 CLI 单文件 dump/restore。高并发 Postgres/MySQL 是未来单独 storage-layer migration。详见 运行目录与数据结构 和 录音保存边界。
API surface
GET /api/heartbeat 是轻量服务心跳,用来判断 Web 服务、SQLite 只读探测和 ASR 状态是否正常;asr.status 会返回 ready、processing 或 degraded,并包含 FunASR 模型是否已热、最近一次识别成功/失败、耗时和输出长度。录音 WebSocket 也会发 asr.stream.processing / asr.stream.heartbeat,前端会显示“模型加载中/识别处理中/失败原因”,不再静默录音无文字。
GET /api/status: redacted backend status, models, ASR channels, sidecar configuration, and route shapes.GET /api/heartbeat: lightweight service/database/ASR heartbeat with model warm-up, processing, and recent error/success state.POST /api/tts: JSON{text, voice, format}->audio/mpegoraudio/wav.GET /api/voice-clone/status: redacted local VoiceClone sidecar status.POST /api/voice-clone/jobs: authenticated multipart{text, lang, duration_factor, reference_audio}-> one-shot local clone job.GET /api/voice-clone/jobs/<id>: authenticated job polling with progress/stage/ETA.GET /api/voice-clone/jobs/<id>/audio: authenticated generated audio download/preview.POST /api/voice-cloning/create: legacy DashScope enrollment endpoint for reusablevoice_idcreation; not the primary browser Voice Studio flow.GET /api/voice-cloning/list: legacy list of server-side voice enrollment ids by prefix.GET /api/asr/channels: available ASR channels.POST /api/asr: programmatic/smoke multipart upload endpoint withchannel=api-server|funasr-gpu|funasr-cpu|stub-local.WS /ws/asr/stream: bounded PCM16 stream used by the recorder.GET /api/realtime/models: Realtime models currently exposed by the configured account.WS /ws/realtime?model=<id>: browser-to-backend Realtime proxy.POST /api/meeting-notes/polish: server-side transcript polish + realtime summary endpoint; supports Token Plan chat completions by default or CRScrs-chat-completionsvia the separate meeting-notes provider settings.POST /api/auth/register: intentionally returns403; self-registration is disabled.POST /api/auth/login,GET /api/auth/session,POST /api/auth/logout: invited-account session lifecycle.GET|PUT|DELETE /api/meetings[/<id>]: authenticated meeting record storage, including optionaltags: string[]metadata. Writes require the session CSRF token.GET|PUT|DELETE /api/conversations[/<id>]: authenticated text-only Realtime conversation storage. Writes require the session CSRF token.GET|POST /api/tokens,DELETE /api/tokens/<id>: signed-in session token management.GET /api/data/meetings[/<id>],GET /api/data/conversations[/<id>]: bearer-token data export for meetings, tags, summaries, and realtime conversations.
Verification
python -m pytest -q
PYTHONPATH=src python -m chatvoice.cli --tree
PYTHONPATH=src python -m chatvoice.cli --tree-brief
python -m mkdocs build --strict
python -m build
python -m twine check dist/*
Metadata
Release files for ChatVoice 0.1.16
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| chatvoice-0.1.16.tar.gz | 276.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| chatvoice-0.1.16-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 408.2 kB
Release files / chatvoice-0.1.16.tar.gz
| Download URL | chatvoice-0.1.16.tar.gz |
|---|---|
| Size | 276.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b17c2530a98212305905c4b46fc570f4d1ffc5516c1e936af5caed26f2ba1c5f
|
|
BLAKE2b-256 checksum How to use checksums |
fad6fd11859bff3d6378494b1766ff0eea10e0724a3a5669b7007ce326a6eb74
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency logRelease files / chatvoice-0.1.16-py3-none-any.whl
| Download URL | chatvoice-0.1.16-py3-none-any.whl |
|---|---|
| Size | 131.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e9508ebd96e62b085db67f1d8a3731845b659b8c6d88b6dc5a5da3349974a3f8
|
|
BLAKE2b-256 checksum How to use checksums |
7252fd43f0512eb8f1a745721a7a08fe368b69a1b31e5e4699d8f31387adea54
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 10, 2026.
Transparency log