Skip to main content

African Speech Corpora MCP

African Speech Corpora MCP

A read-only server implementing the Model Context Protocol, an open standard created by Anthropic, over public speech corpora for African languages: Wolof first, plus Pulaar and Sereer.

The project is Wolof-first. The 2.0 catalog contains 14 variants: 11 variants of original sources and 3 derivatives. Pulaar (ful) and Sereer (srr) are represented only by Kallaama variants without local metrics; the project does not claim Swahili or Amharic coverage. The bundled validation lock, generated on August 26, 2026, makes 6 Hugging Face variants queryable. A later validation run may naturally produce a different state.

The server does not train models, download complete corpora, or write to Hugging Face, OpenSLR, Kaggle, GitHub, or any other remote source.

What the server measures

The quality pipeline keeps the following stages separate:

raw audio → usable audio → transcribed audio → audit-accepted audio → expert-verified audio
  • Raw: a file present in the observed snapshot.
  • Usable: a file remaining after the audit's quantifiable exclusions.
  • Transcribed: audio associated with a transcription.
  • Audit accepted: an explicitly named union of expert-verified material and material only assumed valid by the audit.
  • Expert verified: only expert_audited or source_reported_expert. An “a priori” assessment never enters this level.

Seconds are the canonical duration representation. Decimal hours and HH:MM:SS strings are derived at serialization time. Every audited metric states its scope, source, method, observation date, and confidence. A missing value remains null; it is never converted to zero. Figures published by a project remain under published_metrics, separate from local observations.

2026 Wolof snapshot

The machine-readable source is assets/wolof-audit-2026.csv. The protocol, limitations, and discrepancies with the source table's TOTAL cells are documented in assets/wolof-audit-2026.md.

Seven Wolof variants have a local observation: ALFFA, FLEURS wo_sn, Kallaama Wolof, Urban Bus, Waxal crowdsource, Wolof TTS Baamtu, and WolBanking77. Totals are recomputed from these seven rows and are never stored as a redundant manual total: 148,102 usable files, 64,609 transcribed files, 50,494 audit-accepted files, 2,329,382.06 seconds of audio, 515,185.06 transcribed seconds, and 302,048.06 audit-accepted seconds.

Important limitations:

  • the FLEURS, Urban Bus, Waxal, Wolof TTS, and WolBanking77 observations described as “a priori” are audit_assumed_valid, not expert verified;
  • ALFFA explicitly has zero expert-verified files in this snapshot;
  • the 153 Kallaama files are long radio or interview recordings, not 153 speech turns; 36 files and 12:49:36 are attributed to the local protocol's source_reported_expert basis;
  • Urban Bus contains substantial French content without a quantified rate and therefore has language_purity=mixed_fr;
  • Waxal separates 517:38:05 of raw audio, usable notably for SSL, from only 13:41:28 transcribed for supervised ASR;
  • the local Wolof TTS Baamtu snapshot—36,009 files and 37:04:49—is distinct from the rounded public metric, and this TTS corpus is excluded from ASR by default;
  • WolBanking77 distinguishes 2,563 observed audio files from the 9,791 text phrases reported elsewhere;
  • Afrivoice publishes 530.74 hours of Wolof audio, including 102.96 transcribed hours. These are source-published figures, not measurements from the 2026 local audit and not evidence of expert verification. The dataset is auto-gated on Hugging Face, so its files, schema, splits, and durations cannot be independently checked without accepting its access conditions and supplying a token; it therefore remains audit_status=pending;
  • because the exact source observation date is unavailable, observed_at=2026 intentionally has year-only precision.

Installation

The package requires Python 3.11 or newer. Python 3.12 is recommended and is selected explicitly below so that the virtual environment does not accidentally inherit an older system interpreter such as Python 3.9:

git clone https://github.com/papasega/african-speech-mcp.git
cd african-speech-mcp
python3.12 --version
python3.12 -m venv .venv
source .venv/bin/activate
python --version
python -m pip install -e .

Both version commands should report Python 3.12.x. Creating the environment with python -m venv .venv is safe only when that python executable is already Python 3.11 or newer. If installation fails with an error such as:

ERROR: Package 'african-speech-mcp' requires a different Python: 3.9.6 not in '>=3.11'

then the virtual environment was created with Python 3.9.6. Deactivate it, remove or rename that local .venv, install Python 3.12 if necessary, and recreate the environment with python3.12 -m venv .venv. A virtual environment keeps the interpreter with which it was created; activating it does not upgrade Python.

The stdio server starts with no arguments:

african-speech-mcp

It can also be started explicitly:

african-speech-mcp serve --transport stdio
african-speech-mcp serve --transport streamable-http

Example Claude Desktop configuration after installing the package, preferably using an absolute path:

{
  "mcpServers": {
    "african-speech-corpora": {
      "command": "/absolute/path/to/.venv/bin/african-speech-mcp"
    }
  }
}

This refactor does not publish the package. Do not use uvx or a PyPI identifier until a release has been explicitly published.

MCP tools

All eight tools declare read_only_hint=true, destructive_hint=false, and idempotent_hint=true.

Tool Purpose
list_corpora List variants, audit state, recommended use, and validation state.
audit_corpus Return counts, durations, verification basis, warnings, and separate published figures.
plan_training_set Build a plan without double-counting derivatives or including benchmarks by default.
filter_segments Filter only when a row-level manifest exists; otherwise refuse to promise exclusions.
compare_corpora Compare quality, domain, language purity, license, and metric freshness.
search_segments Search rows in a variant whose remote schema has been validated.
corpus_stats Return sizes for the variant's config only and report schema discrepancies.
cite_corpus Return the license, citation, and parent citations for a derivative.

Conceptual examples:

audit_corpus(corpus="waxal-crowdsource")

plan_training_set(language="wol", task="asr", quality="transcribed")
plan_training_set(language="wol", task="asr", quality="expert_verified")
plan_training_set(
  language="wol",
  task="asr",
  quality="audit_accepted",
  include_mixed_language=true,
)

filter_segments(
  corpus="waxal-crowdsource",
  exclude_duplicates=true,
  exclude_non_wolof=true,
  exclude_corrupt=true,
)

The last call currently returns filter_available=false. Aggregate totals establish that Waxal contains 430 duplicates and 22 corrupt or non-Wolof files, but they provide no row identifiers. The server therefore refuses to pretend that it removed those rows.

plan_training_set semantics

quality accepts:

  • any: select by task and license without requiring transcription;
  • transcribed: use actually transcribed duration;
  • audit_accepted: include expert and assumed-valid material, with a visible breakdown and warning;
  • expert_verified: use only genuinely expert verification bases;
  • wolof_only: exclude mixed_fr, mixed, unknown, and variants without sufficient language-purity evidence.

metric_source is either latest_audit (the default) or published. Variants without a known duration remain listed in hours_unknown_for. FLEURS is a benchmark and stays out of training unless include_benchmarks=true. Urban Bus requires include_mixed_language=true. TTS variants stay out of ASR unless explicitly enabled, and derivatives are never added to their parents. Unknown, non-commercial, share-alike, or unconfirmed licenses produce appropriate warnings; commercial_use=null is never presented as commercially compatible.

Persistent Hugging Face validation

african-speech-mcp validate
african-speech-mcp validate --write
african-speech-mcp validate --output validation-lock.json

Validation checks the dataset identifier, config, splits, transcription column, and any declared language column. The lock is written through atomic replacement and records checked_at, observed schema information, and a stable fingerprint. It is also bound to the catalog hash, so the server rejects a stale lock.

By default, --write creates validation-lock.json in the current directory without modifying the installed package. To use it afterward:

export ASM_VALIDATION_LOCK_PATH="$PWD/validation-lock.json"
african-speech-mcp

Possible states are unverified, verified, gated, unavailable, schema_mismatch, and quality_blocked. A responding identifier is never marked verified when the expected config or column is missing. For Afrivoice, set ASM_HF_TOKEN locally before validation; the token is neither serialized nor logged.

corpus_stats filters strictly by config, so FLEURS wo_sn never includes another language. A multilingual distribution without a separate config must declare a language column and accepted values. When the server cannot guarantee that filter, it refuses the request or returns an explicit warning.

Catalog and architecture

src/african_speech_mcp/
├── audit.py              # audited CSV, duration formats, deterministic aggregation
├── catalog.py            # Pydantic v2 variants and invariants
├── cli.py                # serve, validate, catalog-show
├── config.py             # ASM_* settings
├── hf_client.py          # GET only, retries, bounded concurrency
├── models.py             # structured responses
├── server.py             # eight MCP tools
├── validation.py         # schema inspection and atomic lock files
└── data/
    ├── corpora.json      # single canonical catalog
    └── validation-lock.json

There is no longer a second root-level corpora.json. The audit CSV remains canonical under assets/, and Hatch includes it deterministically in the wheel. Tests check totals and references between the CSV and catalog.

Useful variables include ASM_CATALOG_PATH, ASM_VALIDATION_LOCK_PATH, ASM_AUDIT_LOG_PATH, ASM_HF_TOKEN, ASM_REQUEST_TIMEOUT_S, ASM_REQUEST_RETRIES, and ASM_VALIDATION_CONCURRENCY. Runtime logs go to stderr so that stdout remains reserved for the stdio protocol.

Contributing an audit or manifest

A new audit must:

  1. add unit rows to the CSV, never a manual TOTAL row;
  2. provide audit_id, observed_at, source_reference, method, confidence, and verification basis;
  3. distinguish an observed zero from a missing value;
  4. reference the audit from the matching variant;
  5. document limitations in Markdown and add consistency tests.

To enable real filtering, provide a versioned manifest with at least:

audit_id,corpus_key,revision,row_id,audio_path,is_duplicate,is_empty,is_corrupt,is_wolof,transcript_quality,reason

Identifiers must come from the audited snapshot. They must never be invented from aggregate counts.

Development and release preparation

python -m pip install -e ".[dev]"
ruff check src tests scripts
pytest -q
python -m build

CI runs these checks on Python 3.11 and 3.12, installs the built wheel in a clean environment, and tests both the CLI and a real MCP stdio exchange. Version 2.0.0 is a documented breaking release: the public unit is now a single-language, single-task, single-config variant with more structured quality responses.

License

The server code is licensed under Apache-2.0. Each corpus retains its own license. Catalog entries marked TO BE CONFIRMED or SEE DATASET ... intentionally remain uncertain and must be checked against the primary source before any use, especially commercial use.

Metadata

Release files for african-speech-mcp 2.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for african-speech-mcp 2.0.0
File Size Uploaded
african_speech_mcp-2.0.0.tar.gz 49.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for african-speech-mcp 2.0.0
File Interpreter ABI Platform
african_speech_mcp-2.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 95.9 kB

Release files / african_speech_mcp-2.0.0.tar.gz

Download URL african_speech_mcp-2.0.0.tar.gz
Size 49.2 kB
Tags Source
SHA-256 checksum
How to use checksums
66dbc5de65e1b97cf8e42ad84c73d18165f5098b3ff3893c9898ca0807fcbfca
BLAKE2b-256 checksum
How to use checksums
75125c60f6944b45ab89925c27343678cef89adbdb0540b3833379b9177170d4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release files / african_speech_mcp-2.0.0-py3-none-any.whl

Download URL african_speech_mcp-2.0.0-py3-none-any.whl
Size 46.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
426e7ce41df9c2673e8e05a2380415064e36351baa8cf1423a619ef89b49450d
BLAKE2b-256 checksum
How to use checksums
f6751b6e251ea953d7fb7c51574fb6f0e387309e11556395a7f414303a05d3ea
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 27, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page