parlaibero-mcp
An MCP server that gives AI agents access to ParlaIbero: 10 million plenary interventions from the lower or single chambers of 16 Ibero-American countries, linked to their deputies, deposited in Harvard Dataverse (CC BY 4.0, one DOI per country).
The server downloads each country's dataset once (MD5-verified, straight from Dataverse), loads it into a local DuckDB database and exposes tools to search, count and read it. Nothing is sent anywhere except plain GET requests to Dataverse.
Install
You need uv. The package is on
PyPI; no clone is required. For the development version,
replace parlaibero-mcp with --from "git+https://github.com/rodrodr/parlaibero-tools#subdirectory=mcp" parlaibero-mcp.
Claude Code
claude mcp add parlaibero -s user -- uvx parlaibero-mcp
Codex CLI
codex mcp add parlaibero -- uvx parlaibero-mcp
ChatGPT desktop app and Codex share one configuration (~/.codex/config.toml). The command above
writes it; or add by hand:
[mcp_servers.parlaibero]
command = "uvx"
args = ["parlaibero-mcp"]
ChatGPT on the web only accepts remote servers, so it cannot use this one.
OpenCode (~/.config/opencode/opencode.json):
{
"mcp": {
"parlaibero": { "type": "local", "command": ["uvx", "parlaibero-mcp"], "enabled": true }
}
}
Antigravity (~/.gemini/config/mcp_config.json, or MCP Servers → Manage → View raw config),
Claude Desktop (claude_desktop_config.json), Cursor (~/.cursor/mcp.json), Gemini CLI
(~/.gemini/settings.json), Windsurf and most other clients use the same block:
{
"mcpServers": {
"parlaibero": {
"command": "uvx",
"args": ["parlaibero-mcp"]
}
}
}
Desktop apps do not always see your shell's PATH: if the server fails to start, replace "uvx" with
the full path printed by which uvx (e.g. /Users/you/.local/bin/uvx).
To install it as a regular command instead: uv tool install parlaibero-mcp,
then use parlaibero-mcp as the command.
Tools
Getting and understanding the data
| tool | what it does |
|---|---|
list_countries |
the 16 datasets: DOI, published version, download size, what is already local |
download_country |
download a country from Dataverse and load it (skips files already current) |
import_from_folder |
load CSV files you downloaded by hand from Dataverse |
get_documentation |
README, data dictionary, known limitations, process report, corpus facts (en/es/pt) — no download needed |
describe_data |
schema of the local tables, loaded countries, example queries |
coverage |
check the base before reading a trend: sessions, speech words, linked and sex-known shares per year or legislature, missing years, low-base periods, each corpus's linkage figures |
how_to_cite |
formatted citation of a country's dataset with DOI and version |
Counting words over time
| tool | what it does |
|---|---|
ngram_viewer |
Google Books Ngram-style series: several words or phrases at once (, separates series, + sums variants, * ends a prefix), per million words, raw counts or % of interventions, optional smoothing, one line per country if wanted. Every point carries its base (words, sessions) and low-base years are flagged. Can write the chart as .svg (figure for a paper or slide) or .html (hover read-out, dark mode, data table) |
term_counter |
totals for one or more terms: occurrences, interventions, sessions, speakers, by sex and party with per-million rates, and the first and last use in each country |
term_frequency |
one term grouped by year, decade, country, legislature, party, sex, speaker, session type or session (by='session' shows where a term concentrates — a way to date an event in each chamber) |
Who speaks and how
| tool | what it does |
|---|---|
share_of_voice |
share of words (or turns) and of speakers by sex or party, against the group's share of members on the register in that period, and their ratio (> 1 = speaks more than its weight) |
distinctive_words |
the words that most distinguish two groups — women vs men, party vs party, period vs period — by weighted log-odds with an informative Dirichlet prior (Monroe, Colaresi & Quinn 2008) |
kwic |
keyword-in-context concordance lines from a reproducible random sample |
collocations |
words over-represented around a term (Dunning log-likelihood) |
search_text |
find interventions by word, phrase or regex, with filters, newest first |
get_intervention · get_session |
read one turn with its context, or a whole session |
Anything else, and reproducibility
| tool | what it does |
|---|---|
query_sql |
read-only DuckDB SQL; the connection cannot write, read other files or reach the network |
export_result |
save a query's full result as CSV or Parquet for R, Python or Stata |
query_log |
every analysis run (tool, parameters, dataset versions); can write a Markdown methods note |
Libraries: subsets for focused and comparative work
| tool | what it does |
|---|---|
library_create |
gather the interventions on a topic, in a period or of a group, in one or several countries: one definition per country (terms — min_occurrences keeps those that discuss the topic rather than mention it —, filters, time window), recorded with the edition of the data. event_dates aligns each country's window on its own event |
library_describe |
what is in a library, per country: interventions, sessions, dates, speakers, words, linkage, share of the chair and of very long turns, years too thin to read as trends, the definitions — and warnings for comparing countries. Without a name, lists the libraries |
library_edit |
add or remove interventions (by id or by a query), or attach a note and tags; removed ones stay out when the library is rebuilt |
library_combine |
union, intersection or difference of two libraries |
library_rebuild |
after downloading a new edition: run the definitions again, find hand-added items and notes again by their text, report what came in and what was lost |
library_export · library_import |
for the Diarios Explorer: one .2replib per country (the explorer holds one country at a time) plus an index (.parlaibero-biblioteca.json) that rebuilds the comparative library; optionally the items with metadata as CSV or Parquet. Importing checks every item (row → intervention → date and speaker) and reports what does not match |
delete |
library_delete (asks for confirm=true) |
Every analysis tool takes library= to work inside one; distinctive_words also compares a library
with the rest of its chambers (field='library'). Libraries are kept in their own file
(~/.parlaibero/bibliotecas.duckdb) and attached to every connection as lib, so query_sql can join
lib.library_items with interventions.
⚠ The explorer does not check which country a .2replib belongs to: import each file with its own
country's CSV loaded. The country is in the file name, the library name and its description. The
compatibility of the files is tested against the explorer's real engine, not a copy of it
(diaries_explorer/explorer_src/test/ida_y_vuelta_mcp.mjs).
The resource parlaibero://about summarises the collection.
Two filters that change results
Most analysis tools take exclude_chair and max_turn_words, and warn when you should use them:
max_turn_words— the published corpora keep documents read into the record (committee reports, bills, lists) as speech when the record marks no separator. They sit in very long turns and, measured on 3 October 2026, hold 54 % of the speech words in Argentina, 26 % in Uruguay and 25 % in Mexico, against under 1 % in Spain or Brazil. Any rate per million words or share of words is affected;max_turn_words=10000leaves those turns out, andcoveragereports the share per country (long_turn_words_pct).exclude_chair— the presiding officer's procedural turns (giving the floor, calling votes) can dominate a group. Comparing women and men deputies in Spain 2016-2023, the most distinctive "women's words" are votación, votos, pausa with the chair in, and mujeres, violencia, igualdad, género without it.
Rows without a date (6,187 in Ecuador, 44 in Costa Rica) are left out of anything by year or period, and the tools say so.
How words are counted
A word is a run of letters, digits or combining marks. Counts are whole-word, case- and
accent-insensitive (nacion = Nación), so corrupción does not count anticorrupción — add
it as a variant (corrupción+anticorrupción) if you want both. n_words, the precomputed word
table and every tool use this same definition, so their figures add up. Single words are served
from a precomputed table (instant); phrases and filters by party or deputy scan the text (seconds).
Data model
interventions country · id_session · id_int · legislature · legislative_session · session_number ·
date · session_type · intervention_order · speaker_raw · id_dep · speaker_name ·
sex · party · district · dm_speech · text · n_words · row_n
(row_n = record number in the published CSV, the explorer's speech_id)
unigrams word counts of speech rows by country, year and sex (accent-folded)
vocab display form of each folded word
deputies country + the core columns of every deputy register
deputies_{iso} the full register of one country, with its own extra columns
datasets what is loaded: DOI, version, source, rows, sessions, date range
lib.libraries · lib.library_items · lib.library_parts the user's libraries (bibliotecas.duckdb)
Things worth knowing before drawing conclusions (the server also tells the agent):
dm_speech = 1marks speech. Rows with0are cover pages, summaries, vote tallies and reproduced documents;intervention_order = 0is the session's Prolegomena.sex,partyanddistrictcome from the deputy register and are empty when the speaker is not a linked deputy (ministers, clerks, collective or anonymous voices).- Coverage, linkage and caveats differ by country: read
known_limitationsandcorpus_info. id_session/id_intare stable within a published version, not across versions. Cite the dataset with its version.
Command line
parlaibero-mcp # serve MCP over stdio (what clients run)
parlaibero-mcp download SV ES # download and load countries ('all' for the 16)
parlaibero-mcp import ~/Downloads/parlaibero
parlaibero-mcp status
parlaibero-mcp reindex # after upgrading: reloads each country from the files on disk
parlaibero-mcp remove SV
Large countries (Brazil, Mexico, Portugal, Ecuador, Spain: 1-1.5 GB each) are best downloaded from
the terminal so a client time-out does not interrupt them. All 16 take about 12 GB of downloads
plus 15 GB of database. Keep the downloaded CSVs (~/.parlaibero/data/{ISO}/*.csv): reindex reloads
from them after an upgrade, and without them a country has to be downloaded again.
Configuration
| variable | default | |
|---|---|---|
PARLAIBERO_HOME |
~/.parlaibero |
where downloads and the database live |
PARLAIBERO_DATAVERSE_URL |
https://dataverse.harvard.edu |
Dataverse installation |
PARLAIBERO_LOG |
1 |
0 turns off the query log (~/.parlaibero/query_log.jsonl) |
License
MIT for the code. The data are CC BY 4.0 and are cited per country with their DOI.
Grant PID2022-141706NB-C22 funded by MICIU/AEI/10.13039/501100011033 and by ERDF/EU.
Metadata
Release files for parlaibero-mcp 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| parlaibero_mcp-0.3.0.tar.gz | 65.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| parlaibero_mcp-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 133.0 kB
Release files / parlaibero_mcp-0.3.0.tar.gz
| Download URL | parlaibero_mcp-0.3.0.tar.gz |
|---|---|
| Size | 65.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
d90016c789c9ea809f1cb4d50ef652311380bc55550cd7519809b2e26b1ac52f
|
|
BLAKE2b-256 checksum How to use checksums |
1b912376ea99c0424e423d2eea9fb5dafa958c2fe6282d4e7fb52dfb303ce0c6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / parlaibero_mcp-0.3.0-py3-none-any.whl
| Download URL | parlaibero_mcp-0.3.0-py3-none-any.whl |
|---|---|
| Size | 67.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0beefa54febc63ca7f9158a7386c848217d2c8594297fecd6422ae3ceb0d2331
|
|
BLAKE2b-256 checksum How to use checksums |
0c9bb03b8cfb47fae155a14dc5c15e112ed182b56057a196e64874fa777a8073
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|