Skip to main content

parlaibero-mcp

An MCP server that gives AI agents access to ParlaIbero: 10 million plenary interventions from the lower or single chambers of 16 Ibero-American countries, linked to their deputies, deposited in Harvard Dataverse (CC BY 4.0, one DOI per country).

The server downloads each country's dataset once (MD5-verified, straight from Dataverse), loads it into a local DuckDB database and exposes tools to search, count and read it. Nothing is sent anywhere except plain GET requests to Dataverse.

Install

You need uv. No clone is required.

Claude Code

claude mcp add parlaibero -s user -- uvx --from "git+https://github.com/rodrodr/parlaibero-tools#subdirectory=mcp" parlaibero-mcp

Codex CLI

codex mcp add parlaibero -- uvx --from "git+https://github.com/rodrodr/parlaibero-tools#subdirectory=mcp" parlaibero-mcp

Claude Desktop (claude_desktop_config.json), Cursor (~/.cursor/mcp.json), Gemini CLI (~/.gemini/settings.json), Windsurf and most other clients use the same block:

{
  "mcpServers": {
    "parlaibero": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/rodrodr/parlaibero-tools#subdirectory=mcp", "parlaibero-mcp"]
    }
  }
}

Desktop apps do not always see your shell's PATH: if the server fails to start, replace "uvx" with the full path printed by which uvx (e.g. /Users/you/.local/bin/uvx).

To install it as a regular command instead: uv tool install "git+https://github.com/rodrodr/parlaibero-tools#subdirectory=mcp", then use parlaibero-mcp as the command.

Tools

Getting and understanding the data

tool what it does
list_countries the 16 datasets: DOI, published version, download size, what is already local
download_country download a country from Dataverse and load it (skips files already current)
import_from_folder load CSV files you downloaded by hand from Dataverse
get_documentation README, data dictionary, known limitations, process report, corpus facts (en/es/pt) — no download needed
describe_data schema of the local tables, loaded countries, example queries
coverage check the base before reading a trend: sessions, speech words, linked and sex-known shares per year or legislature, missing years, low-base periods, each corpus's linkage figures
how_to_cite formatted citation of a country's dataset with DOI and version

Counting words over time

tool what it does
ngram_viewer Google Books Ngram-style series: several words or phrases at once (, separates series, + sums variants, * ends a prefix), per million words, raw counts or % of interventions, optional smoothing, one line per country if wanted. Every point carries its base (words, sessions) and low-base years are flagged. Can write the chart as .svg (figure for a paper or slide) or .html (hover read-out, dark mode, data table)
term_counter totals for one or more terms: occurrences, interventions, sessions, speakers, by sex and party with per-million rates, and the first and last use in each country
term_frequency one term grouped by year, decade, country, legislature, party, sex, speaker or session type

Who speaks and how

tool what it does
share_of_voice share of words (or turns) and of speakers by sex or party, against the group's share of members on the register in that period, and their ratio (> 1 = speaks more than its weight)
distinctive_words the words that most distinguish two groups — women vs men, party vs party, period vs period — by weighted log-odds with an informative Dirichlet prior (Monroe, Colaresi & Quinn 2008)
kwic keyword-in-context concordance lines from a reproducible random sample
collocations words over-represented around a term (Dunning log-likelihood)
search_text find interventions by word, phrase or regex, with filters, newest first
get_intervention · get_session read one turn with its context, or a whole session

Anything else, and reproducibility

tool what it does
query_sql read-only DuckDB SQL; the connection cannot write, read other files or reach the network
export_result save a query's full result as CSV or Parquet for R, Python or Stata
query_log every analysis run (tool, parameters, dataset versions); can write a Markdown methods note

The resource parlaibero://about summarises the collection.

Two filters that change results

Most analysis tools take exclude_chair and max_turn_words, and warn when you should use them:

  • max_turn_words — the published corpora keep documents read into the record (committee reports, bills, lists) as speech when the record marks no separator. They sit in very long turns and, measured on 3 October 2026, hold 54 % of the speech words in Argentina, 26 % in Uruguay and 25 % in Mexico, against under 1 % in Spain or Brazil. Any rate per million words or share of words is affected; max_turn_words=10000 leaves those turns out, and coverage reports the share per country (long_turn_words_pct).
  • exclude_chair — the presiding officer's procedural turns (giving the floor, calling votes) can dominate a group. Comparing women and men deputies in Spain 2016-2023, the most distinctive "women's words" are votación, votos, pausa with the chair in, and mujeres, violencia, igualdad, género without it.

Rows without a date (6,187 in Ecuador, 44 in Costa Rica) are left out of anything by year or period, and the tools say so.

How words are counted

A word is a run of letters, digits or combining marks. Counts are whole-word, case- and accent-insensitive (nacion = Nación), so corrupción does not count anticorrupción — add it as a variant (corrupción+anticorrupción) if you want both. n_words, the precomputed word table and every tool use this same definition, so their figures add up. Single words are served from a precomputed table (instant); phrases and filters by party or deputy scan the text (seconds).

Data model

interventions   country · id_session · id_int · legislature · legislative_session · session_number ·
                date · session_type · intervention_order · speaker_raw · id_dep · speaker_name ·
                sex · party · district · dm_speech · text · n_words
unigrams        word counts of speech rows by country, year and sex (accent-folded)
vocab           display form of each folded word
deputies        country + the core columns of every deputy register
deputies_{iso}  the full register of one country, with its own extra columns
datasets        what is loaded: DOI, version, source, rows, sessions, date range

Things worth knowing before drawing conclusions (the server also tells the agent):

  • dm_speech = 1 marks speech. Rows with 0 are cover pages, summaries, vote tallies and reproduced documents; intervention_order = 0 is the session's Prolegomena.
  • sex, party and district come from the deputy register and are empty when the speaker is not a linked deputy (ministers, clerks, collective or anonymous voices).
  • Coverage, linkage and caveats differ by country: read known_limitations and corpus_info.
  • id_session / id_int are stable within a published version, not across versions. Cite the dataset with its version.

Command line

parlaibero-mcp                    # serve MCP over stdio (what clients run)
parlaibero-mcp download SV ES     # download and load countries ('all' for the 16)
parlaibero-mcp import ~/Downloads/parlaibero
parlaibero-mcp status
parlaibero-mcp reindex             # after upgrading from an older version (no re-download)
parlaibero-mcp remove SV

Large countries (Brazil, Mexico, Portugal, Ecuador, Spain: 1-1.5 GB each) are best downloaded from the terminal so a client time-out does not interrupt them. All 16 take about 12 GB of downloads plus 15 GB of database; the downloaded CSVs can be deleted afterwards to save space (~/.parlaibero/data/{ISO}/*.csv), but they are needed to re-import.

Configuration

variable default
PARLAIBERO_HOME ~/.parlaibero where downloads and the database live
PARLAIBERO_DATAVERSE_URL https://dataverse.harvard.edu Dataverse installation
PARLAIBERO_LOG 1 0 turns off the query log (~/.parlaibero/query_log.jsonl)

License

MIT for the code. The data are CC BY 4.0 and are cited per country with their DOI.

Grant PID2022-141706NB-C22 funded by MICIU/AEI/10.13039/501100011033 and by ERDF/EU.

Metadata

Release files for parlaibero-mcp 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for parlaibero-mcp 0.2.0
File Size Uploaded
parlaibero_mcp-0.2.0.tar.gz 43.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for parlaibero-mcp 0.2.0
File Interpreter ABI Platform
parlaibero_mcp-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 89.3 kB

Release files / parlaibero_mcp-0.2.0.tar.gz

Download URL parlaibero_mcp-0.2.0.tar.gz
Size 43.0 kB
Tags Source
SHA-256 checksum
How to use checksums
e2da9e601fb0efab657fd6aada9e086b9e10fc0a7f653e1076d7d6fd71676d0c
BLAKE2b-256 checksum
How to use checksums
9df049c58d2402128202f080e5ce212459e544c52f71394d1a979bb74d5308a6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / parlaibero_mcp-0.2.0-py3-none-any.whl

Download URL parlaibero_mcp-0.2.0-py3-none-any.whl
Size 46.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
5c56d0203b6196747a575e69cca83da56d5b924a9424ddeb2a4e37c08c587a1b
BLAKE2b-256 checksum
How to use checksums
c0c026363f9ae4be6353a82d67124a30fc3d727a21279835a884e393e74055f5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.10.12 {"installer":{"name":"uv","version":"0.10.12","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

0.4.0

2 release files

0.3.0

2 release files

0.2.1

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page