Skip to main content

CrowdOS MCP Server

Simulated focus groups as agent-callable tools. Exposes the CrowdOS developer API as Model Context Protocol tools so AI agents (Claude Desktop, Cursor, Cline, LangGraph, CrewAI, AutoGPT, Devin, etc.) can run simulated public-opinion research with a single tool call.

What this gives you

Your agent can now do things like:

> Run a focus group on whether companies should mandate 4-day weeks.
   Use 200 agents from the us_general_population preset.

[tool: run_focus_group]
{
  "id": "ad4b3736-...",
  "sentiment_summary": {
    "positive_pct": 71.5, "negative_pct": 18.0, "neutral_pct": 10.5,
    "positive": 143, "negative": 36, "neutral": 21
  },
  "sample_responses": [
    {
      "agent_name": "Maria Chen", "age": 34, "occupation": "Software engineer",
      "sentiment": "positive",
      "reasoning": "It would be great for parents — three full days with the kids ..."
    },
    ...
  ]
}

No-install option: the hosted server

The same ten tools are served directly by the API — add https://api.crowdos.ai/mcp as a remote MCP server in any client that supports them (claude.ai connectors, Cursor, agent frameworks). OAuth-capable clients sign in and approve on first connect; the grant shows up as a revocable key under Account → API keys. Header-based clients pass Authorization: Bearer crowd_.... The hosted server is always current with the deploy; this pip package is the stdio alternative for desktop hosts.

Installation

pip install crowdos-mcp

Then mint an API key at https://crowdos.ai/account/api-keys — API access is included with the Pro and Max subscriptions (panel caps 750 and 2,000 respondents per study respectively).

Configure for Claude Desktop

Edit ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "crowdos": {
      "command": "crowdos-mcp",
      "env": {
        "CROWDOS_API_KEY": "crowd_..."
      }
    }
  }
}

Restart Claude Desktop. The CrowdOS tools should appear in the slash-command picker.

Configure for Cursor

Settings → MCP Servers → Add. Same env block as above; command = crowdos-mcp.

Configure for Cline (VS Code)

Settings → Cline → MCP Servers → Edit JSON:

{
  "mcpServers": {
    "crowdos": {
      "command": "crowdos-mcp",
      "env": { "CROWDOS_API_KEY": "crowd_..." }
    }
  }
}

Tools exposed

Tool What it does Auth
run_focus_group Simulated poll on a topic, returns sentiment + quotes required
run_debate Multi-round simulated debate, returns convergence + key arguments required
compare_options Simulated A/B/C/D test on 2-4 text options, returns leaderboard required
test_pricing Van Westendorp / Gabor-Granger pricing study, returns price points or a revenue curve required
rank_items MaxDiff (best-worst scaling) on 4-200 items, returns a rank order with bootstrap rank intervals required
pretest_questions Pretest 1-40 draft survey questions before fielding them — verdict (ok / check / revise) per question required
preview_cost Estimate cents + seconds before running (free to call) required
estimate_reach Preview how many respondents match a custom_audience filter (free to call) none
list_demographic_presets Discover available audience templates required
get_simulation Fetch full results of a previously-run study required
cancel_simulation Stop a running study (refunds the wallet reservation) required
crowd_sample Browse the public CrowdOS crowd (sanitized) none

run_focus_group and run_debate accept an optional image_url (public HTTPS URL of a JPEG/PNG). One vision pass extracts a description, which every text-only agent reacts to alongside the question. ~$0.001 + ~1-3s overhead, regardless of panel size — useful for testing diagrams, product photos, charts, screenshots, or political imagery.

run_focus_group and run_debate block 5–120s depending on population size — that's a real simulated-research call running behind the scenes, not a cached response. The MCP server returns a trimmed envelope (sentiment summary + first 5 representative quotes

  • billing breakdown). Use get_simulation to pull the full payload when you need every agent's full reasoning.

Custom audiences

run_focus_group, run_debate, compare_options, test_pricing, and rank_items all accept an optional custom_audience object that narrows the panel inside the chosen market preset: geography (countries / regions / cities), age band, genders, interest tags, ethnicity + religion (multi-ethnic markets), and income — either as affluence tiers (0–7) or as real annual amounts in a named currency:

{
  "custom_audience": {
    "cities": ["Bangkok"],
    "age_min": 25, "age_max": 45,
    "income_currency": "THB",
    "income_min_amount": 1500000,
    "income_match": "position"
  }
}

income_match: "position" (default) selects the equivalent affluence standing within each market's own income ladder; "absolute" matches actual earnings at indicative FX. Always preview a filter with the free estimate_reach tool first — it reports the match rate and which constraint cuts hardest. Field-by-field docs: GET https://api.crowdos.ai/api/v1/developer/docs.

Live progress

run_focus_group, run_debate, compare_options, test_pricing, and rank_items stream via SSE under the hood. When your MCP host (Claude Desktop / Cursor / Cline) sends a progressToken with the tool call (the default), the server emits MCP notifications/progress as agents complete — so the host displays "47/200 agents responded" live instead of a blank "running tool..." indicator. Debate runs additionally emit "Round 3/5 starting" / "Round 3 complete" message updates. No client work needed; progress just shows up.

Recommended flow

For non-trivial studies the agent should:

  1. list_demographic_presets — if the user didn't pick one, propose one based on the topic.
  2. preview_cost — gets a cents+seconds estimate before committing.
  3. Confirm with the user — show them the cost estimate and the proposed audience/size.
  4. run_focus_group / run_debate — only after confirmation.

When the study is a questionnaire rather than one question, put pretest_questions before step 2: it runs on a small panel, costs a fraction of the study, and returns a verdict per draft question — fix or drop everything marked revise before spending on the real field. Pretesting after fielding is how a bad question becomes a bad dataset.

Defaults match the platform's calibrated quality bars — Standard Pulse (200 agents) for run_focus_group, the platform-standard debate (30 agents × 5 rounds) for run_debate. Cheaper defaults would silently weaken the output, and the moat is calibrated quality. Plan panel caps are 750 (Pro) and 2,000 (Max); an over_plan_cap error names the cap to retry with.

When a study uses custom_audience, add a step 0: call estimate_reach (free) to check the match rate before paying.

Response shapes

Tool responses are mode-aware and field names are stable. Internal QA fields (consistency_score, model routing, harness flags, sampling metadata) are dropped — agents don't need them.

run_focus_group (voting mode):

{
  "id": "...", "status": "complete", "mode": "voting",
  "topic": "...", "demographic_preset": "us_general_population",
  "population_size": 50,
  "sentiment_summary": {
    "positive": 30, "neutral": 12, "negative": 8,
    "positive_pct": 60.0, "neutral_pct": 24.0, "negative_pct": 16.0
  },
  "sample_responses": [
    { "agent_name": "Maria Chen", "age": 34, "occupation": "Software engineer",
      "sentiment": "positive", "reasoning": "..." },
    "..."
  ],
  "total_responses": 50,
  "billing": { "actual_cents": 12, "plan": "pro" }
}

If a stance_statement was provided, the envelope additionally carries stance_statement + sentiment_axis: { positive_label, neutral_label, negative_label } so the agent knows whether positive means "agrees" vs. "supports".

run_debate:

{
  "id": "...", "status": "complete", "mode": "debate",
  "topic": "...", "num_rounds": 5, "agent_count": 30,
  "final_consensus_score": 0.72,
  "summary": "Most agents converged toward ...",
  "key_arguments_for": [ "..." ],
  "key_arguments_against": [ "..." ],
  "dissenting_views": [ "..." ],
  "final_round_responses": [
    { "agent_name": "...", "position": "FOR", "reasoning": "...",
      "confidence": 0.8 },
    "..."
  ],
  "position_shifts_count": 8,
  "billing": { "actual_cents": 18, "plan": "pro" }
}

get_simulation returns the same envelope as the originating tool but with every response (no 5-quote cap), still with internal QA fields stripped. Use this when the trimmed envelope from run_focus_group / run_debate isn't enough.

Errors

API failures come back as a typed envelope so agents can pattern-match and react. The error field is one of:

error When Agent action
auth_error 401, 403 Get a fresh API key
over_plan_cap 400 — population > tier limit Lower population_size or upgrade
validation_error 400 / 422 Fix the request
insufficient_funds 402 — wallet drained Top up at top_up_url
rate_limited 429 Sleep and retry (retry_after seconds when available)
quota_exceeded 429 — monthly token quota Wait for next month or upgrade
simulation_failed 500 — sim crashed mid-run Inspect sim_id; fall back to a smaller run
server_error 500/502/503/504 Retry with backoff
configuration_error local — CROWDOS_API_KEY unset Tell user to fix host config
internal_error bug in this MCP server Report at the GitHub issues link

Insufficient-funds responses also carry balance_cents, needed_cents, and top_up_url so the agent can surface a clear upgrade prompt.

Configuration

Env var Default Required
CROWDOS_API_KEY — yes (except crowd_sample)
CROWDOS_API_BASE_URL https://api.crowdos.ai no

Cost

CrowdOS uses a metered wallet. The MCP server returns the actual debit on every successful call inside billing.actual_cents. Top up at https://crowdos.ai/account/billing.

A monthly developer-API allowance applies, denominated in list-price spend (Pro $500/mo, Max $2,500/mo, pooled across an organization). GET /api/v1/developer/billing reports usage against it, and the complete written API reference lives at GET https://api.crowdos.ai/api/v1/developer/docs.

Programmatic use (without an MCP host)

The server is also a regular Python module:

python -m crowdos_mcp
# stdio MCP server, waits for messages on stdin

Or import and embed:

from crowdos_mcp.server import build_server
server = build_server()
# server is a configured mcp.server.Server instance

Versioning

Follows semver. The MCP tool surface (tool names, input schemas) is stable; additive changes (new tools, new optional fields) ship as minor versions. Removing or renaming a tool is a major version.

Maintainer publish flow

# 1. Bump version in BOTH pyproject.toml and src/crowdos_mcp/__init__.py
# 2. Commit + push
# 3. Publish:
./scripts/publish.sh         # → PyPI
./scripts/publish.sh --test  # → TestPyPI (dry run)

The script verifies the two version numbers agree, runs tests, cleans dist/, builds, and uploads via twine using ~/.pypirc (needs username = __token__ + password = pypi-<token>).

License

MIT.

Issues / questions

https://github.com/bjnagent/crowd/issues

Release files for crowdos-mcp 0.11.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for crowdos-mcp 0.11.0
File Size Uploaded
crowdos_mcp-0.11.0.tar.gz 31.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for crowdos-mcp 0.11.0
File Interpreter ABI Platform
crowdos_mcp-0.11.0-py3-none-any.whl Python 3 none any Details

Total release size: 64.6 kB

Release files / crowdos_mcp-0.11.0.tar.gz

Download URL crowdos_mcp-0.11.0.tar.gz
Size 31.5 kB
Tags Source
SHA-256 checksum
How to use checksums
c49ea78f782d1c47d0b9bd768003183b13cbed534dbc83dd9d363dd1d9bfecf6
BLAKE2b-256 checksum
How to use checksums
54b059095d223bb40ae6d8702158b5977ad410947918ae50365cbd716c748c68
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release files / crowdos_mcp-0.11.0-py3-none-any.whl

Download URL crowdos_mcp-0.11.0-py3-none-any.whl
Size 33.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
0a2b9eade6981042edb08e3e76abc2726a740859e556c493e11a1275306f2673
BLAKE2b-256 checksum
How to use checksums
c18014765699f3606b58af94ab5d2853e3ec497609538841dc53f1dd4e3d0ef6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.15

Release history Release notifications | RSS feed

0.12.5

2 release files

0.12.4

2 release files

This release

0.11.0 This release

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page