Computational tools for anthropological and qualitative research: data collection and analysis as an MCP server and Python package
Project description
AI Anthropology Toolkit
A suite of AI anthropology tools for qualitative research
Overview
The AI Anthropology Toolkit provides computational tools for anthropological and qualitative research. Every component is grounded in the conventions, debates, and craft knowledge of anthropology and cognate qualitative social sciences. Epistemic stance (interpretivist, critical, STS, feminist, applied, etc.) is treated as a first-class design parameter that shapes methods, writing, and analysis.
The toolkit includes standalone notebooks for data collection and qualitative analysis, a Claude Code plugin with research lifecycle skills and agents, and an MCP server that lets Claude run the full pipeline — from data collection through coding and thematic analysis — directly.
What is AI Anthropology?
AI Anthropology is an emerging field that combines:
- Studying AI as cultural artifact — Understanding how AI systems reflect and shape human culture
- Using AI to enhance ethnographic research — Leveraging computational methods to scale qualitative analysis
- Applying anthropological insights to AI development — Bringing cultural understanding to technology design
This toolkit focuses on the second aspect: using AI to enhance traditional anthropological research methods while preserving the interpretive frameworks that make the discipline unique.
Notebooks
Standalone notebooks for computational qualitative analysis. Most can be run directly in Google Colab. Notebooks marked Local should be run on your own machine (see Running Locally below).
| Notebook | Run | Description |
|---|---|---|
| Academic Literature Explorer | Search 250M+ scholarly works across all disciplines via OpenAlex with citation counts and open access detection | |
| Qualitative Codebook Builder | Build qualitative codebooks from source literature with AI-assisted code generation, validation, and structured export | |
| Audio Transcription with Whisper | Transcribe audio and video recordings locally with Whisper — timestamped transcripts, optional speaker diarization, and Chunker-ready export | |
| Interview Transcript Semantic Chunker | Segment interview transcripts into semantically coherent chunks with speaker-aware processing and coherence scoring — fully local, no API key required | |
| Coding and Thematic Analysis | Apply codes to qualitative data and build themes using deductive, inductive, or hybrid approaches, with multi-lens parallel analysis and cross-lens comparison | |
| Text Network Analysis | Build co-occurrence networks from text with community detection, centrality metrics, and interactive visualization | |
| Topic Modeling (BERTopic) | Discover topics in text collections using transformer-based clustering with interactive visualizations and zero-shot mode | |
| Named Entity Recognition (GLiNER2) | Extract people, places, organizations, concepts, and custom entity types from text using zero-shot NER | |
| Google Books Ngram Explorer | Analyze historical word frequency patterns across Google Books corpora (1800-2022) with visualization and export | |
| Google Trends Explorer | Retrieve and visualize Google Trends data with multi-term comparison, regional breakdowns, and related queries | |
| Google News Explorer | Search Google News by keyword, time period, and country with quick or extended date-range modes | |
| Google Scholar Explorer | Local | Search Google Scholar for publications with year filtering, citation counts, and structured export |
| PubMed Literature Harvester | Search PubMed and enrich results with metadata from CrossRef, OpenAlex, and Semantic Scholar | |
| Google Patents Explorer | Search Google Patents for patent metadata including titles, inventors, assignees, and filing dates | |
| YouTube Video Search | Search YouTube and export video metadata including titles, channels, views, and durations | |
| YouTube Transcript Fetcher | Fetch YouTube video transcripts with language selection, segment chunking, and multiple export formats | |
| Podcast RSS Explorer | Pull episode metadata from any podcast RSS feed with titles, dates, durations, and structured export |
Skills
Research skills in the portable SKILL.md format that activate automatically based on context. In Claude Code, the AI Anthropology Toolkit plugin installs all of them at once; other coding agents can install any skill by copying its folder (see Using the Skills in Other Agents below).
| Skill | Description |
|---|---|
| research-question | Five-slot question grammar, evaluation rubric, genre conventions |
| methodology-selection | Method-stance compatibility, evidence need decomposition, multi-method design |
| research-plan | Ten-section plan architecture covering problem through feasibility |
| irb-protocol | 13-section protocol narratives, risk assessment, digital ethnography ethics |
| informed-consent | Consent modes (written, verbal, layered, community-based), cultural adaptation |
| grant-proposal | NSF CA-DDRIG, Wenner-Gren, Fulbright, ERC, SSHRC, Wellcome — funder-specific guidance |
| dissertation-prospectus | Section-by-section prospectus development (8-30 pages) |
| fieldwork-methods | Interview guides, observation protocols, sampling strategies, data management plans |
| qualitative-analysis | Codebook development, deductive/inductive/hybrid coding, thematic analysis, multi-lens comparison |
| research-writing | Article architecture, ethnographic craft, subfield conventions, journal requirements |
| academic-review | Peer review writing, rebuttal letters, revision strategy |
| conference-materials | AAA abstracts, slide decks, posters, speaker notes, oral delivery |
| public-engagement | Op-eds, blog posts, policy briefs, community reports, media preparation |
| job-materials | Academic CVs, cover letters, job talks, application strategy |
| career-statements | Research, teaching, and diversity statements; tenure narratives |
| teaching-materials | Syllabi, lesson plans, assignments, rubrics, discussion guides |
Using the Skills in Other Agents
Each skill is a self-contained folder (SKILL.md plus a references/ directory) in the portable format that most 2026 coding agents read. To install one outside Claude Code, clone the repository and copy the skill folder — plus skills/DESIGN.md, the shared analytical-lens reference the skills consult — into your agent's skills directory:
git clone https://github.com/MattArtzAnthro/AI-Anthropology-Toolkit.git
cp -r AI-Anthropology-Toolkit/skills/qualitative-analysis AI-Anthropology-Toolkit/skills/DESIGN.md ~/.codex/skills/
| Agent | Skills directory |
|---|---|
| Claude Code | ~/.claude/skills/ (or install the plugin — all 16 at once) |
| OpenAI Codex CLI | ~/.codex/skills/ |
| Cursor | ~/.cursor/skills/ |
| GitHub Copilot / VS Code | ~/.copilot/skills/ |
| Shared project-level | .agents/skills/ in your repository |
The skills pair naturally with the MCP server (below): descriptions route the request, the server runs the pipeline.
Agents
Autonomous Claude Code subagents that orchestrate across multiple skills for complex, multi-step tasks.
| Agent | Description |
|---|---|
| research-design | Orchestrates question, methodology, and plan skills for end-to-end research design |
| ethics-reviewer | Reviews research designs for ethics issues, drafts protocols and consent documents |
| proposal-advisor | Translates research designs into persuasive funder-specific narratives |
| fieldwork-advisor | Designs instruments tailored to specific research questions and fieldwork contexts |
| analysis-advisor | Guides qualitative coding, codebook development, and thematic analysis |
| writing-advisor | Guides article/chapter writing and R&R management |
| dissemination-advisor | Handles register translation between academic and public audiences |
| career-advisor | Coordinates application packages and course design |
Commands
| Command | Description |
|---|---|
/ai-anthropology:new-project |
Scaffold a new research project through guided phases |
/ai-anthropology:skills |
List the toolkit's skills, agents, and commands |
MCP Server
The toolkit also ships as a Python package (ai-anthropology-toolkit on PyPI) with an MCP server, so Claude (and other MCP clients) can drive the full research pipeline conversationally: data collection (OpenAlex, CrossRef, PubMed, Google Scholar, Google Trends, Google News, Google Patents, Books Ngram, YouTube search and transcripts, podcast RSS) and analysis (transcript chunking, lens-configured codebook generation, qualitative coding with per-code validation, thematic analysis, and cross-lens comparison).
Installing the Claude Code plugin (above) bundles the server automatically. It also registers in any other MCP-capable agent — the command is the same everywhere:
Claude Code
claude mcp add ai-anthropology -- uvx --from "ai-anthropology-toolkit[data]==2.2.1" ai-anthro-mcp
OpenAI Codex CLI
codex mcp add ai-anthropology -- uvx --from "ai-anthropology-toolkit[data]==2.2.1" ai-anthro-mcp
Google Gemini CLI
gemini mcp add -s user ai-anthropology uvx -- --from "ai-anthropology-toolkit[data]==2.2.1" ai-anthro-mcp
The server is model-agnostic. With ANTHROPIC_API_KEY set, analysis runs autonomously (api mode). Without it, whichever model is orchestrating — Claude, GPT, or Gemini — performs each interpretive step itself through validated work packets (delegated mode): the analysis runs on your model, the methodology and validation run on the server, and every coding decision stays visible to the researcher.
Coding Agents & Sandboxes
AI coding agents — Claude Code and Claude Desktop's Cowork, OpenAI Codex CLI, Gemini CLI — often run in sandboxed environments where MCP connectors are unavailable but pip and Python still work. The toolkit degrades gracefully across three tiers:
-
MCP tools available → use them; the full pipeline runs natively.
-
Code execution only → install the package and call the Python API directly:
pip install "ai-anthropology-toolkit[data]" python -m ai_anthro_toolkit.doctor
The doctor (
python -m ai_anthro_toolkit.doctor, also installed asai-anthro-doctor) probes every data source from the current network and reports which are reachable. Sandbox network policies typically allow the scholarly APIs (OpenAlex, CrossRef, PubMed) and block the Google/YouTube scraping endpoints — collect what is reachable and route each blocked source to local execution or its Colab notebook (the doctor prints the link). The collector functions live inai_anthro_toolkit.datasources; transcript chunking (ai_anthro_toolkit.chunking, fully local) and the 42-lens registry (ai_anthro_toolkit.lenses) work in any environment. Agent-facing instructions for this fallback chain ship in this repository as AGENTS.md and GEMINI.md. -
No code execution either → every capability runs in the browser through the Colab notebooks below.
Getting Started
Notebooks (Colab)
Click any Open in Colab badge above to run a notebook directly in your browser. Each notebook handles its own dependencies — no local installation needed.
Running Locally
Some notebooks (marked Local in the table) need to be run on your own machine. This requires Python and Jupyter.
If you already have Anaconda/Miniconda installed:
pip install scholarly
jupyter notebook
Then open the notebook file from the Jupyter file browser.
If you need to install Jupyter from scratch:
pip install jupyter scholarly
jupyter notebook
Notebooks that run locally will install any other dependencies they need automatically when you run the first cell.
Claude Code Plugin
Install the plugin in Claude Code:
/plugin marketplace add MattArtzAnthro/AI-Anthropology-Toolkit
/plugin install ai-anthropology@ai-anthropology
Skills activate automatically when Claude detects relevant context. Agents handle multi-step tasks across skills. Commands are invoked with slash syntax.
License
This repository — notebooks, Python package, MCP server, plugin content, and documentation — is licensed under the PolyForm Noncommercial License 1.0.0: free for noncommercial use, including research, education, and work by nonprofit and government research organizations, with attribution to Matt Artz appreciated. For commercial licensing, contact Matt Artz.
Citation
If you use this toolkit in your academic research, please cite:
Artz, Matt. 2025. AI Anthropology Toolkit. Software. Zenodo. https://doi.org/10.5281/zenodo.16728812
References
Artz, Matt. 2023. From Machine Learning to Machine Knowing: A Digital Anthropology Approach for the Machine Interpretation of Cultures. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000384902.
Artz, Matt. 2023. "Ten Predictions for AI and the Future of Anthropology." Anthropology News, May 8. https://doi.org/10.1111/AN.1605.
Artz, Matt. 2026. "Artificial Intelligence: The AI Anthropology Lifecycle (of, by, for AI)." In Practicing Digital Ethnography, edited by Devin Proctor. Routledge. https://doi.org/10.4324/9781032672663-29.
Artz, Matt. 2026. "Multi-Agent Ethnography: Post-Conventional Anthropological Practice Through Human-AI Collaboration." Human Organization. https://doi.org/10.1080/00664677.2026.2614501.
Artz, Matt. Forthcoming. "AI Anthropology: The Future of Applied Anthropological Practice." In Routledge Handbook of Applied Anthropology, edited by Christina Wasson, Edward B. Liebow, Karine L. Narahara, Ndukuyakhe Ndlovu, and Alaka Wali. New York: Routledge.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file ai_anthropology_toolkit-2.2.1.tar.gz.
File metadata
- Download URL: ai_anthropology_toolkit-2.2.1.tar.gz
- Upload date:
- Size: 75.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e62e4e5974e97ddac6e452950e967e2fbe0e124403dcdea36ee533df232dd45e
|
|
| MD5 |
59290e534532e104cc1778b13b588c8d
|
|
| BLAKE2b-256 |
0921c028a10e4f965940494dde3e81d1c40fecd822d2d51e132c7e65087c4aac
|
File details
Details for the file ai_anthropology_toolkit-2.2.1-py3-none-any.whl.
File metadata
- Download URL: ai_anthropology_toolkit-2.2.1-py3-none-any.whl
- Upload date:
- Size: 80.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
cb6f5e69ea398ab527f5b5df1825018e72ce7442f7e90a0cf1e67410d96fc633
|
|
| MD5 |
948961ed405e976816667ae769008ef1
|
|
| BLAKE2b-256 |
d40c2965b12c36592dff68962c358839a89844ca106aa37d4c1dad97051eb584
|