Skip to main content

AI Anthropology Toolkit

A suite of AI anthropology tools for qualitative research

Matt Artz | GitHub | ORCID


Overview

The AI Anthropology Toolkit provides computational tools for anthropological and qualitative research. Every component is grounded in the conventions, debates, and craft knowledge of anthropology and cognate qualitative social sciences. Epistemic stance (interpretivist, critical, STS, feminist, applied, etc.) is treated as a first-class design parameter that shapes methods, writing, and analysis.

The toolkit includes standalone notebooks for data collection and qualitative analysis, a Claude Code plugin with research lifecycle skills and agents, and an MCP server that lets Claude run the full pipeline — from data collection through coding and thematic analysis — directly.

What is AI Anthropology?

AI Anthropology is the integrated practice of studying, using, and shaping AI (Artz 2026) — an iterative lifecycle of anthropology of, by, and for AI (Artz 2026):

  • Anthropology of AI — studying AI systems as cultural phenomena: how they reflect, reshape, and redistribute human culture and power
  • Anthropology by AI — using AI and computational methods to extend ethnographic research while preserving interpretive judgment
  • Anthropology for AI — bringing anthropological insight to the design, building, and governance of AI systems

This toolkit primarily operationalizes anthropology by AI — scaling qualitative research while keeping the researcher's interpretive authority at the center — and is itself an act of anthropology for AI: purpose-built research infrastructure designed with anthropological sensibilities. Its governing design philosophy is friction by design: divergence between machine readings is sustained as data rather than resolved as noise, and the interaction itself withholds production wherever a judgment is not the machine's to make (see skills/DESIGN.md). The toolkit is presented as a demonstration of multi-agent ethnography in A Call for an AI Anthropology (General Anthropology) and Multi-Agent Ethnography: Post-Conventional Anthropological Practice Through Human-AI Collaboration (Anthropological Forum).

Notebooks

Standalone notebooks for computational qualitative research. Most run directly in Google Colab; notebooks marked Local should be run on your own machine (see Running Locally below).

Data Collection

Gather scholarly, public, and media data through official APIs and structured collectors.

Notebook Run Description
Academic Literature Explorer Open In Colab Search 250M+ scholarly works across all disciplines via OpenAlex with citation counts and open access detection
CrossRef Reference Verifier Open In Colab Verify reference lists against canonical CrossRef metadata — DOI resolution, text-vs-record comparison, retraction flags — plus journal and author queries
PubMed Literature Harvester Open In Colab Search PubMed and enrich results with metadata from CrossRef, OpenAlex, and Semantic Scholar
World Bank Data Explorer Open In Colab Retrieve development indicators from the World Bank API — 16,000+ series across 200+ economies, no API key required
UN Data Explorer Open In Colab Retrieve UN development data — SDG indicators (no key) and UNDP Human Development indices (free API key)
BLS Labor Statistics Explorer Open In Colab Retrieve US labor and economic series from the Bureau of Labor Statistics — unemployment, prices, wages — keyless
Google Books Ngram Explorer Open In Colab Analyze historical word frequency patterns across Google Books corpora (1800-2022) with visualization and export
Google Trends Explorer Open In Colab Retrieve and visualize Google Trends data with multi-term comparison, regional breakdowns, and related queries
Google News Explorer Open In Colab Search Google News by keyword, time period, and country with quick or extended date-range modes
Google Scholar Explorer Local Search Google Scholar for publications with year filtering, citation counts, and structured export
Google Patents Explorer Open In Colab Search Google Patents for patent metadata including titles, inventors, assignees, and filing dates
YouTube Video Search Open In Colab Search YouTube and export video metadata including titles, channels, views, and durations
YouTube Transcript Fetcher Open In Colab Fetch YouTube video transcripts with language selection, segment chunking, and multiple export formats
Podcast RSS Explorer Open In Colab Pull episode metadata from any podcast RSS feed with titles, dates, durations, and structured export

Qualitative Analysis Pipeline

Four notebooks cover the path from recording to themes — transcribe, segment, build a codebook from your source literature, then code and construct themes. The transcript chunker and the codebook builder both feed coding and thematic analysis. Each also works standalone.

Notebook Run Description
Audio Transcription with Whisper Open In Colab Transcribe audio and video recordings locally with Whisper — timestamped transcripts, optional speaker diarization, and Chunker-ready export
Interview Transcript Semantic Chunker Open In Colab Segment interview transcripts into semantically coherent chunks with speaker-aware processing and coherence scoring — fully local, no API key required
Qualitative Codebook Builder Open In Colab Build qualitative codebooks from source literature with AI-assisted code generation, validation, and structured export
Coding and Thematic Analysis Open In Colab Apply codes to qualitative data and build themes using deductive, inductive, or hybrid approaches, with multi-lens parallel analysis and cross-lens comparison

Multimodal & Visual Analysis

Work with images, audio, video, and mixed-media collections through Gemini's multimodal models (free API key from Google AI Studio).

Notebook Run Description
Multimodal Embedding Explorer Open In Colab Embed text, images, audio, video, and PDFs into one vector space — similarity, clustering, UMAP maps, and cross-modal search
Visual Analysis & Image Annotation Open In Colab Annotate field photos with Gemini vision — ethnographic descriptions, structured tags, and categorization

Computational Text Analysis

Standalone notebooks for analyzing text corpora at scale.

Notebook Run Description
Text Network Analysis Open In Colab Build co-occurrence networks from text with community detection, centrality metrics, and interactive visualization
Topic Modeling (BERTopic) Open In Colab Discover topics in text collections using transformer-based clustering with interactive visualizations and zero-shot mode
Named Entity Recognition (GLiNER2) Open In Colab Extract people, places, organizations, concepts, and custom entity types from text using zero-shot NER

Skills

Research skills in the portable SKILL.md format that activate automatically based on context. In Claude Code, the AI Anthropology Toolkit plugin installs all of them at once; other coding agents can install any skill by copying its folder (see Using the Skills in Other Agents below).

Skill Description
research-question Helps you turn a research interest into a clear, answerable research question. Walks through the five parts a good question needs and checks it against a rubric.
literature-review Helps you plan and write a literature review, whether narrative, scoping, or systematic. Guides your search strategy, logs what you screened and why, and helps you build the bibliography or framework you need.
methodology-selection Helps you choose and justify research methods that fit your question and your stance. Checks that your methods actually match what you are trying to find out, including when you need more than one.
research-plan Helps you write a standalone research plan. Covers every section, from the problem statement through feasibility.
irb-protocol Helps you write your IRB protocol. Covers risk assessment and the specific ethics of digital ethnography, section by section.
informed-consent Helps you design your informed consent process. Covers different consent modes, like written, verbal, layered, or community-based, and how to adapt them to your cultural context.
grant-proposal Helps you write a grant proposal for a specific funder. Covers NSF CA-DDRIG, Wenner-Gren, Fulbright, ERC, SSHRC, and Wellcome, each with its own guidance.
dissertation-prospectus Helps you write your dissertation prospectus, section by section. Works whether your prospectus is short or long.
fieldwork-methods Helps you design your fieldwork instruments. Covers interview guides, observation protocols, sampling strategy, and how you will manage your data.
qualitative-analysis Helps you analyze your qualitative data. Covers building a codebook, coding it, and pulling out themes, including comparing across more than one analytical lens.
ethnographic-generalization Helps you figure out what your confirmed findings actually generalize to. Walks through the kind of claim you can make, builds its warrant, and states its scope and how confident you can be.
digital-computational-methods Helps you design digital or computational research methods. Covers digital ethnography, platform ethics, large-scale text analysis, and working alongside AI tools.
paper-planning Helps you work out what your paper actually argues before you draft it. Extracts your claim, positions it against existing work, and sequences your argument through questions rather than writing.
research-writing Helps you write your article or dissertation chapter. Covers its structure, ethnographic craft, and the conventions of your subfield or journal.
abstract-writing Helps you write the abstract, title, and keywords for a journal or thesis submission. Builds the abstract from your manuscript's own claim, cuts it to the word limit and tells you what the cut cost, checks that it does not promise more than the manuscript delivers or name a site your anonymization protects, and picks keywords that index what the title leaves out.
academic-review Helps you write peer reviews and respond to them. Covers writing the review itself, rebuttal letters, and your revision strategy.
rival-interpretations Helps you test a claim against rival readings before a reviewer does. Argues your material from three other analytical positions, separates what they agree on from what stays genuinely open, and leaves you a record for your methods section.
manuscript-markup Helps you work through a manuscript that came back marked up with comments. Reads each comment in place, sorts them, works through them, and drafts your reply letter.
proof-review Helps you check a publisher's typeset proof against the manuscript you submitted. Compares them word by word, checks that your pseudonyms, quoted speech, and citations survived production, and writes the correction list you send back.
conference-materials Helps you prepare for a conference. Covers abstracts, slide decks, posters, speaker notes, and how you will deliver the talk.
public-engagement Helps you write for a general audience. Covers op-eds, blog posts, policy briefs, and community reports.
job-materials Helps you prepare your job market materials. Covers your CV, cover letter, job talk, and overall application strategy.
career-statements Helps you write your career statements. Covers research, teaching, and diversity statements, and tenure narratives.
teaching-materials Helps you build your course materials. Covers syllabi, lesson plans, assignments, rubrics, and discussion guides.
applied-practice Helps you write client-facing deliverables. Covers statements of work, stakeholder readouts, insight synthesis, and workshop materials.
repeated-work Helps you decide whether work you keep repeating is actually worth turning into a tool. Only if it is, helps you figure out what kind of tool it should be.
tool-building Helps you build your own research tool, like a scraper, MCP server, skill, or agent. Works out a specification with you before writing any code.

Using the Skills in Other Agents

Each skill is a self-contained folder (SKILL.md plus a references/ directory) in the portable format that most 2026 coding agents read. To install one outside Claude Code, clone the repository and copy the skill folder — plus skills/DESIGN.md, the shared analytical-lens reference the skills consult — into your agent's skills directory:

git clone https://github.com/MattArtzAnthro/AI-Anthropology-Toolkit.git
cp -r AI-Anthropology-Toolkit/skills/qualitative-analysis AI-Anthropology-Toolkit/skills/DESIGN.md ~/.codex/skills/
Agent Skills directory
Claude Code ~/.claude/skills/ (or install the plugin — all 27 at once)
OpenAI Codex CLI ~/.codex/skills/
Cursor ~/.cursor/skills/
GitHub Copilot / VS Code ~/.copilot/skills/
Shared project-level .agents/skills/ in your repository

The skills pair naturally with the MCP server (below): descriptions route the request, the server runs the pipeline.

Agents

Autonomous Claude Code subagents that orchestrate across multiple skills for complex, multi-step tasks.

Agent Description
research-design Orchestrates question, methodology, and plan skills for end-to-end research design
ethics-reviewer Reviews research designs for ethics issues, drafts protocols and consent documents
proposal-advisor Translates research designs into persuasive funder-specific narratives
fieldwork-advisor Designs instruments tailored to specific research questions and fieldwork contexts
analysis-advisor Guides qualitative coding, codebook development, and thematic analysis
writing-advisor Guides argument planning, article/chapter writing, and R&R management
dissemination-advisor Handles register translation between academic and public audiences
career-advisor Coordinates application packages and course design
tool-builder Builds a research instrument, or a skill, agent, or MCP tool for this toolkit. The only agent here that writes files

Commands

Command Description
/ai-anthropology:new-project Scaffold a new research project through guided phases
/ai-anthropology:build-tool Build a research instrument, specification first
/ai-anthropology:test-claim Test one interpretive claim against rival readings, and record what stays open
/ai-anthropology:skills List the toolkit's skills, agents, and commands

MCP Server

The toolkit also ships as a Python package (ai-anthropology-toolkit on PyPI) with an MCP server, so Claude (and other MCP clients) can drive the full research pipeline conversationally: data collection (OpenAlex, CrossRef, PubMed, Google Scholar, Google Trends, Google News, Google Patents, Books Ngram, YouTube search and transcripts, podcast RSS) and analysis (transcript chunking, lens-configured codebook generation, qualitative coding with per-code validation, thematic analysis, and cross-lens comparison).

Installing the Claude Code plugin (above) bundles the server automatically. It also registers in any other MCP-capable agent — the command is the same everywhere:

Claude Code

claude mcp add ai-anthropology -- uvx --from "ai-anthropology-toolkit[data]==3.5.0" ai-anthro-mcp

OpenAI Codex CLI

codex mcp add ai-anthropology -- uvx --from "ai-anthropology-toolkit[data]==3.5.0" ai-anthro-mcp

Google Gemini CLI

gemini mcp add -s user ai-anthropology uvx -- --from "ai-anthropology-toolkit[data]==3.5.0" ai-anthro-mcp

The server is model-agnostic. With ANTHROPIC_API_KEY set, analysis runs autonomously (api mode). Without it, whichever model is orchestrating — Claude, GPT, or Gemini — performs each interpretive step itself through validated work packets (delegated mode): the analysis runs on your model, the methodology and validation run on the server, and every coding decision stays visible to the researcher.

Coding Agents & Sandboxes

AI coding agents — Claude Code and Claude Desktop's Cowork, OpenAI Codex CLI, Gemini CLI — often run in sandboxed environments where MCP connectors are unavailable but pip and Python still work. The toolkit degrades gracefully across three tiers:

  1. MCP tools available → use them; the full pipeline runs natively.

  2. Code execution only → install the package and call the Python API directly:

    pip install "ai-anthropology-toolkit[data]"
    python -m ai_anthro_toolkit.doctor
    

    The doctor (python -m ai_anthro_toolkit.doctor, also installed as ai-anthro-doctor) probes every data source from the current network and reports which are reachable. Sandbox network policies typically allow the scholarly APIs (OpenAlex, CrossRef, PubMed) and block the Google/YouTube scraping endpoints — collect what is reachable and route each blocked source to local execution or its Colab notebook (the doctor prints the link). The collector functions live in ai_anthro_toolkit.datasources; transcript chunking (ai_anthro_toolkit.chunking, fully local) and the 42-lens registry (ai_anthro_toolkit.lenses) work in any environment. Agent-facing instructions for this fallback chain ship in this repository as AGENTS.md and GEMINI.md.

  3. No code execution either → every capability runs in the browser through the Colab notebooks below.

Checking an artifact

ai-anthro-check path/to/codebook.json      # or: python -m ai_anthro_toolkit.checks <path>

Runs the standing checks over a codebook or a coded dataset and reports what they found. Most of these run on their own, without being asked, whenever the toolkit produces one of those artifacts; the command exists so you can re-run them later, on an artifact you were handed, without the pipeline that made it.

A check that fires is a question rather than a verdict. It names a commitment the artifact implies — that codes were meant to be mutually exclusive, that every code in the book was meant to earn its place, that the codebook did not move mid-pass — and only you can say whether that commitment is yours. A check that could not run is reported as unrun, never as passed, and a run in which nothing capable of surprising you executed says so.

When you build an instrument through tool-building, it writes instrument-checks.json beside your data from the commitments you settled while specifying it — what an empty result means, whether a repeated identifier is an error, whether a field is required. ai-anthro-check picks that file up automatically, so the checks are about your artifact rather than only the toolkit's.

The file says what it cannot do. A commitment that cannot be checked from the data alone — whether source order was preserved, for instance, which only the source can settle — is listed as unenforceable with the reason. If you settled five commitments and three became checks, it tells you which two did not and why.

One flag matters: --distinct-codes if you hold codes to be mutually exclusive, --overlapping-codes if you keep overlapping codes deliberately, as grounded theory and several interpretive traditions do. Left unsaid, the check that depends on it does not run, because that is a methodological commitment and not the toolkit's to assume.

Releasing

Maintainer checklist for cutting a release, covering the two version tracks, the upload-before-push ordering, and how to verify a release actually resolves: RELEASING.md.

Companion Plugins

Where a specialized tool already owns a capability, the toolkit hands off to it instead of duplicating it here.

Plugin Capability Install
gephi-network-analysis Network analysis in live Gephi Desktop — text-network construction, layouts, centrality and community metrics, structural claim verification claude plugin marketplace add MattArtzAnthro/gephi-ai, then claude plugin install gephi-network-analysis@gephi-ai

With the companion installed, the digital-computational-methods skill hands network-analysis execution to it. Without it, the skill falls back to the Text Network Analysis notebook and a GEXF export for manual work in Gephi.

Getting Started

Notebooks (Colab)

Click any Open in Colab badge above to run a notebook directly in your browser. Each notebook handles its own dependencies — no local installation needed.

Running Locally

Some notebooks (marked Local in the table) need to be run on your own machine. This requires Python and Jupyter.

If you already have Anaconda/Miniconda installed:

pip install scholarly
jupyter notebook

Then open the notebook file from the Jupyter file browser.

If you need to install Jupyter from scratch:

pip install jupyter scholarly
jupyter notebook

Notebooks that run locally will install any other dependencies they need automatically when you run the first cell.

Claude Code Plugin

Install the plugin in Claude Code:

/plugin marketplace add MattArtzAnthro/AI-Anthropology-Toolkit
/plugin install ai-anthropology@ai-anthropology

Skills activate automatically when Claude detects relevant context. Agents handle multi-step tasks across skills. Commands are invoked with slash syntax.

License

This repository — notebooks, Python package, MCP server, plugin content, and documentation — is licensed under the PolyForm Noncommercial License 1.0.0: free for noncommercial use, including research, education, and work by nonprofit and government research organizations, with attribution to Matt Artz appreciated. For commercial licensing, contact Matt Artz.

Citation

If you use this toolkit in your academic research, please cite:

Artz, Matt. 2025. AI Anthropology Toolkit. Software. Zenodo. https://doi.org/10.5281/zenodo.16728812

Research Behind the Toolkit

The toolkit operationalizes a published research program on AI Anthropology:

Artz, Matt. 2026. "A Call for an AI Anthropology." General Anthropology 33(1): 23–28. https://doi.org/10.1111/gena.70007.

Artz, Matt. 2026. "Multi-Agent Ethnography: Post-Conventional Anthropological Practice Through Human-AI Collaboration." Anthropological Forum. https://doi.org/10.1080/00664677.2026.2614501.

Artz, Matt. 2026. "Artificial Intelligence: The AI Anthropology Lifecycle (of, by, for AI)." In Practicing Digital Ethnography, edited by Devin Proctor. Routledge. https://doi.org/10.4324/9781032672663-29.

Koycheva, Lora, Angela K. VandenBroek, and Matt Artz, eds. 2026. Anthropology and AI. New York: Routledge. https://doi.org/10.4324/9781003532750.

Artz, Matt. 2023. From Machine Learning to Machine Knowing: A Digital Anthropology Approach for the Machine Interpretation of Cultures. UNESCO. https://unesdoc.unesco.org/ark:/48223/pf0000384902.

Artz, Matt. 2023. "Ten Predictions for AI and the Future of Anthropology." Anthropology News, May 8. https://doi.org/10.1111/AN.1605.

Artz, Matt. Forthcoming. "AI Anthropology: The Future of Applied Anthropological Practice." In Routledge Handbook of Applied Anthropology, edited by Christina Wasson, Edward B. Liebow, Karine L. Narahara, Ndukuyakhe Ndlovu, and Alaka Wali. New York: Routledge.

Release files for ai-anthropology-toolkit 3.6.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for ai-anthropology-toolkit 3.6.0
File Size Uploaded
ai_anthropology_toolkit-3.6.0.tar.gz 110.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for ai-anthropology-toolkit 3.6.0
File Interpreter ABI Platform
ai_anthropology_toolkit-3.6.0-py3-none-any.whl Python 3 none any Details

Total release size: 226.4 kB

Release files / ai_anthropology_toolkit-3.6.0.tar.gz

Download URL ai_anthropology_toolkit-3.6.0.tar.gz
Size 110.3 kB
Tags Source
SHA-256 checksum
How to use checksums
c7c2db20dca2b629e89fee619772826b116e77c6bc505f6add8c766a0a4d1e49
BLAKE2b-256 checksum
How to use checksums
63216a2910a9911de71c82180e641c10a8b3925c32d3e1eb56214102d1693968
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release files / ai_anthropology_toolkit-3.6.0-py3-none-any.whl

Download URL ai_anthropology_toolkit-3.6.0-py3-none-any.whl
Size 116.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cd3c0fd9814612fe673cc74cfefe2b0fe53c6ea120bb93c9f7df995d4a28af06
BLAKE2b-256 checksum
How to use checksums
6f041d2a28af929234d2efca80489543dd11b35740eea03b10ae7213b444d9d9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.5

Release history Release notifications | RSS feed

3.8.1

2 release files

3.8.0

2 release files

3.7.0

2 release files

This release

3.6.0 This release

2 release files

3.4.0

2 release files

3.3.0

2 release files

3.2.0

2 release files

3.1.0

2 release files

3.0.0

2 release files

2.2.3

2 release files

2.2.2

2 release files

2.2.1

2 release files

2.2.0

2 release files

2.1.1

2 release files

2.1.0

2 release files

2.0.1

2 release files

2.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page