Skip to main content

Potato: The Portable Annotation Tool

Docs & Guides Technical Reference PyPI License Paper (Potato 2.0) Paper (Potato 1.0) Live Demo Website

Potato is a free, self-hosted annotation platform for NLP, Agentic, GenAI, and qualitative research. Annotate text, audio, video, images, documents, agent traces, and more — or run a full qualitative data analysis (QDA) workflow with a living codebook, memos, and cases. Configured entirely through YAML. No coding required.

Try the live demo on HuggingFace Spaces — no installation needed. More at www.potatoannotator.com.


Quick Start

pip install potato-annotation
# The examples/ folder ships with the source repo (see "run from source" below).
# After a PyPI install, clone the repo for the examples, or point `potato start`
# at your own config (see docs/quick-start.md).
potato start examples/classification/single-choice/config.yaml -p 8000

Or run from source (recommended to get the examples/):

git clone https://github.com/davidjurgens/potato.git
cd potato && pip install -r requirements.txt
python potato/flask_server.py start examples/classification/single-choice/config.yaml -p 8000

Open http://localhost:8000 and start annotating. Browse the examples/ directory for ready-to-use templates.


What Can You Annotate?

Potato covers traditional NLP labeling, AI agent evaluation, and interpretive qualitative analysis.

The tables below are a representative sample, not a complete list. Schemes and data types compose freely, custom layouts and raw HTML let you build interfaces beyond these, and new schema types can be added. If you don't see your task here, it's likely still possible.

Data Types

Modality Capabilities
Text Classification, span labeling, entity linking, coreference, pairwise comparison (docs)
Agent Traces Step-by-step evaluation of LLM agents, tool calls, ReAct chains, and multi-agent systems (docs)
Web Agents Screenshot-based review with SVG click/scroll overlays, or live browsing with automatic trace recording (docs)
RAG Pipelines Retrieval relevance, answer faithfulness, citation accuracy, hallucination detection
Audio Waveform visualization, segment labeling, ELAN-style tiered annotation, and 21 transcript/subtitle formats read directly — Whisper, cloud ASR, SRT/VTT, YouTube captions, TextGrid/EAF (docs, transcripts)
Video Frame-by-frame labeling, temporal segments, playback sync (docs)
Images Bounding boxes, polygons, landmarks, classification (docs)
Dialogue Turn-level annotation, conversation trees, interactive chat evaluation, and diarized transcripts synced to their audio (docs)
Documents PDF, Word, Markdown, code, and spreadsheets with coordinate mapping (docs)

Annotation Schemes

Scheme Use Case
Radio / Checkbox / Likert Classification, multi-label, rating scales
Span annotation NER, highlighting, hallucination marking
Pairwise comparison A/B testing, best-worst scaling
Per-step ratings Evaluate individual agent actions or dialogue turns
Free text Open-ended responses with validation
Triage Rapid accept/reject/skip curation (docs)
Conditional logic Adaptive forms that respond to prior answers (docs)

Agent & LLM Evaluation

Potato reads traces from the major agent frameworks and evaluates them at the trajectory, step, span, or comparison level.

Trace Formats

Import traces from any major agent framework with the built-in converter:

python -m potato.trace_converter --input traces.json --input-format openai --output data.jsonl

Supported formats: OpenAI, Anthropic/Claude, ReAct, LangChain, LangFuse, WebArena, SWE-bench, OpenTelemetry, CrewAI/AutoGen/LangGraph, MCP, Aider, Claude Code, ATIF, SWE-Agent, and Web Agent. Auto-detection is available with --auto-detect.

Evaluation Levels

Level What You Annotate Example
Trajectory Overall task success, efficiency, safety "Did the agent complete the task?"
Step Individual action correctness, reasoning quality Per-turn Likert ratings on each agent step
Span Specific text segments within agent output Highlight hallucinated claims, factual errors
Comparison Side-by-side A/B agent evaluation "Which agent performed better?"

Web Agent Viewer

A viewer for GUI agent traces. You step through the screenshots one action at a time, with SVG overlays for clicks, bounding boxes, mouse paths, and scrolls. Annotators rate each step with inline controls, and a filmstrip bar jumps between steps.

Ready-to-Use Agent Examples

Example What It Evaluates
agent-trace-evaluation Text agent traces with MAST error taxonomy + hallucination spans
visual-agent-evaluation GUI agents with screenshot grounding accuracy
agent-comparison Side-by-side A/B agent comparison
rag-evaluation RAG retrieval relevance and citation accuracy
openai-evaluation OpenAI Chat API traces with tool calls
anthropic-evaluation Claude messages with tool_use blocks
swebench-evaluation Coding agents with patch correctness ratings
multi-agent-evaluation Multi-agent coordination (CrewAI, AutoGen, LangGraph)
web-agent-review Pre-recorded web traces with step-by-step overlay viewer
web-agent-creation Live web browsing with automatic trace recording

Qualitative Data Analysis (QDA)

Potato also supports interpretive qualitative research, the kind of work done in NVivo, ATLAS.ti, or MAXQDA, self-hosted and free.

Capability Description
Living codebook The codebook is an evolving markdown document of rules, definitions, examples, and rationales rather than a list of labels. Edit it in a full-page document view or inline while coding, with versioning, diff, and restore; semantic edits can re-flag affected excerpts for review (docs)
In-vivo coding Create codes directly from a highlighted passage, in the participant's own words (example)
Memos Attach analytic notes to excerpts, codes, or the whole project as your interpretation develops (docs)
Cases Group instances into units of analysis — participants, interviews, documents, sites — for case-based comparison (docs)
Search Full-text search across your corpus and annotations to find, revisit, and code recurring patterns (docs)
Codebook distillation Turn the human-authored codebook into an LLM prompt for AI-assisted coding

Enable it with qda_mode, which turns these features on together; see the QDA Mode guide and the runnable qda-mode-example.


AI-Powered Annotation

LLM Label Suggestions

Connect any LLM provider to pre-annotate instances and suggest labels. Annotators then review and correct rather than labeling from scratch, which is the point: correcting a draft is faster than producing one.

Supported backends: OpenAI, Anthropic, Ollama, vLLM, Gemini, HuggingFace, OpenRouter

Active Learning

Potato reorders your annotation queue based on model uncertainty so annotators label the most informative instances first. Supports uncertainty sampling, BADGE, BALD, diversity, and hybrid strategies (docs).

Solo Mode

A human-LLM collaborative workflow where the system learns from annotator feedback and progressively transitions to autonomous LLM labeling as agreement improves (docs).

Chat Assistant

An LLM-powered sidebar where annotators can ask questions about difficult instances. It answers from your task description and annotation guidelines, and never writes a label itself (docs).


Quality Control & Workflows

Quality Assurance

Feature Description
Attention checks Automatically inserted known-answer items to verify engagement
Gold standards Track annotator accuracy against expert labels
Inter-annotator agreement Krippendorff's alpha (general) and Cohen's kappa (step-level agent evaluation)
Training phase Practice annotations with feedback before the real task
Behavioral tracking Timing, click patterns, and annotation change history
Psychometrics Live IRT (multiclass GLAD, Whitehill et al. 2009) fit as annotations arrive: per-item label posteriors, annotator ability with standard errors, item difficulty, and discrimination flags for codebook bugs — no gold labels, no LLM (docs)
Boundary probing Counterfactual probes map each annotator's decision boundary; paraphrase-invariance flags inconsistency (docs)
Truth Serum Surprisingly-popular scoring (Prelec et al., Nature 2017): gold-free verdicts that beat majority vote on hard items, plus annotator calibration (docs)
Paper Mode python -m potato.paper config.yaml emits a compilable LaTeX dataset report — methods paragraphs, booktabs tables, IAA, limitations — ready to cut-paste (docs)
Think-Aloud Mode Speak while you annotate: fully-local speech-to-text, verbatim rationale streams, labels committed by voice via rule-based phrase detection — no LLM (docs)

Annotation Workflows

Workflow Description
Multi-annotator Multiple annotators per item with overlap control and agreement metrics
Adjudication Expert review of annotator disagreements to produce gold labels (docs)
Solo mode Human-LLM collaboration with progressive automation (docs)
Crowdsourcing Prolific and MTurk integration with platform-specific auth (docs)
Triage Rapid accept/reject/skip for data curation (docs)
Multiplayer Rooms Live shared sessions: norming (blind vote → reveal → discuss, with blind vs. post-discussion α), huddle (walk current disagreements together), and shadow (trainees watch the host annotate) (docs)
Pocket Mode Annotate from your phone: installable PWA with a card-stack UI, one-tap labeling, and offline annotation that syncs on reconnect (docs)

Continuous Evaluation Loop

Close the loop from production traces to graded, regression-gated evaluation:

Capability Description
Capture Instrument any agent with the @traceable tracing SDK, or POST traces to the ingestion webhook
Automate Rules (filter → sample → actions) route incoming traces to queues, datasets, evaluators, or webhooks
Curate Versioned datasets & experiments + semantic search/slices to find what to review
Evaluate Programmatic evaluators (trajectory match, tool-use, LLM-judge, heuristics) + a side-by-side model arena
Gate Run evals in pytest and fail CI on score-threshold regressions
Calibrate LLM-judge ↔ human alignment with auto-calibration from human corrections; judges categorical, span, and free-text outputs

Authentication & Deployment

Potato supports six authentication methods:

Method Use Case
In-memory Local development, quick studies
Password + file persistence Team annotation with shared credential files (docs)
Database Production deployments with SQLite or PostgreSQL (docs)
OAuth / SSO Google, GitHub, or institutional OIDC login (docs)
Clerk Managed authentication via Clerk.com (docs)
Passwordless Low-stakes tasks where ease of access matters (docs)

Passwords are hashed with per-user PBKDF2-SHA256 salts. Admins can reset passwords via CLI (potato reset-password) or REST API. Self-service token-based reset is also available.


Example Projects

Ready-to-use templates organized by type in examples/:

Category Examples
Classification Radio, checkbox, Likert, slider, pairwise comparison
Span NER, span linking, coreference, entity linking
Agent Traces LLM agents, web agents, RAG, multi-agent, code agents
Audio Waveform annotation, classification, ELAN-style tiered
Video Frame-level labeling, temporal segments
Image Bounding boxes, PDF/document annotation
Advanced Solo mode, adjudication, quality control, conditional logic
QDA Qualitative analysis: living codebook, in-vivo coding, memos, cases
AI-Assisted LLM suggestions, Ollama integration
Custom Layouts Content moderation, dialogue QA, medical review

Live Demos on HuggingFace

Try Potato in your browser — no installation. A growing catalog of one-click demo Spaces covers classification, span/NER, agent-trace evaluation, multimodal, QDA, and more:

Research Showcase

The Potato Showcase contains annotation projects from published research — sentiment analysis, dialogue evaluation, summarization, and more.


Documentation

Potato has two complementary doc sites: potatoannotator.com/docs for guides, tutorials, and higher-level walkthroughs, and Read the Docs for the complete, version-matched technical reference (every config option, the full HTTP API, and internals). The links below point to the guide pages.

Topic Link
Quick Start docs/quick-start.md
Configuration Reference docs/configuration/configuration.md
Schema Gallery docs/annotation-types/schemas_and_templates.md
Agent Trace Evaluation docs/agent-evaluation/agent_traces.md
Web Agent Annotation docs/agent-evaluation/web_agent_annotation.md
Datasets & Experiments docs/agent-evaluation/datasets_and_experiments.md
Programmatic Evaluators docs/agent-evaluation/evaluators.md
Automation Rules docs/agent-evaluation/automation_rules.md
CI Evaluation (pytest gating) docs/agent-evaluation/ci_evaluation.md
Model Arena docs/agent-evaluation/model_arena.md
Semantic Curation (Catalog) docs/agent-evaluation/semantic_curation.md
Tracing SDK (potato_trace) docs/integrations/tracing_sdk.md
AI Support docs/ai-intelligence/ai_support.md
Using HuggingFace Models docs/ai-intelligence/huggingface_models.md
Potato on HuggingFace docs/data-export/potato_on_huggingface.md
Active Learning docs/ai-intelligence/active_learning_guide.md
Solo Mode docs/solo-mode/solo_mode.md
Qualitative Data Analysis (QDA) docs/advanced/qda.md
Quality Control docs/workflow/quality_control.md
Password Management docs/auth-users/password_management.md
SSO & OAuth docs/auth-users/sso_authentication.md
Admin Dashboard docs/administration/admin_dashboard.md
Crowdsourcing docs/deployment/crowdsourcing.md
Export Formats docs/data-export/export_formats.md
Full Documentation Index docs/index.md

For coding agents

If you point Claude Code, Codex, or Cursor at Potato, give it these generated, machine-checkable specs rather than prose — they are built from the running code, so they cannot drift from it.

Artifact What it gives you
llms.txt Curated index of the docs (llms.txt standard)
llms-full.txt Every documentation page in one file
Config JSON Schema All 159 config keys, 61 annotation types, 24 display types — validates a config.yaml before the server runs
OpenAPI 3.1 spec All 419 HTTP paths, with per-operation auth and config gating

Every config in examples/ carries a # yaml-language-server: $schema=… modeline, so editors validate it live. See Machine-Readable Specs for editor setup, CI validation, and jq recipes.


Development

# Run tests
pytest tests/ -v

# By category
pytest tests/unit/ -v        # Unit tests (fast)
pytest tests/server/ -v      # Integration tests
pytest tests/selenium/ -v    # Browser tests

# With coverage
pytest --cov=potato --cov-report=html

See the Testing guide for which tier to write in, the test-file security rules, the annotation-persistence testing pattern, and the drift tests that keep the generated specs honest.


Support


License

Potato is free software, licensed under the GNU General Public License v3.0 or later (GPLv3+). You are free to use, study, modify, and redistribute it — including for commercial purposes — provided that any distributed derivative works are also licensed under the GPLv3+ and made available with their source code. See the LICENSE file for the full terms.


Citation

If you use Potato in your research, please cite the Potato 2.0 paper (ACL 2026 System Demonstrations):

@inproceedings{jurgens-etal-2026-potato,
    title = "Potato 2.0: A Comprehensive Annotation Platform with {AI}-in-the-Loop Support",
    author = "Jurgens, David  and
      Chen, Michael  and
      Iyer, Lina",
    editor = "Durrett, Greg  and
      Jian, Ping",
    booktitle = "Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.acl-demo.37/",
    pages = "374--386",
    ISBN = "979-8-89176-392-0",
}

To reference the original Potato release, cite the Potato 1.0 paper (EMNLP 2022 System Demonstrations):

@inproceedings{pei-etal-2022-potato,
    title = "{POTATO}: The Portable Text Annotation Tool",
    author = "Pei, Jiaxin  and
      Ananthasubramaniam, Aparna  and
      Wang, Xingyao  and
      Zhou, Naitian  and
      Dedeloudis, Apostolos  and
      Sargent, Jackson  and
      Jurgens, David",
    editor = "Che, Wanxiang  and
      Shutova, Ekaterina",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-demos.33/",
    doi = "10.18653/v1/2022.emnlp-demos.33",
    pages = "327--337",
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

potato_annotation-2.8.0.tar.gz (6.6 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

potato_annotation-2.8.0-py3-none-any.whl (7.5 MB view details)

Uploaded Python 3

File details

Details for the file potato_annotation-2.8.0.tar.gz.

File metadata

  • Download URL: potato_annotation-2.8.0.tar.gz
  • Upload date:
  • Size: 6.6 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.11.5

File hashes

Hashes for potato_annotation-2.8.0.tar.gz
Algorithm Hash digest
SHA256 926b44f6aefdfdffb95d2630e996fbfc293c4b39e5361a3c260074b0d0710cfd
MD5 bd9caf9d430aca849c74ec31c8a2a0ca
BLAKE2b-256 a4552aed2fceb50074e2859963275a2793cbbe11599699edb188126a796f8429

See more details on using hashes here.

File details

Details for the file potato_annotation-2.8.0-py3-none-any.whl.

File metadata

File hashes

Hashes for potato_annotation-2.8.0-py3-none-any.whl
Algorithm Hash digest
SHA256 da8cfd729d5852f26a22f997ab2e3bd1dab860de8ac6d8dc4af76c547d1efc8e
MD5 fae844f8735710bfb69feecb251b2fa6
BLAKE2b-256 3278d159a50a50e3726ade5a84f5599e8edc704bcac3921afb4cb5542a2e5b23

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

2.8.0 This release

2 files

2.7.1

2 files

2.7.0

2 files

2.6.2

2 files

2.6.0

2 files

2.5.0

2 files

2.4.5

2 files

2.4.4

2 files

2.4.3

2 files

2.4.1

2 files

2.4.0

2 files

2.3.0

2 files

2.2.0

2 files

2.1.0

2 files

2.0.2

2 files

2.0.1

2 files

2.0.0

2 files

1.2.2.4

2 files

1.2.2.3

2 files

1.2.2.2

2 files

1.2.2.1

2 files

1.2.2.0

2 files

1.2.1.8

2 files

1.2.1.7

2 files

1.2.1.6

2 files

1.2.1.5

2 files

1.2.1.4

2 files

1.2.1.3

2 files

1.2.1.2

2 files

1.2.1.1

2 files

1.2.1.0

2 files

1.2.0.32

1 file

1.2.0.31

2 files

1.2.0.30

2 files

1.2.0.29

2 files

1.2.0.28

2 files

1.2.0.27

2 files

1.2.0.26

2 files

1.2.0.25

2 files

1.2.0.24

2 files

1.2.0.23

2 files

1.2.0.21

2 files

1.2.0.20

2 files

1.2.0.19

2 files

1.2.0.18

2 files

1.2.0.17

2 files

1.2.0.16

2 files

1.2.0.15

2 files

1.2.0.14

2 files

1.2.0.13

2 files

1.2.0.12

2 files

1.2.0.11

2 files

1.2.0.10

2 files

1.2.0.9

2 files

1.2.0.8

2 files

1.2.0.7

2 files

1.2.0.6

2 files

1.2.0.5

2 files

1.2.0.4

2 files

1.2.0.3

2 files

1.2.0.2

2 files

1.2.0.1

2 files

1.2.0

3 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page