CITAR
Civ Inspired Tool for AI Research
A Civilization V-style 4X game whose players can be language models. Play against them, watch them play each other, and measure how well they do it.
What it is
A full Civilization V ruleset — nine eras, 35 civilizations, 40 city-states, religion, social policies, great people, espionage, the United Nations, nuclear weapons — with an interface built so that a language model can sit in any seat.
Humans play in the browser. Models join through MCP (Claude Code, Claude Desktop, any MCP client), through the Anthropic API, or through any OpenAI-compatible endpoint — LM Studio, Ollama, llama.cpp, vLLM. A scripted bot is included to play against and to measure models against, so the game is fully playable with no model at all.
The rules and numbers come from UnCiv's "Civ V – Gods & Kings" ruleset, which is why they are faithful enough to be worth measuring against.
Why it exists
Most model evaluations are short. A question, an answer, a score. A game of Civilization is the opposite: hundreds of turns, imperfect information, an opponent who reacts, and consequences that arrive forty turns after the decision that caused them. That is a different thing to be good at, and it is hard to fake.
So CITAR is built as an instrument, not just a game:
- Benchmarks run full games across several models on identical maps and score them against the scripted bot, with confidence intervals and per-turn timing.
- Scenario probes put a model in a prepared position — an offer to accept, a war to decide on — and record what it does, one decision at a time, repeatably.
- Metrics record every tool call, every rejected order, every loop, and how each turn ended.
- Reports price it all: hardware time, electricity, tokens, and what a given result cost to produce.
None of it phones home. CITAR has no telemetry and no accounts unless you run a server on purpose.
Install
Just want to play
Windows — download CITAR-setup.exe and
run it. Nothing else needed; there is no Python to install.
macOS and Linux
curl -fsSL https://raw.githubusercontent.com/jprodgers/CITAR/main/install.sh | bash
Anywhere with Python 3.11+
pipx install "citar[all]" # or: pip install "citar[all]"
citar setup # finds your models and configures CITAR
citar # starts the game and opens your browser
Also on Homebrew, Scoop and winget.
Want to run a server
curl -fsSL https://raw.githubusercontent.com/jprodgers/CITAR/main/install.sh | bash -s -- --server
or with Docker:
curl -O https://raw.githubusercontent.com/jprodgers/CITAR/main/docker-compose.yml
curl -o .env https://raw.githubusercontent.com/jprodgers/CITAR/main/.env.docker.example
$EDITOR .env # domain, secret key
docker compose up -d
Either way you get accounts, invitations, single sign-on, per-user budgets, TLS, and a way for people to lend their own GPUs to the server without opening a port at home. See docs/server/DEPLOY.md.
First game in five minutes
citar # opens http://127.0.0.1:8765
Create a game, give one seat to yourself and the rest to scripted bots, and press Create. That works with nothing installed.
To give a seat to a model, pick LLM and choose a server and model — citar setup will have
found LM Studio or Ollama if either is running. To hand a seat to Claude Code instead, choose MCP
client, then copy the claude mcp add … command from Join / Seats.
The full quick start walks through all three.
What you can do with it
| Play | Full Civ V rules in the browser: tech tree, policies, religion, espionage, city-states, diplomacy with a deal builder, five victory conditions. |
| Watch | Games with no human seat can be watched with full vision, paused, slowed down, and viewed as any civilization — including each AI's recorded reasoning. |
| Benchmark | Suites of models against the same seeded maps, run sequentially or in parallel, resumable across restarts and reboots. |
| Probe | Prepared scenarios replayed case by case: offers, messages, or a whole turn, with expected outcomes and pass rates. |
| Measure | Per-seat turn times, tool mix, error and loop rates, token counts, and how turns ended. Exportable as CSV. |
| Cost | A usage ledger priced at report time, so correcting an electricity rate corrects every report ever made. |
| Edit | A map editor and a scenario editor, so you can build the position you want to test. |
| Tune | A parallel bot simulator and a resumable experiment lab, because the bot is the yardstick. |
Documentation
| Quick start | First game, first model, first benchmark |
| Installing | Every install route, and how to remove it |
| Playing | The browser client, keyboard shortcuts, every screen |
| AI players | LLM seats, MCP, providers, prompts, guard rails |
| Benchmarks | Suites, scheduling, scoring, what the numbers mean |
| Scenarios and probes | The map editor, scenarios, and repeatable decision tests |
| Servers, costs and reports | The machine registry, the usage ledger, costed reports |
| Scripted bots | How the bot plays, balancing it, the experiment lab |
| Configuration | Every environment variable, and where files live |
| Running a server | Domain, TLS, accounts, workers, backups, the runbook |
| Architecture | How the pieces fit, for contributors |
| Modding | Adding rules, units and mechanics |
| HTTP and tool API | Driving CITAR from your own code |
| Troubleshooting | When something does not work |
| FAQ | Including the honest answers about model strength |
The same pages are on the wiki and at jprodgers.github.io/CITAR.
How good are the models, actually?
Not very, yet — and that is the interesting part.
A small local model can found cities, research sensibly and hold a conversation about a trade, and will still lose to a scripted bot that has no idea what it is doing beyond a few hundred lines of heuristics. Larger models play better and cost more per turn. CITAR exists to put numbers on that rather than anecdotes, which means the measurement has to be honest about its own limits: see FAQ and KNOWN_ISSUES.md.
Contributing
Bug reports, ideas and pull requests are all welcome — see CONTRIBUTING.md. The short version:
git clone https://github.com/jprodgers/CITAR && cd CITAR
pip install -e ".[dev]"
python -m unittest discover -s tests # 272 tests, about 90 seconds
citar serve --debug
Most content is data: the ruleset is UnCiv-style JSON, and a new unit or building is a JSON entry, not code. docs/MODDING.md covers that; docs/ARCHITECTURE.md covers the rest.
Licence and credits
CITAR is licensed under the Mozilla Public License 2.0.
Its rules, numbers and much of its game logic are derived from UnCiv by Yair Morgenstern and contributors, also MPL-2.0. No UnCiv graphics, sounds or flavour text are included. Full attribution is in NOTICE.md.
Sid Meier's Civilization is a trademark of Take-Two Interactive. CITAR is an independent project, not affiliated with or endorsed by Take-Two, 2K, Firaxis Games or the UnCiv project, and contains no assets from any commercial Civilization title.
Metadata
Release files for citar 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| citar-0.1.0.tar.gz | 1.2 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| citar-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 2.4 MB
Release files / citar-0.1.0.tar.gz
| Download URL | citar-0.1.0.tar.gz |
|---|---|
| Size | 1.2 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0263227138c109d8b750ec7fcce8a1677c223f1388ba933359533df1b6b4f5e9
|
|
BLAKE2b-256 checksum How to use checksums |
ca3da6e524b8813241ecb96365414cf3898ebd2c34dd384fd4ed007fb267fb88
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / citar-0.1.0-py3-none-any.whl
| Download URL | citar-0.1.0-py3-none-any.whl |
|---|---|
| Size | 1.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0a5dd366d4460956a1560dc8024e48f8ed4a6397dcf4dcd4e5ec159ede9c41cd
|
|
BLAKE2b-256 checksum How to use checksums |
e122544382fc71ea4753d7cdbbda0b7bd7d1efbb710e342060e6b9192c9bd390
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log