Polytopia LLM Benchmark
Project description
Polytopia LLM Benchmark
Fully automated UI benchmark for The Battle of Polytopia (Perfection mode, strict 30 actions).
It captures the screen, OCRs game state, clicks actions, and reads the final score from the game.
Install
pip install polytopia-bench
Requirements:
- Windows
- Tesseract OCR installed (set
tesseract_cmdincalibration.jsonif not on PATH)
Calibration (one-time)
Create calibration.json with your UI coordinates and grid settings. Minimal example:
{
"window_title": "The Battle of Polytopia",
"tesseract_cmd": "C:/Program Files/Tesseract-OCR/tesseract.exe",
"regions": {
"turn": [10, 10, 120, 40],
"score": [1700, 10, 200, 40],
"map": [200, 120, 1500, 800],
"end_screen": [600, 200, 700, 500]
},
"tile_grid": {
"origin": [260, 170],
"dx": 64,
"dy": 56,
"rows": 11,
"cols": 11
},
"buttons": {
"end_turn": [1780, 960],
"confirm": [1200, 900],
"close_popup": [1700, 120],
"tech_tree": [80, 960]
},
"unit_buttons": {
"warrior": [600, 900]
},
"build_buttons": {
"farm": [600, 900]
},
"tech_buttons": {
"riding": [900, 500]
},
"result_rules": {
"win_score": 1
}
}
Use examples/calibration_template.json as a starting point.
Run
polybench run --difficulty easy --opponents 1 --games 1 --calibration calibration.json
LLM command (stdin prompt → stdout JSON):
polybench run --difficulty hard --opponents 7 --games 1 --calibration calibration.json --llm-cmd "python examples\\echo_llm.py"
OpenAI-compatible HTTP endpoint:
polybench run --difficulty hard --opponents 7 --games 1 --calibration calibration.json --llm-host http://localhost:8000 --llm-model your-model --llm-api-key YOUR_KEY
Env alternative:
$env:POLYBENCH_LLM_HOST = "http://localhost:8000"
$env:POLYBENCH_LLM_MODEL = "your-model"
$env:POLYBENCH_LLM_API_KEY = "YOUR_KEY"
polybench run --difficulty hard --opponents 7 --games 1 --calibration calibration.json
Kaggle LLM Bridge (remote)
You can run the LLM inside a Kaggle notebook and call it over HTTP from your Windows machine (where Polytopia + UI automation runs).
Kaggle notebook
- Open a Kaggle notebook with
kaggle-benchmarksavailable. - Create a cell with the bridge server:
!pip -q install kaggle-benchmarks
import json
from http.server import BaseHTTPRequestHandler, HTTPServer
import kaggle_benchmarks as kbench
class Handler(BaseHTTPRequestHandler):
def do_POST(self):
if self.path != "/prompt":
self.send_response(404)
self.end_headers()
return
length = int(self.headers.get("Content-Length", "0"))
body = self.rfile.read(length) if length else b"{}"
data = json.loads(body.decode("utf-8"))
prompt = data.get("prompt", "")
model = data.get("model")
llm = kbench.llm if not model else kbench.llms[model]
response = llm.prompt(prompt)
payload = json.dumps({"content": response}).encode("utf-8")
self.send_response(200)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(payload)))
self.end_headers()
self.wfile.write(payload)
server = HTTPServer(("0.0.0.0", 8000), Handler)
server.serve_forever()
- Expose port 8000 with a tunnel (ngrok or cloudflared) and copy the public URL.
Windows run (using Kaggle LLM)
polybench run --difficulty hard --opponents 7 --games 1 --calibration calibration.json ^
--llm-provider kaggle ^
--llm-host https://YOUR_TUNNEL_URL ^
--llm-model google/gemini-2.5-flash
If --llm-model is omitted, the bridge uses Kaggle’s default kbench.llm.
Python API
import polybench
cfg = polybench.RunConfig(
difficulty="easy",
opponents=1,
games=1,
calibration_path="calibration.json",
llm_host="http://localhost:8000",
llm_model="your-model",
llm_api_key="YOUR_KEY",
)
polybench.run_benchmark(cfg)
Game API
from polybench import UIAutomationGameAPI
api = UIAutomationGameAPI("calibration.json")
api.reset("easy", 1, 1)
state = api.get_state()
# ...call your LLM and produce an action...
# api.apply_action(action, run_dir="runs/tmp", turn_index=1)
How it works
- Captures screen → OCRs
turnandscore→ samples a color grid for the map. - Builds prompt → LLM returns JSON action.
- Clicks UI to execute the action.
- Stops after 30 actions (Perfection limit), then reads the final score and writes summary.
Action schema
Allowed action types:
end_turnmoveattacktrainbuildresearch
For UI automation, unit_id and city_id must be coordinates:
{ "type": "move", "unit_id": { "x": 3, "y": 5 }, "to": { "x": 4, "y": 5 } }
Output
Each run writes:
runs/<timestamp>_<difficulty>_<opponents>/game_###/turn_###_prompt.txtruns/<timestamp>_<difficulty>_<opponents>/game_###/turn_###_response.txtruns/<timestamp>_<difficulty>_<opponents>/game_###/turn_###_action.jsonruns/<timestamp>_<difficulty>_<opponents>/game_###/turn_###_ui_log.txtruns/<timestamp>_<difficulty>_<opponents>/summary.json
Options
polybench run supports:
--difficultyeasy | normal | hard | crazy--opponents1 | 7 | 15--games(default 1)--calibrationpath to calibration.json--llm-cmdexternal command that reads prompt on stdin and returns JSON on stdout--llm-provideropenai | kaggle--llm-hostHTTP base URL (OpenAI-compatible/v1/chat/completions)--llm-modelmodel name for HTTP LLM--llm-api-keyAPI key for HTTP LLM (or setPOLYBENCH_LLM_API_KEY)--k-factorELO K (default 32)--opponent-elo(default 1000)--start-elo(default 1000)
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file polytopia_bench-0.2.1.tar.gz.
File metadata
- Download URL: polytopia_bench-0.2.1.tar.gz
- Upload date:
- Size: 15.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4d65799531e00e7b5a7b698d1bb8c6a2cce5db23337b36b7f4fd0b88dae230cb
|
|
| MD5 |
dfd7fa9d7cf8ce650916b48b54772119
|
|
| BLAKE2b-256 |
f64466d2584ee7dc7165cf12670e12c5b129848ae503c2470e148d2334bac8db
|
File details
Details for the file polytopia_bench-0.2.1-py3-none-any.whl.
File metadata
- Download URL: polytopia_bench-0.2.1-py3-none-any.whl
- Upload date:
- Size: 17.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.0
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6bade4846619ec0e054b70422876cfe8936bd690420ab61a5bdd9c361278dfa5
|
|
| MD5 |
f00ab71e6195be6502018d4ffa8c8f8c
|
|
| BLAKE2b-256 |
1199e94647b8a988c17026c62c86c17efeeb77c153d0d4d3f05859e18e3abad2
|