Skip to main content

PALACE

License: EUPL-1.2 Python 3.13+ PyPI version CI

PALACE logo

PALACE is an open benchmark format for LLM evaluation, with native support for agentic tasks.

Benchmarks should be portable. A PALACE benchmark is a self-contained folder with two JSON files: info.json (metadata and evaluation config) and tasks.json (the actual tasks). No framework lock-in, no Python decorators, no code to run. Just data that any tool can read.

palace-eval is the reference implementation.

Quick start

1. Install:

uv tool install palace-eval   # recommended
# or: pip install palace-eval

2. Configure your API:

palace config set url https://api.openai.com/v1
palace config set key sk-your-api-key
palace config set judge_model gpt-4o

3. Run your first evaluation:

palace download SimpleQA
palace run SimpleQA -m gpt-4o -l 10

That's it. Results go to ~/.cache/palace/results/.

Run palace config anytime to see your current settings.

Why another benchmark format?

Most benchmarks today are tied to specific frameworks. LM-eval-harness tasks are Python code. Inspect-AI uses @task decorators. If you want to run a benchmark, you need their infrastructure.

PALACE takes a different approach:

  • Self-contained: everything travels together (tasks, expected outputs, judge criteria, evaluation logic)
  • HuggingFace-distributed: download with palace download, standard hosting and discovery
  • Agentic-native: tool-using agents are a first-class task type
  • Framework-agnostic: palace-eval runs them, but so could anything else

We ship 27+ benchmarks covering reasoning, knowledge, safety, multilingual, multimodal, and agentic capabilities.

The format

A PALACE benchmark is just two JSON files:

SimpleQA/
├── info.json      # what kind of benchmark, how to evaluate
└── tasks.json     # the actual tasks

info.json defines the benchmark:

{
  "name": "SimpleQA",
  "id": "openai/SimpleQA", 
  "task_type": "QA",
  "category": "Knowledge"
}

tasks.json contains the tasks:

[
  {"id": "q1", "objective": "What is the capital of France?", "expected": "Paris"},
  {"id": "q2", "objective": "Who wrote Romeo and Juliet?", "expected": "Shakespeare"}
]

The format supports five task types (QA, Classification, Criteria Evaluation, Instruction Following, and Agentic), each with its own verification logic. See the full spec.

Configuration

Palace stores settings in ~/.config/palace/config.yaml. The quick start covers the essentials.

For CI/Docker, use environment variables instead. Run palace config env to see the mapping.

Included benchmarks

Category Benchmarks
Knowledge & Reasoning SimpleQA, HotpotQA, Humanity's Last Exam, GPQA Diamond, MUSR
Academic & Math MMLU, MMLU-Pro, MATH-500, AIME 2025, BBH, HellaSwag
Multilingual MMMLU (14 languages), MGSM, Belebele (122 languages)
Long Context BABILong (4k-128k), LongBench v2
Multimodal MMMU, MMMU Pro, VLSBench
Instruction Following IFEval
Agentic GAIA, AssistantBench
palace list                    # see all available
palace download MMLU           # download one
palace download --all          # download everything

Running evaluations

palace run MMLU -m gpt-4o
palace run MMLU -m gpt-4o -l 50           # limit to 50 tasks
palace run "GPQA Diamond" -m claude-sonnet
palace run SWE-bench -m o3-mini --agentic # agentic benchmark

Python API (for integration):

from palace import evaluate

evaluate(
    run_name="my-eval",
    url="https://api.openai.com/v1",
    token="sk-...",
    name="gpt-4o",
    tasklist="SimpleQA",
    limit=100,
)

Works with any OpenAI-compatible endpoint: OpenAI, Anthropic, Azure, Mistral, local deployments, etc.

Agentic evaluation

For benchmarks where the model uses tools (GAIA, AssistantBench, etc.), you need Vivarium, a sandboxed Docker runtime for agents.

pip install vivarium-ai

Vivarium starts automatically when you run an agentic tasklist. Requires Docker 24+.

Note: Vivarium is currently pending open-source release (expected within weeks). Contact massimiliano.altieri@ec.europa.eu for early access.

Creating your own benchmarks

palace init my-benchmark      # interactive wizard
palace validate my-benchmark  # check for errors
palace publish my-benchmark   # publish to HuggingFace

See Your First Benchmark for a walkthrough.

Documentation

Full docs at palace.pages.code.europa.eu/palace-eval.

Contributing

Issues and PRs welcome. See CONTRIBUTING.md.

License

Copyright 2025 European Union. Licensed under EUPL-1.2.

Developed by the European Commission Joint Research Centre.

Citation

@software{palace2025,
  title = {PALACE: An Open Benchmark Format for LLM Evaluation},
  author = {Altieri, Massimiliano},
  year = {2025},
  institution = {European Commission Joint Research Centre},
  url = {https://code.europa.eu/palace/palace-eval},
  license = {EUPL-1.2}
}

Contact

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

palace_eval-1.0.7.tar.gz (135.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

palace_eval-1.0.7-py3-none-any.whl (187.4 kB view details)

Uploaded Python 3

File details

Details for the file palace_eval-1.0.7.tar.gz.

File metadata

  • Download URL: palace_eval-1.0.7.tar.gz
  • Upload date:
  • Size: 135.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.15

File hashes

Hashes for palace_eval-1.0.7.tar.gz
Algorithm Hash digest
SHA256 94da5349faa161031f3a592a77d2d7296e661a8a2e1c1e036cd60a67a423a16d
MD5 a7993521032e87e13b107362c025f428
BLAKE2b-256 a1f286dc50c36f906107721fbfd74ae6b9d07ae98c9785e80c8f79ea100564f4

See more details on using hashes here.

File details

Details for the file palace_eval-1.0.7-py3-none-any.whl.

File metadata

  • Download URL: palace_eval-1.0.7-py3-none-any.whl
  • Upload date:
  • Size: 187.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.13.15

File hashes

Hashes for palace_eval-1.0.7-py3-none-any.whl
Algorithm Hash digest
SHA256 00074fa4fbc4545a59e1fe9646c03d0911a8ab650cc65cb7a95b546f5c338851
MD5 a5b81a607350e1a1ff84e41475b665ad
BLAKE2b-256 149a320f01a6159431825abffc99b62f7b8ed51502bd195c490094ff598bfafb

See more details on using hashes here.

Release history Release notifications | RSS feed

1.0.9

2 files

1.0.8

2 files

This release

1.0.7 This release

2 files

1.0.6

2 files

1.0.5

2 files

1.0.4

2 files

1.0.3

2 files

1.0.2

2 files

1.0.1

2 files

1.0.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page