Skip to main content

PALACE

License: EUPL-1.2 Python 3.13+ PyPI version CI

PALACE logo

PALACE is an open benchmark format for LLM evaluation, with native support for agentic tasks.

Benchmarks should be portable. A PALACE benchmark is a self-contained folder with two JSON files: info.json (metadata and evaluation config) and tasks.json (the actual tasks). No framework lock-in, no Python decorators, no code to run. Just data that any tool can read.

palace-eval is the reference implementation.

Quick start

1. Install:

uv tool install palace-eval   # recommended
# or: pip install palace-eval

2. Configure your API:

palace config set url https://api.openai.com/v1
palace config set key sk-your-api-key
palace config set judge_model gpt-4o

3. Run your first evaluation:

palace download SimpleQA
palace run SimpleQA -m gpt-4o -l 10

That's it. Results go to ~/.cache/palace/results/.

Run palace config anytime to see your current settings.

Why another benchmark format?

Most benchmarks today are tied to specific frameworks. LM-eval-harness tasks are Python code. Inspect-AI uses @task decorators. If you want to run a benchmark, you need their infrastructure.

PALACE takes a different approach:

  • Self-contained: everything travels together (tasks, expected outputs, judge criteria, evaluation logic)
  • HuggingFace-distributed: download with palace download, standard hosting and discovery
  • Agentic-native: tool-using agents are a first-class task type
  • Framework-agnostic: palace-eval runs them, but so could anything else

We ship 27+ benchmarks covering reasoning, knowledge, safety, multilingual, multimodal, and agentic capabilities.

The format

A PALACE benchmark is just two JSON files:

SimpleQA/
├── info.json      # what kind of benchmark, how to evaluate
└── tasks.json     # the actual tasks

info.json defines the benchmark:

{
  "name": "SimpleQA",
  "id": "openai/SimpleQA", 
  "task_type": "QA",
  "category": "Knowledge"
}

tasks.json contains the tasks:

[
  {"id": "q1", "objective": "What is the capital of France?", "expected": "Paris"},
  {"id": "q2", "objective": "Who wrote Romeo and Juliet?", "expected": "Shakespeare"}
]

The format supports five task types (QA, Classification, Criteria Evaluation, Instruction Following, and Agentic), each with its own verification logic. See the full spec.

Configuration

Palace stores settings in ~/.config/palace/config.yaml. The quick start covers the essentials.

For CI/Docker, use environment variables instead. Run palace config env to see the mapping.

Included benchmarks

Category Benchmarks
Knowledge & Reasoning SimpleQA, HotpotQA, Humanity's Last Exam, GPQA Diamond, MUSR
Academic & Math MMLU, MMLU-Pro, MATH-500, AIME 2025, BBH, HellaSwag
Multilingual MMMLU (14 languages), MGSM, Belebele (122 languages)
Long Context BABILong (4k-128k), LongBench v2
Multimodal MMMU, MMMU Pro, VLSBench
Instruction Following IFEval
Agentic GAIA, AssistantBench
palace list                    # see all available
palace download MMLU           # download one
palace download --all          # download everything

Running evaluations

palace run MMLU -m gpt-4o
palace run MMLU -m gpt-4o -l 50           # limit to 50 tasks
palace run "GPQA Diamond" -m claude-sonnet
palace run SWE-bench -m o3-mini --agentic # agentic benchmark

Python API (for integration):

from palace import evaluate

evaluate(
    run_name="my-eval",
    url="https://api.openai.com/v1",
    token="sk-...",
    name="gpt-4o",
    tasklist="SimpleQA",
    limit=100,
)

Works with any OpenAI-compatible endpoint: OpenAI, Anthropic, Azure, Mistral, local deployments, etc.

Agentic evaluation

For benchmarks where the model uses tools (GAIA, AssistantBench, etc.), you need Vivarium, a sandboxed Docker runtime for agents.

pip install vivarium-ai

Vivarium starts automatically when you run an agentic tasklist. Requires Docker 24+.

Note: Vivarium is currently pending open-source release (expected within weeks). Contact massimiliano.altieri@ec.europa.eu for early access.

Creating your own benchmarks

palace init my-benchmark      # interactive wizard
palace validate my-benchmark  # check for errors
palace publish my-benchmark   # publish to HuggingFace

See Your First Benchmark for a walkthrough.

Documentation

Full docs at palace.pages.code.europa.eu/palace-eval.

Contributing

Issues and PRs welcome. See CONTRIBUTING.md.

License

Copyright 2025 European Union. Licensed under EUPL-1.2.

Developed by the European Commission Joint Research Centre.

Citation

@software{palace2025,
  title = {PALACE: An Open Benchmark Format for LLM Evaluation},
  author = {Altieri, Massimiliano},
  year = {2025},
  institution = {European Commission Joint Research Centre},
  url = {https://code.europa.eu/palace/palace-eval},
  license = {EUPL-1.2}
}

Contact

Release files for palace-eval 1.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for palace-eval 1.1.3
File Size Uploaded
palace_eval-1.1.3.tar.gz 140.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for palace-eval 1.1.3
File Interpreter ABI Platform
palace_eval-1.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 333.6 kB

Release files / palace_eval-1.1.3.tar.gz

Download URL palace_eval-1.1.3.tar.gz
Size 140.9 kB
Tags Source
SHA-256 checksum
How to use checksums
7fd1da9731d986c0f8d72d874712cba6f8ccba28f220575a33b50f4e34076cfb
BLAKE2b-256 checksum
How to use checksums
e74549619fe73086891bb65fa2660f411a646ec5a32c83aef23eead6db1e5192
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / palace_eval-1.1.3-py3-none-any.whl

Download URL palace_eval-1.1.3-py3-none-any.whl
Size 192.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
07949a0933994039edc8cd06809872b88fef89780e783c7742fde9bb4eeb3090
BLAKE2b-256 checksum
How to use checksums
816f5be593b08cee59acb776d70b0efb7877cd40e0be6c55135a484d3e5e14ce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

1.1.4

2 release files

This release

1.1.3 This release

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.10

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page