Skip to main content

PALACE

License: EUPL-1.2 Python 3.13+ PyPI version CI

PALACE logo

PALACE is an open benchmark format for LLM evaluation, with native support for agentic tasks.

Benchmarks should be portable. A PALACE benchmark is a self-contained folder with two JSON files: info.json (metadata and evaluation config) and tasks.json (the actual tasks). No framework lock-in, no Python decorators, no code to run. Just data that any tool can read.

palace-eval is the reference implementation.

Quick start

1. Install:

uv tool install palace-eval   # recommended
# or: pip install palace-eval

2. Configure your API:

palace config set url https://api.openai.com/v1
palace config set key sk-your-api-key
palace config set judge_model gpt-4o

3. Run your first evaluation:

palace download SimpleQA
palace run SimpleQA -m gpt-4o -l 10

That's it. Results go to ~/.cache/palace/results/.

Run palace config anytime to see your current settings.

Why another benchmark format?

Most benchmarks today are tied to specific frameworks. LM-eval-harness tasks are Python code. Inspect-AI uses @task decorators. If you want to run a benchmark, you need their infrastructure.

PALACE takes a different approach:

  • Self-contained: everything travels together (tasks, expected outputs, judge criteria, evaluation logic)
  • HuggingFace-distributed: download with palace download, standard hosting and discovery
  • Agentic-native: tool-using agents are a first-class task type
  • Framework-agnostic: palace-eval runs them, but so could anything else

We ship 27+ benchmarks covering reasoning, knowledge, safety, multilingual, multimodal, and agentic capabilities.

The format

A PALACE benchmark is just two JSON files:

SimpleQA/
├── info.json      # what kind of benchmark, how to evaluate
└── tasks.json     # the actual tasks

info.json defines the benchmark:

{
  "name": "SimpleQA",
  "id": "openai/SimpleQA", 
  "task_type": "QA",
  "category": "Knowledge"
}

tasks.json contains the tasks:

[
  {"id": "q1", "objective": "What is the capital of France?", "expected": "Paris"},
  {"id": "q2", "objective": "Who wrote Romeo and Juliet?", "expected": "Shakespeare"}
]

The format supports five task types (QA, Classification, Criteria Evaluation, Instruction Following, and Agentic), each with its own verification logic. See the full spec.

Configuration

Palace stores settings in ~/.config/palace/config.yaml. The quick start covers the essentials.

For CI/Docker, use environment variables instead. Run palace config env to see the mapping.

Included benchmarks

Category Benchmarks
Knowledge & Reasoning SimpleQA, HotpotQA, Humanity's Last Exam, GPQA Diamond, MUSR
Academic & Math MMLU, MMLU-Pro, MATH-500, AIME 2025, BBH, HellaSwag
Multilingual MMMLU (14 languages), MGSM, Belebele (122 languages)
Long Context BABILong (4k-128k), LongBench v2
Multimodal MMMU, MMMU Pro, VLSBench
Instruction Following IFEval
Agentic GAIA, AssistantBench
palace list                    # see all available
palace download MMLU           # download one
palace download --all          # download everything

Running evaluations

palace run MMLU -m gpt-4o
palace run MMLU -m gpt-4o -l 50           # limit to 50 tasks
palace run "GPQA Diamond" -m claude-sonnet
palace run SWE-bench -m o3-mini --agentic # agentic benchmark

Python API (for integration):

from palace import evaluate

evaluate(
    run_name="my-eval",
    url="https://api.openai.com/v1",
    token="sk-...",
    name="gpt-4o",
    tasklist="SimpleQA",
    limit=100,
)

Works with any OpenAI-compatible endpoint: OpenAI, Anthropic, Azure, Mistral, local deployments, etc.

Agentic evaluation

For benchmarks where the model uses tools (GAIA, AssistantBench, etc.), you need Vivarium, a sandboxed Docker runtime for agents.

pip install vivarium-ai

Vivarium starts automatically when you run an agentic tasklist. Requires Docker 24+.

Note: Vivarium is currently pending open-source release (expected within weeks). Contact massimiliano.altieri@ec.europa.eu for early access.

Creating your own benchmarks

palace init my-benchmark      # interactive wizard
palace validate my-benchmark  # check for errors
palace publish my-benchmark   # publish to HuggingFace

See Your First Benchmark for a walkthrough.

Documentation

Full docs at palace.pages.code.europa.eu/palace-eval.

Contributing

Issues and PRs welcome. See CONTRIBUTING.md.

License

Copyright 2025 European Union. Licensed under EUPL-1.2.

Developed by the European Commission Joint Research Centre.

Citation

@software{palace2025,
  title = {PALACE: An Open Benchmark Format for LLM Evaluation},
  author = {Altieri, Massimiliano},
  year = {2025},
  institution = {European Commission Joint Research Centre},
  url = {https://code.europa.eu/palace/palace-eval},
  license = {EUPL-1.2}
}

Contact

Release files for palace-eval 1.0.10

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for palace-eval 1.0.10
File Size Uploaded
palace_eval-1.0.10.tar.gz 136.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for palace-eval 1.0.10
File Interpreter ABI Platform
palace_eval-1.0.10-py3-none-any.whl Python 3 none any Details

Total release size: 325.4 kB

Release files / palace_eval-1.0.10.tar.gz

Download URL palace_eval-1.0.10.tar.gz
Size 136.8 kB
Tags Source
SHA-256 checksum
How to use checksums
510eeed015edf172d1408ac5b5db45e778a73c0a9b15a2b07620fb53d7c9c046
BLAKE2b-256 checksum
How to use checksums
06d3bf4f67cde5cad5c28d6ee27ec8056813a1481164d800e5699f1a8c800203
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release files / palace_eval-1.0.10-py3-none-any.whl

Download URL palace_eval-1.0.10-py3-none-any.whl
Size 188.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3eb85e0f8ec3e013c3dc5fe65bcf557d87b92e1f1e09fbd58d73ae21d3988d06
BLAKE2b-256 checksum
How to use checksums
01df1b7b668c05fa231f87413922c11eb0797e572f2bd6ec0ddb297fa091b5fe
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.13.15

Release history Release notifications | RSS feed

1.1.4

2 release files

1.1.3

2 release files

1.1.2

2 release files

1.1.1

2 release files

1.1.0

2 release files

This release

1.0.10 This release

2 release files

1.0.9

2 release files

1.0.8

2 release files

1.0.7

2 release files

1.0.6

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page