Skip to main content

Add your description here

Project description

Craftax LM

A wrapper around the Craftax agent benchmark, for evaluating digital agents over extremely long time horizons.

Craftax-Classic

LM Algorithm Score (% max) Code
claude-3-7-sonnet-latest (default) ReAct 18.0
claude-3-5-sonnet-20241022 ReAct 17.8
claude-3-5-sonnet-20240620 ReAct 15.7
o3-mini ReAct 12.6
gpt-4o ReAct 7.0
  • Note - this is a limited evaluation where trajectories are terminated after 30 api calls, or roughly 150 in-game steps. 10 trajectories are rolled-out, yielding a log-weighted score as per the Crafter paper. Reproducible code forthcoming.

Usage

First, download the package with pip install craftaxlm. Next, import the agent-computer interface of your choice via

from craftaxlm import CraftaxACI, CraftaxClassicACI

This package is early in development, so for implementation examples, please refer to the baseline ReAct implementation

Leaderboard

In order to make experiments reasonable to run across a range of LMs, currently the leaderboard evaluates agents in the following manner:

  1. Five rollouts are sampled from the agent, with a hard cap of 300 actions per rollout.
  2. The agent is evaluated using a modified version of the original Crafter score -
    sum(ln(1 + P(1_achievement_obtained)) for achievement in achievements) / (sum(ln(2) * len(achievements)))
    
    where P(1_achievement_obtained) is the probability of the achievement being obtained in a single rollout. The key idea is that incremental progress towards difficult achievements ought to weigh more heavily in the score.

Craftax-Full

LM Algorithm Score (% max) Code

Dev Instructions

pyenv virtualenv craftax_env
poetry install

When in doubt

from jax import debug
...
debug.breakpoint()

📚 Citation

To learn more about Craftax, check out the paper website here. To cite the underlying Craftax environment, see:

@inproceedings{matthews2024craftax,
    author={Michael Matthews and Michael Beukman and Benjamin Ellis and Mikayel Samvelyan and Matthew Jackson and Samuel Coward and Jakob Foerster},
    title = {Craftax: A Lightning-Fast Benchmark for Open-Ended Reinforcement Learning},
    booktitle = {International Conference on Machine Learning ({ICML})},
    year = {2024}
}

To cite the Crafter benchmark, see:

@article{hafner2021crafter,
  title={Benchmarking the Spectrum of Agent Capabilities},
  author={Danijar Hafner},
  year={2021},
  journal={arXiv preprint arXiv:2109.06780},
}

Contributing

Setup

uv venv craftaxlm-dev
source craftaxlm-dev/bin/activate
uv sync
uv run ruff format .

Help Wanted

  • General code quality suggestions or improvements. Especially those that improve speed or reduce tokens.
  • PRs to fix issues or add afforances that help your LM agent perform well
  • Leaderboard submissions that demonstrate improved performance using algorithms for learning from data

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

craftaxlm-0.0.34.tar.gz (31.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

craftaxlm-0.0.34-py3-none-any.whl (31.7 kB view details)

Uploaded Python 3

File details

Details for the file craftaxlm-0.0.34.tar.gz.

File metadata

  • Download URL: craftaxlm-0.0.34.tar.gz
  • Upload date:
  • Size: 31.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.11.8

File hashes

Hashes for craftaxlm-0.0.34.tar.gz
Algorithm Hash digest
SHA256 17e0d2f2bd19d1bd6361d9cdf994d2f03a479048be26fea8a431ec6dfee4a02e
MD5 5bad27b62d6ccdc443574d2f5c43d247
BLAKE2b-256 1481194fbc9dc2655af6fccfd08f5789d36505704ea54a78f3b97b30c330a35c

See more details on using hashes here.

File details

Details for the file craftaxlm-0.0.34-py3-none-any.whl.

File metadata

  • Download URL: craftaxlm-0.0.34-py3-none-any.whl
  • Upload date:
  • Size: 31.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.11.8

File hashes

Hashes for craftaxlm-0.0.34-py3-none-any.whl
Algorithm Hash digest
SHA256 f6927b6a5f35178b09f460217c76d726512be29b63359e67adc6d534ac1b2398
MD5 9cbbcae436b886eb6eb6ddfb1cf11a60
BLAKE2b-256 7a72f59b1a81ad7edf7aa57c646e58e8e0f612f77c5693a20ffc644115703b9f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page