Skip to main content

Interactive-ARC

An interactive benchmark for evaluating LLM abstract reasoning on ARC-AGI tasks.

Instead of producing output grids directly, models construct solutions step-by-step using tool calls. Every action is recorded, producing interpretable reasoning traces that reveal how a model solves a task, not just whether it does.

Features

  • Interactive evaluation: models build solutions incrementally using grid-editing tools
  • Full action traces: every tool call, grid state, and token count is recorded
  • 1,120 public tasks: ships with ARC-AGI-1 (800) and ARC-AGI-2 (320)
  • Multiple providers: supports Anthropic, Amazon Bedrock, and any OpenAI-compatible endpoint (including vLLM)
  • Concurrent execution: evaluates tasks in parallel with configurable concurrency
  • Checkpointing: interrupted runs resume from where they left off

Installation

pip install interactive-arc

Requires Python 3.12+.

Quick Start

With a cloud provider

# Anthropic
export ANTHROPIC_API_KEY=your-key
interactive-arc run --provider anthropic --model claude-sonnet-4-20250514

# Amazon Bedrock (uses default AWS credentials)
interactive-arc run --provider bedrock --model anthropic.claude-sonnet-4-20250514-v1:0

With a local model (vLLM, Ollama, etc.)

interactive-arc run \
    --provider openai \
    --base-url http://localhost:8000/v1 \
    --model Qwen/Qwen3.6-27B \
    --dataset arc-agi-1 \
    --split training \
    --sample 50 --seed 42

Inspect a single task

interactive-arc task --task-id 08ed6ac7 --provider anthropic --model claude-sonnet-4-20250514

CLI Reference

interactive-arc run [OPTIONS]
Option Default Description
--dataset arc-agi-1 Dataset (arc-agi-1 or arc-agi-2)
--split training Split (training or evaluation)
--provider bedrock LLM provider (anthropic, bedrock, openai)
--model Model identifier
--base-url Base URL for OpenAI-compatible endpoints
--renderer text Grid format sent to model (text, json, markdown)
--sample all Number of tasks to sample
--seed Random seed for reproducible sampling
--output Path for summary statistics JSON
--traces ./traces Directory for full trace files
--max-attempts 2 Submission attempts per task (1-10)
--enabled-tools all Comma-separated subset of tools to enable
--grid-feedback both Grid state shown after actions (both, output, none)

Tools

Models interact with the grid through these tools:

Tool Description
set_cell(x, y, color) Set a single cell
set_width(width) Resize grid width
set_height(height) Resize grid height
flood_fill(x, y, color) Fill connected region
copy_input() Copy test input to output grid
copy_region(x, y, w, h) Copy a rectangular region to clipboard
paste_region(x, y) Paste clipboard at position
undo() Undo last operation
reset() Reset grid to initial state
submit(explanation) Submit current grid as answer

Python API

from interactive_arc.environment.loader import TaskLoader
from interactive_arc.environment.tools import ToolExecutor
from interactive_arc.agent.loop import AgentLoop
from interactive_arc.agent.providers.anthropic import AnthropicLLM
from interactive_arc.agent.renderers.text_renderer import TextRenderer

# Load a task
loader = TaskLoader("arc-agi-1", "training")
task = loader.load_task("08ed6ac7")

# Create an agent and solve
llm = AnthropicLLM(model="claude-sonnet-4-20250514")
loop = AgentLoop(task=task, llm=llm, renderer=TextRenderer())
result = loop.run()

print(f"Solved: {result.success}")
print(f"Actions: {result.total_tool_calls}")

Output

Each run produces:

  • Summary JSON: success rate, action efficiency, token usage, cost estimates
  • Trace files: one JSON per task with the full interaction history (every tool call, grid state, LLM response, and timestamps)

Architecture

The codebase follows a three-layer architecture with strict one-directional dependencies:

  1. Environment: grid state machine, tool execution, task loading
  2. Agent: multi-turn LLM interaction loop, provider adapters, grid renderers
  3. Runner: concurrent orchestration, checkpointing, metrics, CLI

Components are swappable via Protocol classes. Adding a new LLM provider or grid renderer requires implementing a single interface with no changes to other layers.

Development

git clone https://github.com/interactive-arc/interactive-arc.git
cd interactive-arc
uv sync --dev
uv run pytest tests/
uv run ruff check src/ tests/

Licence

MIT

Release files for interactive-arc 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for interactive-arc 0.1.2
File Size Uploaded
interactive_arc-0.1.2.tar.gz 2.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for interactive-arc 0.1.2
File Interpreter ABI Platform
interactive_arc-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 2.2 MB

Release files / interactive_arc-0.1.2.tar.gz

Download URL interactive_arc-0.1.2.tar.gz
Size 2.2 MB
Tags Source
SHA-256 checksum
How to use checksums
2fd1c913c8a411bc2e8a2616a17d1310e2b1259d7b93e0ce90d4fdc184abcf08
BLAKE2b-256 checksum
How to use checksums
c573f2b1a105823e59a1b912788d469db09fe749f31aa4ba208465104c4f281a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 18, 2026.

Transparency log

Release files / interactive_arc-0.1.2-py3-none-any.whl

Download URL interactive_arc-0.1.2-py3-none-any.whl
Size 43.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
68c7f8b0d43b69a99d33f286d1fca31ca45bc5afb20321c13d33b2a56cb304ed
BLAKE2b-256 checksum
How to use checksums
cfdddce003aa6fa13b527f7b3fabcc6de862c9347baf109596346f3f5254dbbb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.12

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on May 18, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page