9 projects
harbor
A framework for evaluating and optimizing agents and models using sandboxed environments.
harbor-rewardkit
Lightweight grading toolkit for environment-based tasks.
harbor-atif2otel
Convert ATIF agent trajectories to OpenTelemetry protobuf spans with pluggable uploaders.
harbor-langsmith
LangSmith plugin for Harbor jobs.
harbor-hub
Add your description here
terminus-ai
Terminus CLI: An autonomous AI agent for terminal-based task execution
terminal-bench
Terminal-bench is a collection of tasks and evaluation harness for evaluating AI agents' ability to complete complex tasks in terminal environments.
sandboxes
Add your description here
benchmarks
A library for building agentic benchmarks.