Skip to main content

Python SDK for evaluating in ScaleWoB: Scalable world-of-bit that revolutionizes the evaluation of Computer-Use Agents.

Project description

ScaleWoB Python SDK

PyPI version Python 3.12+ License: MIT REUSE status

Python SDK for evaluating in ScaleWoB: Scalable world-of-bit that revolutionizes the evaluation of Computer-Use Agents.

🔥 Use this SDK to plug your computer-use agent to our upcoming benchmark!

Installation

pip install scalewob

Quick Start

from scalewob import ScaleWoBAutomation

# Initialize automation for a specific environment
auto = ScaleWoBAutomation(env_id='booking-hotel-simple')

# Start browser and load environment
auto.start()

# Start evaluation mode
auto.start_evaluation()

# Perform actions using coordinates
auto.click(x=300, y=150)  # Click at coordinates
auto.type('New York')      # Type into focused element

# Finish evaluation and get results
result = auto.finish_evaluation({'destination': 'New York'})
print(result)

# Clean up
auto.close()

Discovering Tasks

Fetch available tasks from the ScaleWoB registry as a flat list:

from scalewob import fetch_tasks

# Get all available tasks
tasks = fetch_tasks()
print(f"Found {len(tasks)} tasks")

# Filter by difficulty
expert_tasks = fetch_tasks(difficulty="Expert")

# Filter by platform and tags
time_selection_tasks = fetch_tasks(
    platform="Mobile Interfaces",
    tags=["Time Selection"]
)

# Each task includes environment context
for task in tasks[:3]:
    print(f"[{task['env_id']}:{task['task_id']}] {task['description']}")

See Task Discovery in the API Reference for more details.

Usage

Context Manager

with ScaleWoBAutomation(env_id='booking-hotel-simple') as auto:
    auto.start()
    auto.start_evaluation()
    auto.click(x=300, y=150)
    auto.type('New York')
    result = auto.finish_evaluation({'destination': 'New York'})

Configuration

auto = ScaleWoBAutomation(
    env_id='booking-hotel-simple',
    headless=False,             # Run in headless mode
    base_url='https://niumascript.com/scalewob-env',
    timeout=5000,               # Default timeout in milliseconds
    screenshot_quality='high',  # 'low' (1x) or 'high' (3x) scale on mobile
    platform='mobile'           # 'mobile' for iPhone emulation, 'desktop' for standard browser
)

API Reference

Initialization

ScaleWoBAutomation(env_id, headless=False, base_url='https://niumascript.com/scalewob-env', timeout=5000, screenshot_quality='high', platform='mobile')

Initialize automation interface for ScaleWoB environments.

Parameters:

  • env_id (str): Environment ID to launch
  • headless (bool): Run browser in headless mode (default: False). Uses Chrome browser.
  • base_url (str): Base URL for ScaleWoB environments (default: 'https://niumascript.com/scalewob-env')
  • timeout (int): Default timeout for operations in milliseconds (default: 5000)
  • screenshot_quality (str): Screenshot quality - 'low' for 1x scale, 'high' for 3x scale on mobile (default: 'high')
  • platform (str): Platform type - 'mobile' for iPhone emulation, 'desktop' for standard browser (default: 'mobile')

Note: Currently only Chrome browser is supported. The browser runs with stealth mode options to avoid detection. Mobile mode uses iPhone viewport (390x844) with 3x pixel ratio and touch interactions. Desktop mode uses standard browser window (1280x800) with mouse interactions.

Core Methods

start()

Initialize Chrome browser and navigate to the environment page. Must be called before any other automation methods. Waits for DOM to be fully loaded before returning.

start_evaluation()

Start evaluation mode. Ensures the environment is fully initialized and clears the trajectory for a fresh evaluation. The environment loads ready to interact without requiring UI button clicks.

finish_evaluation(task_id=0, params=None)

Finish evaluation and get results.

Parameters:

  • task_id (int, optional): Task index within the environment (default: 0). Used to identify which task in the environment's tasks array is being evaluated.
  • params (dict, optional): Evaluation parameters (environment-specific)

Returns: Evaluation result dictionary

Interaction Methods

click(x, y)

Click at coordinates (x, y).

Parameters:

  • x (int): Horizontal coordinate
  • y (int): Vertical coordinate

type(text, append=False)

Type text into the currently focused element. An element must be focused first (e.g., via click).

Parameters:

  • text (str): Text to type
  • append (bool): If True, append to existing text; if False, clear field first (default: False)

scroll(x, y, direction='down', distance=100)

Scroll in direction from coordinates (x, y).

Parameters:

  • x (int): Horizontal coordinate
  • y (int): Vertical coordinate
  • direction (str): Scroll direction ('up', 'down', 'left', 'right')
  • distance (int): Distance to scroll in pixels

long_press(x, y, duration=1000)

Long press at coordinates (x, y).

Note: This is a mobile-specific gesture and will raise CommandError on desktop platform.

Parameters:

  • x (int): Horizontal coordinate
  • y (int): Vertical coordinate
  • duration (int): Duration of press in milliseconds

drag(x, y, end_x, end_y)

Drag from start coordinates to end coordinates.

Parameters:

  • x (int): Starting horizontal coordinate
  • y (int): Starting vertical coordinate
  • end_x (int): Ending horizontal coordinate
  • end_y (int): Ending vertical coordinate

back()

Go back in navigation history.

State and Information Methods

take_screenshot(format='base64')

Capture screenshot of the environment.

Parameters:

  • format (str): Return format - "base64" for raw base64 string, "pil" for PIL Image object

Returns: Base64 string or PIL Image object

get_evaluation_result()

Get the last evaluation result.

Returns: Last evaluation result or None

get_trajectory()

Get current action trajectory.

Returns a copy of the trajectory history containing all actions performed since start_evaluation() was called.

Returns: List of trajectory entries with timestamp, type, and data

Example:

trajectory = auto.get_trajectory()
print(f"Collected {len(trajectory)} actions")
for action in trajectory:
    print(f"{action['type']} at {action['timestamp']}")

clear_trajectory()

Clear the current trajectory history.

This is useful if you want to reset the trajectory without restarting the evaluation. Note that start_evaluation() automatically clears the trajectory.

Example:

auto.clear_trajectory()
print(len(auto.get_trajectory()))  # 0

close()

Close browser and cleanup resources.

Task Discovery

fetch_tasks(difficulty=None, platform=None, tags=None, force_refresh=False)

Fetch all tasks from ScaleWoB registry as a flat list with optional filtering.

Each task includes its environment context, making it easy to iterate through all available tasks without nested loops.

Parameters:

  • difficulty (str, optional): Filter by difficulty level (e.g., "Basic", "Advanced", "Expert")
  • platform (str, optional): Filter by platform (e.g., "Mobile Interfaces")
  • tags (list, optional): Filter by tags (returns tasks from environments matching any tag)
  • force_refresh (bool): Bypass cache and fetch fresh data (default: False)

Returns: List of task dictionaries, each containing:

  • env_id: Environment ID
  • env_name: Environment display name
  • task_id: Task index within the environment (for use with finish_evaluation())
  • task_name: Task name (if available)
  • description: Task description/instruction
  • difficulty: Environment difficulty level
  • platform: Environment platform
  • tags: Environment tags
  • params: Task parameters (if any)

Raises: NetworkError if fetching or parsing fails

Example:

from scalewob import fetch_tasks, ScaleWoBAutomation

# Get all tasks
all_tasks = fetch_tasks()

# Filter by multiple criteria
filtered = fetch_tasks(
    difficulty="Expert",
    platform="Mobile Interfaces"
)

# Force refresh cache
fresh = fetch_tasks(force_refresh=True)

# Iterate through tasks and run evaluations
for task in fetch_tasks(difficulty="Basic"):
    auto = ScaleWoBAutomation(task['env_id'])
    auto.start()
    auto.start_evaluation()
    # ... perform actions based on task['description'] ...
    result = auto.finish_evaluation(task_id=task['task_id'])
    auto.close()

Exception Handling

from scalewob import (
    ScaleWoBError,      # Base exception
    TimeoutError,       # Operation timeout
    CommandError,       # Command execution failure
    EvaluationError,    # Evaluation failure
    BrowserError,       # Browser automation failure
    NetworkError        # Network operation failure
)

try:
    auto = ScaleWoBAutomation(env_id='booking-hotel-simple')
    auto.start()
    auto.start_evaluation()
    result = auto.finish_evaluation()
except TimeoutError as e:
    print(f"Operation timed out: {e}")
except EvaluationError as e:
    print(f"Evaluation failed: {e}")
except ScaleWoBError as e:
    print(f"ScaleWoB error: {e}")
finally:
    auto.close()

Development

Setup

# Clone the repo and enter the directory first
uv sync

# Install pre-commit hooks
uv pre-commit install

Code Quality

# Format code
uv run poe format

# Run checks (format, lint, type checking)
uv run poe check

# Fix linting issues
uv run poe fix

License

MIT License - see LICENSE file for details.

Links

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scalewob-0.7.1.tar.gz (12.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scalewob-0.7.1-py3-none-any.whl (13.9 kB view details)

Uploaded Python 3

File details

Details for the file scalewob-0.7.1.tar.gz.

File metadata

  • Download URL: scalewob-0.7.1.tar.gz
  • Upload date:
  • Size: 12.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for scalewob-0.7.1.tar.gz
Algorithm Hash digest
SHA256 fe47e76c9088ec00705d70e56e2d22aa614fbf8becfb1289dd028b516b7c649d
MD5 404fb445ae2fd9d133caa09f89a82c0e
BLAKE2b-256 224cb363477405548dbf5899e041c56d1e94934e3fe0033d6f07dd36bf23c21d

See more details on using hashes here.

File details

Details for the file scalewob-0.7.1-py3-none-any.whl.

File metadata

  • Download URL: scalewob-0.7.1-py3-none-any.whl
  • Upload date:
  • Size: 13.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}

File hashes

Hashes for scalewob-0.7.1-py3-none-any.whl
Algorithm Hash digest
SHA256 7dd0d848f97034b09ebd61a7264c701e88b8f4b86f0667d98ab35738452a07de
MD5 5c4521cb18329c2fc92f33606261bbdc
BLAKE2b-256 9369a383ca8e70a50194656554ddbda2887b7b3365a20cc974e2ef0b574d92f4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page