Python SDK for evaluating in ScaleWoB: Scalable world-of-bit that revolutionizes the evaluation of Computer-Use Agents.
Project description
ScaleWoB Python SDK
Python SDK for evaluating in ScaleWoB: Scalable world-of-bit that revolutionizes the evaluation of Computer-Use Agents.
🔥 Use this SDK to plug your computer-use agent to our upcoming benchmark!
Installation
pip install scalewob
Quick Start
from scalewob import ScaleWoBAutomation
# Initialize automation for a specific environment
auto = ScaleWoBAutomation(env_id='booking-hotel-simple')
# Start browser and load environment
auto.start()
# Start evaluation mode
auto.start_evaluation()
# Perform actions using coordinates
auto.click(x=300, y=150) # Click at coordinates
auto.type('New York') # Type into focused element
# Finish evaluation and get results
result = auto.finish_evaluation({'destination': 'New York'})
print(result)
# Clean up
auto.close()
Discovering Tasks
Fetch available tasks from the ScaleWoB registry as a flat list:
from scalewob import fetch_tasks
# Get all available tasks
tasks = fetch_tasks()
print(f"Found {len(tasks)} tasks")
# Filter by difficulty
expert_tasks = fetch_tasks(difficulty="Expert")
# Filter by platform and tags
time_selection_tasks = fetch_tasks(
platform="Mobile Interfaces",
tags=["Time Selection"]
)
# Each task includes environment context
for task in tasks[:3]:
print(f"[{task['env_id']}:{task['task_id']}] {task['description']}")
See Task Discovery in the API Reference for more details.
Usage
Context Manager
with ScaleWoBAutomation(env_id='booking-hotel-simple') as auto:
auto.start()
auto.start_evaluation()
auto.click(x=300, y=150)
auto.type('New York')
result = auto.finish_evaluation({'destination': 'New York'})
Configuration
auto = ScaleWoBAutomation(
env_id='booking-hotel-simple',
headless=False, # Run in headless mode
base_url='https://niumascript.com/scalewob-env',
timeout=5000, # Default timeout in milliseconds
screenshot_quality='high', # 'low' (1x) or 'high' (3x) scale on mobile
platform='mobile' # 'mobile' for iPhone emulation, 'desktop' for standard browser
)
API Reference
Initialization
ScaleWoBAutomation(env_id, headless=False, base_url='https://niumascript.com/scalewob-env', timeout=5000, screenshot_quality='high', platform='mobile')
Initialize automation interface for ScaleWoB environments.
Parameters:
env_id(str): Environment ID to launchheadless(bool): Run browser in headless mode (default: False). Uses Chrome browser.base_url(str): Base URL for ScaleWoB environments (default: 'https://niumascript.com/scalewob-env')timeout(int): Default timeout for operations in milliseconds (default: 5000)screenshot_quality(str): Screenshot quality - 'low' for 1x scale, 'high' for 3x scale on mobile (default: 'high')platform(str): Platform type - 'mobile' for iPhone emulation, 'desktop' for standard browser (default: 'mobile')
Note: Currently only Chrome browser is supported. The browser runs with stealth mode options to avoid detection. Mobile mode uses iPhone viewport (390x844) with 3x pixel ratio and touch interactions. Desktop mode uses standard browser window (1280x800) with mouse interactions.
Core Methods
start()
Initialize Chrome browser and navigate to the environment page. Must be called before any other automation methods. Waits for DOM to be fully loaded before returning.
start_evaluation()
Start evaluation mode. Ensures the environment is fully initialized and clears the trajectory for a fresh evaluation. The environment loads ready to interact without requiring UI button clicks.
finish_evaluation(task_id=0, params=None)
Finish evaluation and get results.
Parameters:
task_id(int, optional): Task index within the environment (default: 0). Used to identify which task in the environment's tasks array is being evaluated.params(dict, optional): Evaluation parameters (environment-specific)
Returns: Evaluation result dictionary
Interaction Methods
click(x, y)
Click at coordinates (x, y).
Parameters:
x(int): Horizontal coordinatey(int): Vertical coordinate
type(text, append=False)
Type text into the currently focused element. An element must be focused first (e.g., via click).
Parameters:
text(str): Text to typeappend(bool): If True, append to existing text; if False, clear field first (default: False)
scroll(x, y, direction='down', distance=100)
Scroll in direction from coordinates (x, y).
Parameters:
x(int): Horizontal coordinatey(int): Vertical coordinatedirection(str): Scroll direction ('up', 'down', 'left', 'right')distance(int): Distance to scroll in pixels
long_press(x, y, duration=1000)
Long press at coordinates (x, y).
Note: This is a mobile-specific gesture and will raise CommandError on desktop platform.
Parameters:
x(int): Horizontal coordinatey(int): Vertical coordinateduration(int): Duration of press in milliseconds
drag(x, y, end_x, end_y)
Drag from start coordinates to end coordinates.
Parameters:
x(int): Starting horizontal coordinatey(int): Starting vertical coordinateend_x(int): Ending horizontal coordinateend_y(int): Ending vertical coordinate
back()
Go back in navigation history.
State and Information Methods
take_screenshot(format='base64')
Capture screenshot of the environment.
Parameters:
format(str): Return format - "base64" for raw base64 string, "pil" for PIL Image object
Returns: Base64 string or PIL Image object
get_evaluation_result()
Get the last evaluation result.
Returns: Last evaluation result or None
get_trajectory()
Get current action trajectory.
Returns a copy of the trajectory history containing all actions performed since start_evaluation() was called.
Returns: List of trajectory entries with timestamp, type, and data
Example:
trajectory = auto.get_trajectory()
print(f"Collected {len(trajectory)} actions")
for action in trajectory:
print(f"{action['type']} at {action['timestamp']}")
clear_trajectory()
Clear the current trajectory history.
This is useful if you want to reset the trajectory without restarting the evaluation. Note that start_evaluation() automatically clears the trajectory.
Example:
auto.clear_trajectory()
print(len(auto.get_trajectory())) # 0
close()
Close browser and cleanup resources.
Task Discovery
fetch_tasks(difficulty=None, platform=None, tags=None, force_refresh=False)
Fetch all tasks from ScaleWoB registry as a flat list with optional filtering.
Each task includes its environment context, making it easy to iterate through all available tasks without nested loops.
Parameters:
difficulty(str, optional): Filter by difficulty level (e.g., "Basic", "Advanced", "Expert")platform(str, optional): Filter by platform (e.g., "Mobile Interfaces")tags(list, optional): Filter by tags (returns tasks from environments matching any tag)force_refresh(bool): Bypass cache and fetch fresh data (default: False)
Returns: List of task dictionaries, each containing:
env_id: Environment IDenv_name: Environment display nametask_id: Task index within the environment (for use withfinish_evaluation())task_name: Task name (if available)description: Task description/instructiondifficulty: Environment difficulty levelplatform: Environment platformtags: Environment tagsparams: Task parameters (if any)
Raises: NetworkError if fetching or parsing fails
Example:
from scalewob import fetch_tasks, ScaleWoBAutomation
# Get all tasks
all_tasks = fetch_tasks()
# Filter by multiple criteria
filtered = fetch_tasks(
difficulty="Expert",
platform="Mobile Interfaces"
)
# Force refresh cache
fresh = fetch_tasks(force_refresh=True)
# Iterate through tasks and run evaluations
for task in fetch_tasks(difficulty="Basic"):
auto = ScaleWoBAutomation(task['env_id'])
auto.start()
auto.start_evaluation()
# ... perform actions based on task['description'] ...
result = auto.finish_evaluation(task_id=task['task_id'])
auto.close()
Exception Handling
from scalewob import (
ScaleWoBError, # Base exception
TimeoutError, # Operation timeout
CommandError, # Command execution failure
EvaluationError, # Evaluation failure
BrowserError, # Browser automation failure
NetworkError # Network operation failure
)
try:
auto = ScaleWoBAutomation(env_id='booking-hotel-simple')
auto.start()
auto.start_evaluation()
result = auto.finish_evaluation()
except TimeoutError as e:
print(f"Operation timed out: {e}")
except EvaluationError as e:
print(f"Evaluation failed: {e}")
except ScaleWoBError as e:
print(f"ScaleWoB error: {e}")
finally:
auto.close()
Development
Setup
# Clone the repo and enter the directory first
uv sync
# Install pre-commit hooks
uv pre-commit install
Code Quality
# Format code
uv run poe format
# Run checks (format, lint, type checking)
uv run poe check
# Fix linting issues
uv run poe fix
License
MIT License - see LICENSE file for details.
Links
- Homepage: https://github.com/ScaleWoB/ScaleWoB.github.io
- Documentation: https://github.com/ScaleWoB/ScaleWoB#readme
- Bug Tracker: https://github.com/ScaleWoB/ScaleWoB/issues
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file scalewob-0.7.1.tar.gz.
File metadata
- Download URL: scalewob-0.7.1.tar.gz
- Upload date:
- Size: 12.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
fe47e76c9088ec00705d70e56e2d22aa614fbf8becfb1289dd028b516b7c649d
|
|
| MD5 |
404fb445ae2fd9d133caa09f89a82c0e
|
|
| BLAKE2b-256 |
224cb363477405548dbf5899e041c56d1e94934e3fe0033d6f07dd36bf23c21d
|
File details
Details for the file scalewob-0.7.1-py3-none-any.whl.
File metadata
- Download URL: scalewob-0.7.1-py3-none-any.whl
- Upload date:
- Size: 13.9 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: uv/0.9.18 {"installer":{"name":"uv","version":"0.9.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":true}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7dd0d848f97034b09ebd61a7264c701e88b8f4b86f0667d98ab35738452a07de
|
|
| MD5 |
5c4521cb18329c2fc92f33606261bbdc
|
|
| BLAKE2b-256 |
9369a383ca8e70a50194656554ddbda2887b7b3365a20cc974e2ef0b574d92f4
|