WebAI: The AI-Ready Internet Standard & Token Benchmarking Suite
WebAI is an open architectural standard, benchmarking suite, and agent identity framework designed to establish an AI-Agent-Ready Internet operating alongside the modern human-first web.
🌍 The WebAI Manifesto: Why Build This Now?
- The 57% Reality: Major internet infrastructure providers like Cloudflare have documented that over 57% of all web traffic is now generated by automated systems and AI agents.
- The Structural Flaw: Despite AI being the majority of internet consumers, the web remains engineered exclusively for human visual perception (deeply nested HTML DOM trees, CSS stylesheets, layout wrappers, and tracking scripts).
- The Downstream Impact: AI agents waste up to 95%+ of their I/O token budgets simply parsing presentation boilerplate. This inflates task completion times, drives up inference costs, and needlessly ties up global GPU clusters that could be solving valuable societal tasks.
- The Dual-Web Solution: Humans keep their rich visual web interfaces, while websites provide a parallel Agent-Ready Web with zero presentation bloat.
The Two Pillars of WebAI
- Pillar 1: Content & Representation Standards
- Standardized
llms.txtand semantic JSON contracts. - Storage4gaming.com: Complete rewrite into the reference AI-agent-ready site.
- Empirical Benchmarks: Quantifying token and latency savings across open-weight models on Storage4gaming and the Steam Storefront.
- Standardized
- Pillar 2: AI Agent Identity & Authentication (AIAID)
- Cryptographic machine identities for AI agents (analogous to MAC addresses / UUIDs).
- A public depository and challenge-response handshake that unlocks high-speed semantic endpoints for verified agents while filtering malicious scrapers.
🌟 Key Highlights
- ~76% - 95% Token Reduction: Cuts prompt tokens from thousands of tokens down to ~100 tokens.
- 1.5x - 2.5x Faster Latency: Accelerates task completion by stripping away visual rendering layers.
- Embedded SQLite Telemetry: All benchmark runs, tokenizer token counts, throughput (tok/s), latency, and peak RAM are persisted to
webai_benchmarks.dband synchronized withbenchmark_results.csv. - Apple Silicon Optimized: 100% offline inference on Apple Silicon using
mlx-lmwith sequential execution and automatic memory cache clearing to protect 16 GB Unified Memory. - Auto-Updating Documentation: Live benchmark tables and model cache statuses in this README and
ARCHITECTURE.mdstay automatically synchronized with your local benchmark runs.
📊 Live Benchmark Performance Summary
The table below is automatically synchronized with
webai_benchmarks.dbbydoc_updater.py:
| Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
|---|---|---|---|---|---|---|---|
Llama-3.2-1B-Instruct-4bit |
Apex Legends | 565 tok | 132 tok | 76.64% | 0.56s | 0.42s | 1.3x faster |
Llama-3.2-1B-Instruct-4bit |
Counter-Strike 2 | 567 tok | 136 tok | 76.01% | 0.99s | 0.45s | 2.2x faster |
Llama-3.2-1B-Instruct-4bit |
Halo Infinite | 565 tok | 132 tok | 76.64% | 0.50s | 0.42s | 1.2x faster |
gemma-4-e4b-it-OptiQ-4bit |
Apex Legends | 635 tok | 125 tok | 80.31% | 4.69s | 3.45s | 1.4x faster |
gemma-4-e4b-it-OptiQ-4bit |
Counter-Strike 2 | 637 tok | 132 tok | 79.28% | 6.96s | 4.77s | 1.5x faster |
gemma-4-e4b-it-OptiQ-4bit |
Halo Infinite | 636 tok | 126 tok | 80.19% | 3.61s | 3.31s | 1.1x faster |
🤖 Supported Model Matrix & Local Cache Status
The suite evaluates models across two buckets, all optimized in 4-bit quantization for 16 GB Unified Memory:
- Bucket 1 (< 4B): Ultra-lightweight models for fast edge inference.
- Bucket 2 (4B – 12B): Desktop-class models including Google's Gemma 4 12B and Gemma 4 E4B (8B).
Check your local cache anytime with
python download_models.py --list:
| Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
|---|---|---|---|
| [CACHED] | <4B | ~1.0 GB | mlx-community/Llama-3.2-1B-Instruct-4bit |
| [NOT CACHED] | <4B | ~2.2 GB | mlx-community/Llama-3.2-3B-Instruct-4bit |
| [NOT CACHED] | <4B | ~2.1 GB | mlx-community/Qwen2.5-3B-Instruct-4bit |
| [NOT CACHED] | <4B | ~2.5 GB | mlx-community/Phi-4-mini-instruct-4bit |
| [CACHED] | 4B-12B | ~7.5 GB | mlx-community/gemma-4-12B-it-qat-4bit |
| [CACHED] | 4B-12B | ~3.5 GB | mlx-community/gemma-4-e4b-it-OptiQ-4bit |
| [NOT CACHED] | 4B-12B | ~4.8 GB | mlx-community/Qwen2.5-7B-Instruct-4bit |
| [NOT CACHED] | 4B-12B | ~5.2 GB | mlx-community/Llama-3.1-8B-Instruct-4bit |
| [NOT CACHED] | 4B-12B | ~5.2 GB | mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit |
🚀 Quickstart
Prerequisites
- macOS on Apple Silicon (M-series).
- Homebrew installed (
/opt/homebrew/bin/brew). - Python 3.11 (
brew install python@3.11).
One-Command Setup & Benchmark
Execute the master runner to set up the environment, run unit tests, and launch a benchmark:
chmod +x run.sh
./run.sh
🛠️ CLI Usage Guide
1. Download / Install Models Without Running Inference
Download model weights directly to your local Hugging Face cache (~/.cache/huggingface/hub/) without loading them into RAM:
# Check cache status of all models:
.venv/bin/python3 download_models.py --list
# Download a specific model (e.g. Gemma 4 12B):
.venv/bin/python3 download_models.py --models mlx-community/gemma-4-12B-it-qat-4bit
# Download all models in a bucket:
.venv/bin/python3 download_models.py --bucket "<4B"
2. Run Comparative Token Benchmarks
Run side-by-side comparative evaluation between human-first HTML and agent-first JSON:
# Benchmark specific models:
.venv/bin/python3 benchmark_runner.py --models mlx-community/gemma-4-e4b-it-OptiQ-4bit mlx-community/Llama-3.2-1B-Instruct-4bit
# Benchmark a specific game title:
.venv/bin/python3 benchmark_runner.py --game "Cyberpunk 2077"
# Enable live web scraping from https://www.storage4gaming.com:
.venv/bin/python3 benchmark_runner.py --live
3. Scan Local Steam Game Catalog
Discover locally installed games and owned account library titles from Steam (Windows & macOS):
# Standard discovery:
.venv/bin/python3 catalog_scanner.py
# Test with synthetic sample catalog:
.venv/bin/python3 catalog_scanner.py --sample
4. Steam Storefront Specs & Hardware Capacity Calculator
Query the Steam Store API vs. storefront HTML, and calculate total library storage and RAM:
# Benchmark single Steam game (e.g. Apex Legends):
.venv/bin/python3 steam_client.py --appid 1172470
# Calculate aggregate storage & RAM for your entire scanned Steam library:
.venv/bin/python3 steam_client.py --scan
5. Update Documentation Automatically
Keep README.md and ARCHITECTURE.md updated with the latest SQLite benchmark records and model cache status:
.venv/bin/python3 doc_updater.py
🗄️ Database & Telemetry Inspection
All metrics are stored in SQLite (webai_benchmarks.db) and CSV (benchmark_results.csv).
Query Recent Runs
sqlite3 webai_benchmarks.db "SELECT model_id, site_format, prompt_tokens, token_savings_pct, total_time_sec, peak_memory_mb FROM benchmarks ORDER BY id DESC LIMIT 6;"
View CSV Log
cat benchmark_results.csv
📚 Central Documentation (docs/)
All architectural designs, research, and design decision logs are organized in the docs/ folder:
- docs/decisions.md: Design log covering Skills vs. Scripts,
llms.txt, local filesystem discovery, PyPI packaging, readability scoring metrics, and the AIAID protocol. - docs/research.md: Research on publishing platforms (GitHub, Hugging Face Datasets & Spaces, Model Context Protocol / MCP,
llms.txtdirectories, and PyPI). - docs/ARCHITECTURE.md: Full system design, Mermaid data flow, AIAID authentication handshake, and Storage4gaming rewrite blueprint.
- docs/WebAI_specification.md: Master project specification and criteria.
- docs/starter.md: Task requirements and model testing matrix.
📂 Repository Layout
├── docs/ # Central documentation folder
│ ├── decisions.md # Pending design decisions and architectural trade-offs
│ ├── research.md # Research on publishing platforms (GitHub, Hugging Face, MCP)
│ ├── ARCHITECTURE.md # Technical design, data flow, and database schema
│ ├── WebAI_specification.md # Core project specification
│ └── starter.md # Task requirements and model evaluation matrix
├── llms.txt # Standard machine-readable AI agent spec index
├── llms-full.txt # Comprehensive agent knowledge base and sizing formulas
├── README.md # Project overview, manifesto, quickstart, and live scoreboard
├── ARCHITECTURE.md # Root link to technical design
├── telemetry_db.py # SQLite telemetry database engine and CSV exporter
├── storage4gaming_client.py # Storage4gaming comparative payload generator (HTML vs JSON vs llms.txt)
├── steam_client.py # Steam Storefront API client & aggregate hardware calculator
├── benchmark_runner.py # MLX sequential inference and benchmarking harness
├── download_models.py # Standalone model downloader/installer
├── catalog_scanner.py # Cross-platform Steam game library detector (Windows & macOS)
├── doc_updater.py # Documentation auto-synchronizer
├── test_telemetry.py # Automated test suite for database and telemetry
├── run.sh # Master setup and execution script
├── webai_benchmarks.db # SQLite database storing benchmark records
└── benchmark_results.csv # Exported benchmark CSV dataset
Metadata
Release files for webai-scanner 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| webai_scanner-0.1.0.tar.gz | 65.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| webai_scanner-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 80.9 kB
Release files / webai_scanner-0.1.0.tar.gz
| Download URL | webai_scanner-0.1.0.tar.gz |
|---|---|
| Size | 65.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e8c48c4cd543b0c469d0ab85a9b696c473bd46100dc75bccafc65a2c2025dc6e
|
|
BLAKE2b-256 checksum How to use checksums |
c389653782771ee14ca441807f8c740ce675210ab82c129db093cfe69848b98f
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / webai_scanner-0.1.0-py3-none-any.whl
| Download URL | webai_scanner-0.1.0-py3-none-any.whl |
|---|---|
| Size | 15.3 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ada33df5cd4136b99f5e4a290a3e545995130656cabc30b98892e10306d9fffd
|
|
BLAKE2b-256 checksum How to use checksums |
99ba1198fd1d40cfd97b37889aa70554a1902ea0bae24c53288d0ee875f0ddc1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|