Skip to main content

WebAI: The AI-Ready Internet Standard & Token Benchmarking Suite

WebAI is an open architectural standard, benchmarking suite, and agent identity framework designed to establish an AI-Agent-Ready Internet operating alongside the modern human-first web.


🌍 The WebAI Manifesto: Why Build This Now?

  • The 57% Reality: Major internet infrastructure providers like Cloudflare have documented that over 57% of all web traffic is now generated by automated systems and AI agents.
  • The Structural Flaw: Despite AI being the majority of internet consumers, the web remains engineered exclusively for human visual perception (deeply nested HTML DOM trees, CSS stylesheets, layout wrappers, and tracking scripts).
  • The Downstream Impact: AI agents waste up to 95%+ of their I/O token budgets simply parsing presentation boilerplate. This inflates task completion times, drives up inference costs, and needlessly ties up global GPU clusters that could be solving valuable societal tasks.
  • The Dual-Web Solution: Humans keep their rich visual web interfaces, while websites provide a parallel Agent-Ready Web with zero presentation bloat.

The Two Pillars of WebAI

  1. Pillar 1: Content & Representation Standards
    • Standardized llms.txt and semantic JSON contracts.
    • Storage4gaming.com: Complete rewrite into the reference AI-agent-ready site.
    • Empirical Benchmarks: Quantifying token and latency savings across open-weight models on Storage4gaming and the Steam Storefront.
  2. Pillar 2: AI Agent Identity & Authentication (AIAID)
    • Cryptographic machine identities for AI agents (analogous to MAC addresses / UUIDs).
    • A public depository and challenge-response handshake that unlocks high-speed semantic endpoints for verified agents while filtering malicious scrapers.

🌟 Key Highlights

  • ~76% - 95% Token Reduction: Cuts prompt tokens from thousands of tokens down to ~100 tokens.
  • 1.5x - 2.5x Faster Latency: Accelerates task completion by stripping away visual rendering layers.
  • Embedded SQLite Telemetry: All benchmark runs, tokenizer token counts, throughput (tok/s), latency, and peak RAM are persisted to webai_benchmarks.db and synchronized with benchmark_results.csv.
  • Apple Silicon Optimized: 100% offline inference on Apple Silicon using mlx-lm with sequential execution and automatic memory cache clearing to protect 16 GB Unified Memory.
  • Auto-Updating Documentation: Live benchmark tables and model cache statuses in this README and ARCHITECTURE.md stay automatically synchronized with your local benchmark runs.

📊 Live Benchmark Performance Summary

The table below is automatically synchronized with webai_benchmarks.db by doc_updater.py:

Model Target Game Human Tokens Agent Tokens Token Savings Human Latency Agent Latency Speedup
Llama-3.2-1B-Instruct-4bit Apex Legends 565 tok 132 tok 76.64% 0.56s 0.42s 1.3x faster
Llama-3.2-1B-Instruct-4bit Counter-Strike 2 567 tok 136 tok 76.01% 0.99s 0.45s 2.2x faster
Llama-3.2-1B-Instruct-4bit Halo Infinite 565 tok 132 tok 76.64% 0.50s 0.42s 1.2x faster
gemma-4-e4b-it-OptiQ-4bit Apex Legends 635 tok 125 tok 80.31% 4.69s 3.45s 1.4x faster
gemma-4-e4b-it-OptiQ-4bit Counter-Strike 2 637 tok 132 tok 79.28% 6.96s 4.77s 1.5x faster
gemma-4-e4b-it-OptiQ-4bit Halo Infinite 636 tok 126 tok 80.19% 3.61s 3.31s 1.1x faster

🤖 Supported Model Matrix & Local Cache Status

The suite evaluates models across two buckets, all optimized in 4-bit quantization for 16 GB Unified Memory:

  • Bucket 1 (< 4B): Ultra-lightweight models for fast edge inference.
  • Bucket 2 (4B – 12B): Desktop-class models including Google's Gemma 4 12B and Gemma 4 E4B (8B).

Check your local cache anytime with python download_models.py --list:

Status Bucket Est. RAM (4-bit) Hugging Face Model ID
[CACHED] <4B ~1.0 GB mlx-community/Llama-3.2-1B-Instruct-4bit
[NOT CACHED] <4B ~2.2 GB mlx-community/Llama-3.2-3B-Instruct-4bit
[NOT CACHED] <4B ~2.1 GB mlx-community/Qwen2.5-3B-Instruct-4bit
[NOT CACHED] <4B ~2.5 GB mlx-community/Phi-4-mini-instruct-4bit
[CACHED] 4B-12B ~7.5 GB mlx-community/gemma-4-12B-it-qat-4bit
[CACHED] 4B-12B ~3.5 GB mlx-community/gemma-4-e4b-it-OptiQ-4bit
[NOT CACHED] 4B-12B ~4.8 GB mlx-community/Qwen2.5-7B-Instruct-4bit
[NOT CACHED] 4B-12B ~5.2 GB mlx-community/Llama-3.1-8B-Instruct-4bit
[NOT CACHED] 4B-12B ~5.2 GB mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit

🚀 Quickstart

Prerequisites

  • macOS on Apple Silicon (M-series).
  • Homebrew installed (/opt/homebrew/bin/brew).
  • Python 3.11 (brew install python@3.11).

One-Command Setup & Benchmark

Execute the master runner to set up the environment, run unit tests, and launch a benchmark:

chmod +x run.sh
./run.sh

🛠️ CLI Usage Guide

1. Download / Install Models Without Running Inference

Download model weights directly to your local Hugging Face cache (~/.cache/huggingface/hub/) without loading them into RAM:

# Check cache status of all models:
.venv/bin/python3 download_models.py --list

# Download a specific model (e.g. Gemma 4 12B):
.venv/bin/python3 download_models.py --models mlx-community/gemma-4-12B-it-qat-4bit

# Download all models in a bucket:
.venv/bin/python3 download_models.py --bucket "<4B"

2. Run Comparative Token Benchmarks

Run side-by-side comparative evaluation between human-first HTML and agent-first JSON:

# Benchmark specific models:
.venv/bin/python3 benchmark_runner.py --models mlx-community/gemma-4-e4b-it-OptiQ-4bit mlx-community/Llama-3.2-1B-Instruct-4bit

# Benchmark a specific game title:
.venv/bin/python3 benchmark_runner.py --game "Cyberpunk 2077"

# Enable live web scraping from https://www.storage4gaming.com:
.venv/bin/python3 benchmark_runner.py --live

3. Scan Local Steam Game Catalog

Discover locally installed games and owned account library titles from Steam (Windows & macOS):

# Standard discovery:
.venv/bin/python3 catalog_scanner.py

# Test with synthetic sample catalog:
.venv/bin/python3 catalog_scanner.py --sample

4. Steam Storefront Specs & Hardware Capacity Calculator

Query the Steam Store API vs. storefront HTML, and calculate total library storage and RAM:

# Benchmark single Steam game (e.g. Apex Legends):
.venv/bin/python3 steam_client.py --appid 1172470

# Calculate aggregate storage & RAM for your entire scanned Steam library:
.venv/bin/python3 steam_client.py --scan

5. Update Documentation Automatically

Keep README.md and ARCHITECTURE.md updated with the latest SQLite benchmark records and model cache status:

.venv/bin/python3 doc_updater.py

🗄️ Database & Telemetry Inspection

All metrics are stored in SQLite (webai_benchmarks.db) and CSV (benchmark_results.csv).

Query Recent Runs

sqlite3 webai_benchmarks.db "SELECT model_id, site_format, prompt_tokens, token_savings_pct, total_time_sec, peak_memory_mb FROM benchmarks ORDER BY id DESC LIMIT 6;"

View CSV Log

cat benchmark_results.csv

📚 Central Documentation (docs/)

All architectural designs, research, and design decision logs are organized in the docs/ folder:

  • docs/decisions.md: Design log covering Skills vs. Scripts, llms.txt, local filesystem discovery, PyPI packaging, readability scoring metrics, and the AIAID protocol.
  • docs/research.md: Research on publishing platforms (GitHub, Hugging Face Datasets & Spaces, Model Context Protocol / MCP, llms.txt directories, and PyPI).
  • docs/ARCHITECTURE.md: Full system design, Mermaid data flow, AIAID authentication handshake, and Storage4gaming rewrite blueprint.
  • docs/WebAI_specification.md: Master project specification and criteria.
  • docs/starter.md: Task requirements and model testing matrix.

📂 Repository Layout

├── docs/                        # Central documentation folder
│   ├── decisions.md             # Pending design decisions and architectural trade-offs
│   ├── research.md              # Research on publishing platforms (GitHub, Hugging Face, MCP)
│   ├── ARCHITECTURE.md          # Technical design, data flow, and database schema
│   ├── WebAI_specification.md   # Core project specification
│   └── starter.md               # Task requirements and model evaluation matrix
├── llms.txt                     # Standard machine-readable AI agent spec index
├── llms-full.txt                # Comprehensive agent knowledge base and sizing formulas
├── README.md                    # Project overview, manifesto, quickstart, and live scoreboard
├── ARCHITECTURE.md              # Root link to technical design
├── telemetry_db.py              # SQLite telemetry database engine and CSV exporter
├── storage4gaming_client.py     # Storage4gaming comparative payload generator (HTML vs JSON vs llms.txt)
├── steam_client.py              # Steam Storefront API client & aggregate hardware calculator
├── benchmark_runner.py          # MLX sequential inference and benchmarking harness
├── download_models.py           # Standalone model downloader/installer
├── catalog_scanner.py           # Cross-platform Steam game library detector (Windows & macOS)
├── doc_updater.py               # Documentation auto-synchronizer
├── test_telemetry.py            # Automated test suite for database and telemetry
├── run.sh                       # Master setup and execution script
├── webai_benchmarks.db          # SQLite database storing benchmark records
└── benchmark_results.csv        # Exported benchmark CSV dataset

Metadata

Release files for webai-scanner 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for webai-scanner 0.1.0
File Size Uploaded
webai_scanner-0.1.0.tar.gz 65.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for webai-scanner 0.1.0
File Interpreter ABI Platform
webai_scanner-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 80.9 kB

Release files / webai_scanner-0.1.0.tar.gz

Download URL webai_scanner-0.1.0.tar.gz
Size 65.6 kB
Tags Source
SHA-256 checksum
How to use checksums
e8c48c4cd543b0c469d0ab85a9b696c473bd46100dc75bccafc65a2c2025dc6e
BLAKE2b-256 checksum
How to use checksums
c389653782771ee14ca441807f8c740ce675210ab82c129db093cfe69848b98f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / webai_scanner-0.1.0-py3-none-any.whl

Download URL webai_scanner-0.1.0-py3-none-any.whl
Size 15.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ada33df5cd4136b99f5e4a290a3e545995130656cabc30b98892e10306d9fffd
BLAKE2b-256 checksum
How to use checksums
99ba1198fd1d40cfd97b37889aa70554a1902ea0bae24c53288d0ee875f0ddc1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page