Skip to main content

⟠ GEO-Scope

Empirical AI Answer Visibility Measurement Framework

An open-source framework for empirical measurement of AI answer visibility, entity mentions, recommendations, and citations across generative AI systems.

Language فارسی Türkçe Azərbaycan العربية

CI License: MIT Scientific Foundation Measurement Contract Research Paper Outline Golden Parser Security Audit

Introduction • What It Measures • Scientific Foundation • Benchmarks • Reproducibility • Research • Quickstart • MCP


1. Introduction

Generative AI systems and search-grounded answer engines are rapidly becoming the primary discovery layer for users seeking products, vendors, services, and factual insights.

GEO-Scope is an evidence-first, open-source measurement framework designed to empirically quantify and preserve auditable evidence of how generative AI systems surface entities. It records, normalizes, and analyzes observable AI completions under documented, neutral prompt sets without relying on speculative ranking algorithms or ungrounded claims.

Core Observable Outputs Measured:

  • Entity Mentions: Observable presence of brands, products, technologies, and public figures in generated text.
  • Recommendations: Explicit linguistic endorsements and ordered top-position recommendations.
  • Citations: Grounding source URLs and referenced web domains returned by search-augmented models.
  • Attribution: Textual credit linking specific facts, statistics, or claims to source entities.
  • Provider Differences: Distributional shifts between live search-grounded answer engines and parametric foundation LLMs.
  • Multilingual Behavior: Cross-lingual response variations across 26+ evaluated languages.

2. What GEO-Scope Measures

GEO-Scope enforces a strict taxonomic separation between four independent visibility dimensions:

┌─────────────────────────────────────────────────────────────────────────┐
│                        AI RESPONSE VISIBILITY MATRIX                    │
├───────────────────┬─────────────────────────────────────────────────────┤
│ Mention           │ Did the entity appear anywhere in the completion?   │
│ Recommendation    │ Was the entity explicitly endorsed or recommended?  │
│ Citation          │ Was a source URL or grounding domain link provided? │
│ Attribution       │ Was specific data/claim textually credited to it?  │
│ Rank              │ Extracted ONLY when a valid ordered list exists.   │
└───────────────────┴─────────────────────────────────────────────────────┘
  1. Mention (mentioned: true/false):
    • Captures whether the target entity (or associated canonical aliases/founders) appeared in the generated completion.
    • Evaluated via Unicode NFKC normalization, Arabic/Persian letter unification, Zero-Width Non-Joiner (ZWNJ) handling, and negative homonym collision filtering.
  2. Recommendation (recommended: true/false):
    • Strict rule: mentioned != recommended.
    • Evaluated based on explicit linguistic recommendation markers (e.g., "We recommend...", "Top pick", "گزینه پیشنهادی") or inclusion in an ordered list answering a recommendation query.
  3. Citation (cited: true/false):
    • Identifies presence of target entity web domains in grounding references, markdown hyperlinks, or structured provider citation chunks.
  4. Attribution (attributed: true/false):
    • Distinct from citation: detects explicit textual sourcing phrasing (e.g., "According to [Entity]...", "طبق گزارش [موجودیت]") even if an active URL link was omitted by the model.
  5. Rank (rank: 1..N | null):
    • Extracted strictly from numbered lists or ordinal items. If an informational question yields an unranked mention, rank is set to null to prevent artificial ranking bias.

3. What GEO-Scope Does NOT Measure

To maintain scientific integrity, GEO-Scope clearly outlines its epistemic boundaries:

  • ❌ It does NOT reverse-engineer internal ranking algorithms: GEO-Scope observes external API completions; it cannot inspect internal model weights, attention matrices, or proprietary ranking formulas.
  • ❌ It does NOT inspect hidden training data: Observed entity knowledge reflects generated outputs, not full visibility into private training corpora.
  • ❌ It does NOT claim causal ranking factors: All reported metrics represent descriptive statistical associations under documented prompts, not causal guarantees.
  • ❌ It does NOT guarantee SEO or AI visibility improvements: Measurements provide historical observation, not predictive visibility promises.
  • ❌ It does NOT treat simulation as live empirical data: Simulation fixtures are strictly quarantined for testing and CI.

4. Architecture

GEO-Scope operates as a modular, six-stage evidence pipeline:

  ┌─────────────────────────────────────────────────────────────┐
  │                       Provider Layer                        │
  │   ┌──────────────────────────┐  ┌────────────────────────┐  │
  │   │  Search Answer Engines   │  │ Parametric Base LLMs   │  │
  │   │  (Perplexity, Gemini...) │  │ (OpenAI, Claude...)    │  │
  │   └──────────────────────────┘  └────────────────────────┘  │
  └──────────────────────────────┬──────────────────────────────┘
                                 │
                                 ▼
  ┌─────────────────────────────────────────────────────────────┐
  │                     Measurement Engine                      │
  │     (Prompt Provenance · Zero Silent Fallback · Runs)       │
  └──────────────────────────────┬──────────────────────────────┘
                                 │
                                 ▼
  ┌─────────────────────────────────────────────────────────────┐
  │                    Raw Response Storage                     │
  │    (Unparsed API Payloads · Latency · Token Usage)          │
  └──────────────────────────────┬──────────────────────────────┘
                                 │
                                 ▼
  ┌─────────────────────────────────────────────────────────────┐
  │                     Observation Parser                      │
  │    (Multi-Lingual Normalizer · Homonyms · Citations)        │
  └──────────────────────────────┬──────────────────────────────┘
                                 │
                                 ▼
  ┌─────────────────────────────────────────────────────────────┐
  │                     Metrics Calculation                     │
  │    (OMR · Rec Share · Citation Rate · Honest Denominators)  │
  └──────────────────────────────┬──────────────────────────────┘
                                 │
                                 ▼
  ┌─────────────────────────────────────────────────────────────┐
  │                 Reports + Replayable Bundle                 │
  │    (JSONL Bundles · SHA-256 Checksums · Markdown Summaries) │
  └─────────────────────────────────────────────────────────────┘

5. Execution Modes

GEO-Scope provides three mutually exclusive execution modes:

demo (Simulation Fixture)

  • Purpose: Rapid offline testing, development fixtures, and CI validation.
  • Behavior: Uses local mock completions with prefixed IDs (simulated_*) and a clear simulation banner.
  • Guarantee: Simulation data is strictly rejected by the release quality gate and can never enter published empirical benchmarks.

measure (Live Empirical Execution)

  • Purpose: Real-world observation runs against live generative AI endpoints.
  • Behavior: Dispatches neutral prompt bundles to configured API providers with zero silent fallback.
  • Preservation: Saves full unmodified payloads to raw_responses.jsonl with exact model governance metadata (requested_provider, actual_provider, search_grounded).

replay (Deterministic Offline Replay)

  • Purpose: Independent auditability and benchmark verification without API calls or cost.
  • Behavior: Reruns the observation parser and metric calculations directly against preserved raw_responses.jsonl.
  • Integrity: Verifies that recomputed metrics match published results bit-for-bit.

6. Measurement Contract v1 & Scientific Foundation

GEO-Scope does not claim universal AI visibility truth. It measures empirical observations under declared, reproducible measurement configurations.

Core Measurement Principles

  1. Mention Definition: A response-level binary observation indicating whether the target entity appears at least once in the completion. Multiple mentions in a single answer do not artificially inflate response-level mention counts.
  2. Citation Separation: Strict 4-way separation between entity_mentioned in text, target_domain_cited (root domain), target_url_cited (deep link), and third_party_source_cited (external authority/review links). Mention and citation are never treated as equivalent.
  3. Recommendation Semantics: Evaluated as true only when the model semantically recommends, selects, or endorses the entity. Ambiguous detections are gated and marked experimental.
  4. Comparability Rules: Machine-readable comparability verification. Two studies are marked comparable: true only when prompt universe, market/language, provider/model family, measurement definitions, and observation windows match.
  5. Raw Evidence Traceability: Every public observation is linked to prompt ID, raw response or cryptographic SHA-256 hash (response_hash_sha256), extracted entities, citations, and execution configuration hash.

7. Published Benchmark Releases

GEO-Scope maintains immutable, peer-review-ready benchmark releases under benchmark/releases/ (see complete Releases & Milestones Timeline):

Benchmark Release Prompt Count Observations Providers Cryptographic Status Documentation
global-ai-answers-2026.2 500 prompts (50 countries) 45,698 obs 4 models SHA-256 Verified Paper Draft
global-ai-answers-2026.2-pilot 100 prompts (10 countries) 8,940 obs 4 models SHA-256 Verified Pilot Report
global-ai-answers-2026.1 34 prompts (7 regions) Baseline obs 4 models SHA-256 Verified Methodology
geo-seo-digital-agency-iran-2026.1 30 prompts (5 intent strata) 120 completions 4 models SHA-256 Verified Agency Report

Every release bundle contains:

  • manifest.json: Dataset metadata, provider matrix, and schema version (measurement_contract_version: "1.0").
  • prompts.jsonl: Neutral, categorized prompts.
  • raw_responses.jsonl: Verbatim API completion payloads.
  • observations.jsonl: Granular extracted observation records.
  • metrics.json: Aggregated metrics with explicit failure denominators.
  • checksums.sha256: SHA-256 hashes of all artifacts.

8. Reproducibility & Auditability

1. Cryptographic SHA-256 Verification

Verify that dataset files have not been modified or corrupted:

geo-scope benchmark verify --dataset benchmark/releases/global-ai-answers-2026.2

2. Zero-Network Replay Workflow

Replay metrics directly from preserved raw responses without executing live API calls:

geo-scope replay \
  --bundle benchmark/releases/global-ai-answers-2026.2 \
  --out-dir output/replay_2026_2

3. Golden Parser Evaluation

Evaluate the deterministic multi-lingual parser against human-labeled ground truth:

geo-scope parser evaluate --golden-set benchmark/golden_sets/v1

Golden Set Benchmark Results (v1, 220 examples across 5 languages):

  • Mention F1: 99.75% (Precision: 99.51%, Recall: 100.00%)
  • Recommendation F1: 100.00% (Precision: 100.00%, Recall: 100.00%)
  • Citation F1: 100.00% (Precision: 100.00%, Recall: 100.00%)
  • Attribution F1: 91.56% (Precision: 100.00%, Recall: 84.44%)
  • Wrong Entity (Homonym) F1: 96.97%
  • Rank Accuracy: 100.00%

9. Research & Documentation


10. Installation & Usage

Installation

# Clone the repository
git clone https://github.com/tmolavi/geo-scope.git
cd geo-scope

# Install package in editable mode
pip install -e .

Quick Commands

# 1. Run local simulation fixture demo
geo-scope demo

# 2. Execute live empirical measurement (requires API credentials)
geo-scope measure \
  --entities entities/iran-seo-agencies.json \
  --prompts examples/research_run/prompts.jsonl \
  --mode live \
  --providers perplexity_sonar,gemini_grounding \
  --out-dir output/live_run_01

# 3. Deterministic offline replay
geo-scope replay \
  --bundle output/live_run_01 \
  --out-dir output/replay_run_01

# 4. Launch interactive local research dashboard
geo-scope serve --host 127.0.0.1 --port 8000

11. Model Context Protocol (MCP)

GEO-Scope includes a native MCP Server (stdio), enabling AI coding assistants and agents (Claude Desktop, Cursor, Antigravity) to query visibility benchmarks and inspect entity evidence chains directly:

{
  "mcpServers": {
    "geo-scope": {
      "command": "/absolute/path/to/geo-scope/.venv/bin/geo-scope",
      "args": ["mcp"]
    }
  }
}

See the Client Integrations Guide for full configuration details.


Citation & Authorship

Developed by Taqi Molavi (Senior SEO Strategist & GEO Systems Architect).
Part of the Molavi GEO Pyramid research initiative.

@software{molavi2026geoscope,
  author = {Molavi, Taqi},
  title = {GEO-Scope: Empirical AI Answer Visibility Measurement Framework},
  year = {2026},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/tmolavi/geo-scope}},
  note = {Personal Homepage: https://molavi.pro/}
}

License

This project is licensed under the MIT License — Copyright (c) 2026 تقی مولوی (Taqi Molavi).

Metadata

Release files for geo-scope 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for geo-scope 0.3.0
File Size Uploaded
geo_scope-0.3.0.tar.gz 324.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for geo-scope 0.3.0
File Interpreter ABI Platform
geo_scope-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 625.0 kB

Release files / geo_scope-0.3.0.tar.gz

Download URL geo_scope-0.3.0.tar.gz
Size 324.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c47c81be888ba3e9cf65fac657abb8afb4274b57e538f592aa1dbd34afc613fb
BLAKE2b-256 checksum
How to use checksums
c77d49b1f412a6863bd17b5c3e179c57529fbf5f1484458c7e5080c1586e303b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release files / geo_scope-0.3.0-py3-none-any.whl

Download URL geo_scope-0.3.0-py3-none-any.whl
Size 301.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
51344f1ed9072b640dd287f173c20567621ad73d24ef514ac364018876788147
BLAKE2b-256 checksum
How to use checksums
73188c71bc6c4029df92ff5f91fedc8c4939cf454d67a5c02faf9a63bc7fab54
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.13

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page