Skip to main content

AI-powered geographic information crawler with multi-model support

Project description

GEO Crawler

中文文档

GEO Crawler is an installable Python CLI for running standardized sampling on AI platforms such as Kimi, Doubao, Yuanbao, DeepSeek, and Qianwen. It saves auditable raw outputs and includes the data models, judges, and metrics used for brand evaluation.

Recommended Runtime

For regular users, pipx is the recommended way to run the CLI. It installs the command into an isolated virtual environment, so you can run geo-crawler from any project folder without cloning this repository or managing dependencies manually.

python3 -m pip install --user pipx
python3 -m pipx ensurepath

pipx install geo-crawler
geo-crawler --help

Upgrade an installed version:

pipx upgrade geo-crawler

Install from Git before a version is published to PyPI:

pipx install git+https://github.com/fengclient/geo-crawler.git

Quick Start

For a step-by-step walkthrough of the full CLI loop, see CLI_WALKTHROUGH.md. The Chinese version is CLI_WALKTHROUGH.zh-CN.md.

1. Run Setup

Sampling needs an OpenAI-compatible model for the browser controller and a browser runtime. Run setup once in the project folder where you want .geo-auth/ and runs/ to live.

geo-crawler setup

setup writes a private .env.geo-crawler, creates .geo-auth/, and adds both to .gitignore. Later runtime commands discover .env.geo-crawler automatically. The generic .env file is still ignored by default.

2. Initialize Platform Auth

Before the first run, prepare a browser-use/Playwright storage_state for each target platform. The default auth directory is .geo-auth/ under the current working directory, so no extra flag is usually needed.

Interactive initialization:

geo-crawler auth init kimi
geo-crawler doctor auth kimi

Initialize every supported platform:

geo-crawler auth init all
geo-crawler doctor auth all

Import an existing storage state in CI, on a remote machine, or when you already have the file:

geo-crawler auth import kimi ./kimi_storage_state.json
geo-crawler doctor auth kimi

auth import overwrites the same platform file in the default auth directory. For example, geo-crawler auth import kimi ./kimi_storage_state.json writes .geo-auth/kimi_storage_state.json. It affects only kimi, not other platforms.

Remote Gem Browser initialization:

geo-crawler setup --preset volcengine-gem --discover
geo-crawler auth init kimi \
  --browser remote

Remote auth init uses Gem Browser managed lifecycle: the CLI creates or obtains a remote browser session, shows the viewer URL for QR code or phone verification login, exports .geo-auth/, and later reuses the same provider family for sampling. --yes only skips CLI confirmation prompts; it does not bypass platform login.

3. Run Sampling

Run commands from the project folder that owns your dataset, auth directory, and outputs. That keeps .geo-auth/, env files, and runs/ in the same workspace.

Single query:

geo-crawler sample run \
  --platform kimi \
  --query "20万左右电动车推荐" \
  --mode quick \
  --repeat 1 \
  --out runs

Dataset:

geo-crawler sample run \
  --platform kimi \
  --dataset dataset.json \
  --mode quick \
  --repeat 3 \
  --out runs

Config file:

geo-crawler sample run --config sampling.json

After sampling, outputs are written to runs/run-YYYYMMDD-HHMMSS-xxxxxx/:

runs/run-20260614-143025-a1b2c3/
├── manifest.json
├── inputs.json
├── outputs.json
├── summary.json
└── DONE

Configuration Files

Dataset

[
  {
    "id": "q1",
    "query": "20万左右电动车推荐"
  },
  {
    "id": "q2",
    "query": "适合家庭使用的混动车有哪些"
  }
]

Sampling Config

A config file is useful when platforms, inputs, and browser settings should be repeatable. The auth directory defaults to .geo-auth/; set browser.profile_dir or pass --profile-dir only when you need an override.

{
  "platforms": ["kimi", "qianwen"],
  "inputs": [
    {
      "id": "q1",
      "query": "20万左右电动车推荐"
    }
  ],
  "mode": "quick",
  "repeat": 3,
  "out": "runs",
  "controller": {
    "model": "gpt-4.1-mini",
    "api_key_env": "OPENAI_API_KEY"
  },
  "browser": {
    "cdp_url": "wss://example.com/browser"
  }
}

Env File

geo-crawler setup manages .env.geo-crawler for the ordinary path. This file is project-local, private, and discovered automatically when --env-file is omitted. Do not commit .geo-auth/, storage state JSON files, or env files containing real tokens.

OPENAI_MODEL=gpt-4.1-mini
OPENAI_API_KEY=sk-...
OPENAI_API_BASE=https://api.openai.com/v1
BROWSER_USE_LOGGING_LEVEL=warning
BROWSER_PROVIDER=local

Advanced users can still pass --env-file <path> to use a different environment file. The explicit file replaces .env.geo-crawler discovery for that command.

Remote CDP attach is for trusted private endpoints only:

geo-crawler setup --preset remote-cdp

For Gem Browser remote runs, use the Gem preset. The recommended path is managed lifecycle; raw CDP/viewer attach is only for debugging an existing session.

geo-crawler setup --preset volcengine-gem --discover

The older BROWSER_PROVIDER=volcengine Browser Session path remains available for programmatic remote sampling, but it is no longer the recommended human-login path because it does not provide a first-class viewer takeover flow.

Common Commands

# Show CLI help
geo-crawler --help
geo-crawler sample run --help

# Auth status
geo-crawler auth status all
geo-crawler doctor auth all

# Force re-login
geo-crawler auth init kimi --force

# Import from stdin
geo-crawler auth import kimi -

# Print JSON summary
geo-crawler sample run \
  --platform kimi \
  --query "20万左右电动车推荐" \
  --format json

Machine-Readable JSON

For scripts, CI, and AI agents, pass --format json explicitly. JSON mode writes a single JSON document to stdout with schema_version and status; progress, Rich tables, and human prompts stay out of stdout. sample run JSON includes artifact paths, counts, input source metadata, failed platforms, and safe recommended next actions when work remains.

geo-crawler auth status kimi --format json
geo-crawler doctor auth kimi --format json
geo-crawler doctor browser --format json
geo-crawler config inspect --format json
geo-crawler runs inspect runs/run-20260614-143025-a1b2c3 --format json
geo-crawler version --format json
geo-crawler sample run --platform kimi --query "20万左右电动车推荐" --format json

Use config inspect before a run to see effective controller, browser, and sampling config with source provenance. Use doctor browser for local, remote CDP, and Gem Browser readiness checks; live browser attachment only runs when --probe is passed. Use runs inspect after a run to summarize artifact health, platform counts, empty answers, diagnostics, and next actions without re-running sampling.

When a Gem Browser managed sampling run fails and the remote session can be retained, sample run --format json returns a short-lived live_browser.viewer_url and an open_live_browser recommended action. This is for inspecting the failed browser state only; manual actions in the viewer do not resume the current run or automatically retry failed inputs.

JSON output redacts secrets such as API keys, AK/SK, session tokens, cookies, storage state, and full CDP WebSocket URLs. Short-lived viewer URLs may appear only in commands that explicitly hand them to a user for login or failure inspection. Long-lived artifacts and runs inspect keep only redacted live-browser summaries; ordinary page URLs may appear unless credential-like query parameters need redaction.

Auth State Management

The recommended auth directory is .geo-auth/. The legacy browser_profiles/ path is still read for compatibility, but new users and automation should use auth init or auth import.

Supported platforms:

  • kimi
  • doubao
  • yuanbao
  • deepseek
  • qianwen
  • all for auth init, auth status, and doctor auth

Directory example:

.geo-auth/
├── kimi_storage_state.json
├── doubao_storage_state.json
├── yuanbao_storage_state.json
├── deepseek_storage_state.json
└── qianwen_storage_state.json

Developer Workflow

Clone the repository and use uv only for development, testing, or source changes. For regular operation, prefer pipx.

git clone https://github.com/fengclient/geo-crawler.git
cd geo-crawler

uv sync --all-extras --dev
uv run python -m pytest
uv run python -m geo_crawler.cli --help

In a development checkout, the same CLI can be run through the source entry point:

uv run python -m geo_crawler.cli sample run \
  --platform kimi \
  --query "20万左右电动车推荐"

The compatibility entry point python examples/geo_cli.py ... remains for old scripts, but it is no longer the recommended README path.

Architecture

The CLI has four main layers: command entry point, sampling configuration, platform query execution, and result storage. Auth state is managed by src/auth_state.py, and browser storage state is checked before sampling starts.

geo_crawler.cli
└── examples/geo_cli.py compatibility wrapper

src/sampling/
├── config.py
├── runner.py
└── storage.py

src/<platform>/
├── query.py
├── schema.py
├── tools.py
├── prompt_quick.md
└── prompt_deep.md

src/orchestrator/
├── base_judge.py
├── base_metric.py
├── judges/
└── metrics/

Troubleshooting

geo-crawler: command not found

Make sure pipx ensurepath has been run, then reopen your terminal.

python3 -m pipx ensurepath

doctor auth reports missing or expired auth

Reinitialize interactively, or import a fresh storage state.

geo-crawler auth init kimi --force
geo-crawler doctor auth kimi

--env-file does not take effect

Confirm the command explicitly passes --env-file, and that variable names match supported CLI configuration. .env is not loaded automatically.

Does a remote browser still need browser_profiles/?

No. Remote browser workflows should still use .geo-auth/<platform>_storage_state.json. browser_profiles/ is only a legacy compatibility path, not a dependency for the new flow.

Check installed package metadata

pipx list
pipx runpip geo-crawler show geo-crawler

Security Notes

.geo-auth/ and storage state JSON files are browser login credentials. Treat them as secrets. Do not commit them, paste them into chats or tickets, or expose them in logs. In CI, prefer private files or controlled secret injection.

License

MIT

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

geo_crawler-0.4.0.tar.gz (278.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

geo_crawler-0.4.0-py3-none-any.whl (226.0 kB view details)

Uploaded Python 3

File details

Details for the file geo_crawler-0.4.0.tar.gz.

File metadata

  • Download URL: geo_crawler-0.4.0.tar.gz
  • Upload date:
  • Size: 278.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.10 {"installer":{"name":"uv","version":"0.9.10"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for geo_crawler-0.4.0.tar.gz
Algorithm Hash digest
SHA256 593e6f504e3c7f879f5f89129988ad208aa149db6cabdcb7d2eeb862a4a3837d
MD5 9c457470a9b75a2cf98e09c7e9059558
BLAKE2b-256 1823fcad924db05356c5b5d24baffbaa96e9acadc75be6ad105f2b750cb38e10

See more details on using hashes here.

File details

Details for the file geo_crawler-0.4.0-py3-none-any.whl.

File metadata

  • Download URL: geo_crawler-0.4.0-py3-none-any.whl
  • Upload date:
  • Size: 226.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.9.10 {"installer":{"name":"uv","version":"0.9.10"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for geo_crawler-0.4.0-py3-none-any.whl
Algorithm Hash digest
SHA256 4fa010e8b1d87bd9c51ed23d611729af2786a66711c9f329f0b63e674af3e6eb
MD5 5943cf6ded640fbe22eb1fe0da3b1a7d
BLAKE2b-256 609a028f61d7dd77f5219653abc67be2627c2794cfc7655edf990476814acc88

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page