AI-powered geographic information crawler with multi-model support
Project description
GEO Crawler
GEO Crawler is an installable Python CLI for running standardized sampling on AI platforms such as Kimi, Doubao, Yuanbao, DeepSeek, and Qianwen. It saves auditable raw outputs and includes the data models, judges, and metrics used for brand evaluation.
Recommended Runtime
For regular users, pipx is the recommended way to run the CLI. It installs the command into an isolated virtual environment, so you can run geo-crawler from any project folder without cloning this repository or managing dependencies manually.
python3 -m pip install --user pipx
python3 -m pipx ensurepath
pipx install geo-crawler
geo-crawler --help
Upgrade an installed version:
pipx upgrade geo-crawler
Install from Git before a version is published to PyPI:
pipx install git+https://github.com/fengclient/geo-crawler.git
Quick Start
For a step-by-step walkthrough of the full CLI loop, see CLI_WALKTHROUGH.md. The Chinese version is CLI_WALKTHROUGH.zh-CN.md.
1. Run Setup
Sampling needs an OpenAI-compatible model for the browser controller and a browser runtime. Run setup once in the project folder where you want .geo-auth/ and runs/ to live.
geo-crawler setup
setup writes a private .env.geo-crawler, creates .geo-auth/, and adds both to .gitignore. Later runtime commands discover .env.geo-crawler automatically. The generic .env file is still ignored by default.
2. Initialize Platform Auth
Before the first run, prepare a browser-use/Playwright storage_state for each target platform. The default auth directory is .geo-auth/ under the current working directory, so no extra flag is usually needed.
Interactive initialization:
geo-crawler auth init kimi
geo-crawler doctor auth kimi
Initialize every supported platform:
geo-crawler auth init all
geo-crawler doctor auth all
Import an existing storage state in CI, on a remote machine, or when you already have the file:
geo-crawler auth import kimi ./kimi_storage_state.json
geo-crawler doctor auth kimi
auth import overwrites the same platform file in the default auth directory. For example, geo-crawler auth import kimi ./kimi_storage_state.json writes .geo-auth/kimi_storage_state.json. It affects only kimi, not other platforms.
Remote Gem Browser initialization:
geo-crawler setup --preset volcengine-gem --discover
geo-crawler auth init kimi \
--browser remote
Remote auth init uses Gem Browser managed lifecycle: the CLI creates or obtains a remote browser session, shows the viewer URL for QR code or phone verification login, exports .geo-auth/, and later reuses the same provider family for sampling. --yes only skips CLI confirmation prompts; it does not bypass platform login.
3. Run Sampling
Run commands from the project folder that owns your dataset, auth directory, and outputs. That keeps .geo-auth/, env files, and runs/ in the same workspace.
Single query:
geo-crawler sample run \
--platform kimi \
--query "20万左右电动车推荐" \
--mode quick \
--repeat 1 \
--out runs
Dataset:
geo-crawler sample run \
--platform kimi \
--dataset dataset.json \
--mode quick \
--repeat 3 \
--out runs
Config file:
geo-crawler sample run --config sampling.json
After sampling, outputs are written to runs/run-YYYYMMDD-HHMMSS-xxxxxx/:
runs/run-20260614-143025-a1b2c3/
├── manifest.json
├── inputs.json
├── outputs.json
├── summary.json
└── DONE
Configuration Files
Dataset
[
{
"id": "q1",
"query": "20万左右电动车推荐"
},
{
"id": "q2",
"query": "适合家庭使用的混动车有哪些"
}
]
Sampling Config
A config file is useful when platforms, inputs, and browser settings should be repeatable. The auth directory defaults to .geo-auth/; set browser.profile_dir or pass --profile-dir only when you need an override.
{
"platforms": ["kimi", "qianwen"],
"inputs": [
{
"id": "q1",
"query": "20万左右电动车推荐"
}
],
"mode": "quick",
"repeat": 3,
"out": "runs",
"controller": {
"model": "gpt-4.1-mini",
"api_key_env": "OPENAI_API_KEY"
},
"browser": {
"cdp_url": "wss://example.com/browser"
}
}
Env File
geo-crawler setup manages .env.geo-crawler for the ordinary path. This file is project-local, private, and discovered automatically when --env-file is omitted. Do not commit .geo-auth/, storage state JSON files, or env files containing real tokens.
OPENAI_MODEL=gpt-4.1-mini
OPENAI_API_KEY=sk-...
OPENAI_API_BASE=https://api.openai.com/v1
BROWSER_USE_LOGGING_LEVEL=warning
BROWSER_PROVIDER=local
Advanced users can still pass --env-file <path> to use a different environment file. The explicit file replaces .env.geo-crawler discovery for that command.
Remote CDP attach is for trusted private endpoints only:
geo-crawler setup --preset remote-cdp
For Gem Browser remote runs, use the Gem preset. The recommended path is managed lifecycle; raw CDP/viewer attach is only for debugging an existing session.
geo-crawler setup --preset volcengine-gem --discover
The older BROWSER_PROVIDER=volcengine Browser Session path remains available for programmatic remote sampling, but it is no longer the recommended human-login path because it does not provide a first-class viewer takeover flow.
Common Commands
# Show CLI help
geo-crawler --help
geo-crawler sample run --help
# Auth status
geo-crawler auth status all
geo-crawler doctor auth all
# Force re-login
geo-crawler auth init kimi --force
# Import from stdin
geo-crawler auth import kimi -
# Print JSON summary
geo-crawler sample run \
--platform kimi \
--query "20万左右电动车推荐" \
--format json
Machine-Readable JSON
For scripts, CI, and AI agents, pass --format json explicitly. JSON mode writes a single JSON document to stdout with schema_version and status; progress, Rich tables, and human prompts stay out of stdout. sample run JSON includes artifact paths, counts, input source metadata, failed platforms, and safe recommended next actions when work remains.
geo-crawler auth status kimi --format json
geo-crawler doctor auth kimi --format json
geo-crawler doctor browser --format json
geo-crawler config inspect --format json
geo-crawler runs inspect runs/run-20260614-143025-a1b2c3 --format json
geo-crawler version --format json
geo-crawler sample run --platform kimi --query "20万左右电动车推荐" --format json
Use config inspect before a run to see effective controller, browser, and sampling config with source provenance. Use doctor browser for local, remote CDP, and Gem Browser readiness checks; live browser attachment only runs when --probe is passed. Use runs inspect after a run to summarize artifact health, platform counts, empty answers, diagnostics, and next actions without re-running sampling.
When a Gem Browser managed sampling run fails and the remote session can be retained, sample run --format json returns a short-lived live_browser.viewer_url and an open_live_browser recommended action. This is for inspecting the failed browser state only; manual actions in the viewer do not resume the current run or automatically retry failed inputs.
JSON output redacts secrets such as API keys, AK/SK, session tokens, cookies, storage state, and full CDP WebSocket URLs. Short-lived viewer URLs may appear only in commands that explicitly hand them to a user for login or failure inspection. Long-lived artifacts and runs inspect keep only redacted live-browser summaries; ordinary page URLs may appear unless credential-like query parameters need redaction.
Auth State Management
The recommended auth directory is .geo-auth/. The legacy browser_profiles/ path is still read for compatibility, but new users and automation should use auth init or auth import.
Supported platforms:
kimidoubaoyuanbaodeepseekqianwenallforauth init,auth status, anddoctor auth
Directory example:
.geo-auth/
├── kimi_storage_state.json
├── doubao_storage_state.json
├── yuanbao_storage_state.json
├── deepseek_storage_state.json
└── qianwen_storage_state.json
Developer Workflow
Clone the repository and use uv only for development, testing, or source changes. For regular operation, prefer pipx.
git clone https://github.com/fengclient/geo-crawler.git
cd geo-crawler
uv sync --all-extras --dev
uv run python -m pytest
uv run python -m geo_crawler.cli --help
In a development checkout, the same CLI can be run through the source entry point:
uv run python -m geo_crawler.cli sample run \
--platform kimi \
--query "20万左右电动车推荐"
The compatibility entry point python examples/geo_cli.py ... remains for old scripts, but it is no longer the recommended README path.
Architecture
The CLI has four main layers: command entry point, sampling configuration, platform query execution, and result storage. Auth state is managed by src/auth_state.py, and browser storage state is checked before sampling starts.
geo_crawler.cli
└── examples/geo_cli.py compatibility wrapper
src/sampling/
├── config.py
├── runner.py
└── storage.py
src/<platform>/
├── query.py
├── schema.py
├── tools.py
├── prompt_quick.md
└── prompt_deep.md
src/orchestrator/
├── base_judge.py
├── base_metric.py
├── judges/
└── metrics/
Troubleshooting
geo-crawler: command not found
Make sure pipx ensurepath has been run, then reopen your terminal.
python3 -m pipx ensurepath
doctor auth reports missing or expired auth
Reinitialize interactively, or import a fresh storage state.
geo-crawler auth init kimi --force
geo-crawler doctor auth kimi
--env-file does not take effect
Confirm the command explicitly passes --env-file, and that variable names match supported CLI configuration. .env is not loaded automatically.
Does a remote browser still need browser_profiles/?
No. Remote browser workflows should still use .geo-auth/<platform>_storage_state.json. browser_profiles/ is only a legacy compatibility path, not a dependency for the new flow.
Check installed package metadata
pipx list
pipx runpip geo-crawler show geo-crawler
Security Notes
.geo-auth/ and storage state JSON files are browser login credentials. Treat them as secrets. Do not commit them, paste them into chats or tickets, or expose them in logs. In CI, prefer private files or controlled secret injection.
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file geo_crawler-0.4.0.tar.gz.
File metadata
- Download URL: geo_crawler-0.4.0.tar.gz
- Upload date:
- Size: 278.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.10 {"installer":{"name":"uv","version":"0.9.10"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
593e6f504e3c7f879f5f89129988ad208aa149db6cabdcb7d2eeb862a4a3837d
|
|
| MD5 |
9c457470a9b75a2cf98e09c7e9059558
|
|
| BLAKE2b-256 |
1823fcad924db05356c5b5d24baffbaa96e9acadc75be6ad105f2b750cb38e10
|
File details
Details for the file geo_crawler-0.4.0-py3-none-any.whl.
File metadata
- Download URL: geo_crawler-0.4.0-py3-none-any.whl
- Upload date:
- Size: 226.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.9.10 {"installer":{"name":"uv","version":"0.9.10"},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4fa010e8b1d87bd9c51ed23d611729af2786a66711c9f329f0b63e674af3e6eb
|
|
| MD5 |
5943cf6ded640fbe22eb1fe0da3b1a7d
|
|
| BLAKE2b-256 |
609a028f61d7dd77f5219653abc67be2627c2794cfc7655edf990476814acc88
|