Skip to main content

go-spider

Intent-driven web research: discover sources, collect native-language documents, follow evidence gaps, and return a source-cited answer—or an explicit partial result when the evidence or budget is insufficient.

v0.2.0 is a standalone Python library release. The distribution stays go-spider; its redesigned implementation is imported as chimera. It is not backward-compatible with v0.1.0's spider_core API or spider CLI. Python 3.11+ is required. Some planned browser/document adapters and public-corpus acceptance are still in progress; see the limitations below.

go-spider is an independent library. Supply your own search provider, self-hosted model services, extraction policies and graph profile.

What is implemented

  • Goal and intent loops. Collect from configured seeds with GoalLoop, or use ResearchLoop to plan questions, discover sources through an injected search provider, assess gaps, and draft/review an evidence-cited answer.
  • Local models first. Configured, already-served self-hosted models supply planning, judging, answer generation, review and embeddings. Compatible HTTP interfaces are supported; no external LLM fallback, model weights or model server are included.
  • Bounded direct and Tor HTTP. Native v3 onion collection and open-web requests through Tor share the same route policy. Public-network validation, pinned DNS for direct requests, per-hop redirect checks, robots policy, concurrency/rate limits and retries are accounted for before returning data. A failed Tor route never silently falls back to direct access.
  • Isolated JavaScript rendering. Patchright runs in a network-isolated Linux worker; the parent fetch boundary handles its permitted HTTP resources, redirects and accounting. Browser binaries are explicitly configured and verified, not downloaded on import. Camoufox/nodriver adapters remain planned.
  • Native extraction. Configured HTML fit-Markdown, adaptive locator profiles, language detection, DOCX tables and native PDF text preserve raw bytes beside extracted native-language text. Full PDF/OCR and Marker acceptance remain open.
  • Relevance and deduplication. An injected self-hosted encoder scores native text/windows and observed links against pinned reference vectors. Keyword/semantic ranking, encoding budgets, canonical URL handling, SHA-256 and configured near-duplicate grouping retain source-qualified evidence. Similarity is not a calibrated probability or a substitute for a verdict.
  • Graphs and audit records. A configurable research graph starts with the intent. Typed harvests, receipts and ledger rows retain configuration, transport, model-call spend, omissions, verdicts, refusals and source hashes. Operational discovery traces are distinct from evidence-supported claims.

Every candidate reaching acceptance receives an accept/reject/hold verdict. A hold gets a second model pass. An intent is marked answered only after the coverage, citations, review and configured confidence checks pass; budget exhaustion is not silently presented as success.

Installation

Use a dedicated virtual environment:

python3.11 -m venv .venv
. .venv/bin/activate
python -m pip install 'go-spider==0.2.0'

Install the adapters you intend to configure:

python -m pip install 'go-spider[html,documents,browser]==0.2.0'

The base package contains the typed core, HTTP/Tor transport, research/search and self-hosted model/embedding clients. Extras add pinned HTML, document and Patchright dependencies. They do not install an inference server, browser binary, Tor daemon or PDF/OCR model artifacts. The isolated browser adapter requires Linux, a compatible explicitly supplied Chromium binary and Bubblewrap; other operating systems have not been accepted for that adapter.

Configuration and API

Operational choices are typed, versioned configuration—not Python constants: scope, budgets, endpoints, model identities/revisions, thresholds, private worker directories, browser provenance and direct/Tor policy. Parse once with ChimeraConfig.from_toml(Path(...)); inject the matching collaborators.

The examples are non-active templates. Replace invalid endpoints, contact information, private paths and model identifiers; reference-vector fixtures are not production relevance data. Adapter blocks belong in the main configuration under their named keys, not as unrelated root settings. Supply any model credential separately in memory, only to its authorized exact endpoint; configuration is not credential or destination approval.

Download the starter configuration, or copy it from the repository:

curl --fail --proto '=https' --tlsv1.2 \
  https://raw.githubusercontent.com/ginkorea/spider/v0.2.0/examples/chimera.toml \
  --output chimera.toml

An entirely offline smoke example, using explicitly named test doubles:

import asyncio
from pathlib import Path

from chimera import ChimeraConfig, Goal, GoalLoop, Harvest, Scope
from chimera.doubles import FakeExtractor, FakeJudge, FakeRoute, KeywordScorer
from chimera.fetch import FetchLadder


async def main() -> None:
    config = ChimeraConfig.from_toml(Path("chimera.toml"))
    collector = GoalLoop(
        config=config,
        fetcher=FetchLadder((FakeRoute(),)),
        extractor=FakeExtractor(),
        scorer=KeywordScorer(),
        judge=FakeJudge(),
    )
    result = await collector.run(
        Goal(text="ports", seeds=("https://example.org/start",)),
        Scope(allowed_hosts=("example.org",), max_depth=1, content_types=("text/html",)),
    )
    # Reader revalidates retained sources, provenance and receipt accounting.
    restored = Harvest.model_validate_json(result.model_dump_json())
    print(restored.receipt.stop_reason)


asyncio.run(main())

This example makes no network or real model calls and proves no research accuracy. For actual collection bind CurlRoute, configured extraction and scoring adapters, and the self-hosted judge through their ports. For intent-only research, inject those into ResearchLoop alongside GroundedSearch, IntentPlanner, ResearchAnalyst and AnswerReviewer, then call run(ResearchRequest(intent="your research question")). SearxSearch is the implemented search adapter. The following guides cover the concrete wiring:

Area Guide
Intent, discovery, coverage and answer review Intent research
Model roles, credentials and native evidence context Self-hosted models
Reference vectors, encoding and frontier ranking Embedding scoring
Open web and native onion routing Tor policy
Browser isolation, resources and redirects Browser rendering
HTML extraction and adaptive locators HTML extraction
DOCX/native PDF and offline artifacts Document extraction
Canonical and near-duplicate source evidence Deduplication
Configurable research graphs Research graph
Architecture, contracts and completion tracker Specification

Scope and limitations

Robots are honored by default. An override requires a recorded, reasoned, exact-host configuration decision; it does not disable CAPTCHA, login, paywall or challenge refusal. There is no challenge solver or authenticated-site bypass. Tor routing is a transport capability, not a guarantee of anonymity or authority to access a source.

Source/search requests run on the host executing the crawler. Model control is a separate private-service boundary. Your application owns authorization, deployment and storage; no external scheduler or registry is required to import or use the library.

Still required for the complete planned spider: the remaining browser adapters, representative publisher/locator acceptance, full PDF/OCR and Marker validation, one-hop references/cited-by expansion, real served-model quality/admission and calibrated decision policy and live runtime/egress acceptance. The repository's detailed tracker retains those requirements; this release does not erase them or describe fixture results as real-world model accuracy.

Development

git clone https://github.com/ginkorea/spider.git
cd spider
uv sync --locked --extra html --extra documents --extra browser --python 3.11
# Explicit browser/isolation paths are required for the complete gate.
export CHIMERA_TEST_BROWSER=/absolute/path/to/compatible/chrome
export CHIMERA_TEST_ISOLATOR=/absolute/path/to/bwrap
bash scripts/gate.sh
uv build --no-sources

The gate checks the actual interpreter/import path, lockfile, Ruff, strict mypy and the entire test suite. Browser tests refuse absent acceptance prerequisites rather than pretending they ran. Evidence records distinguish protocol fixtures, installed-artifact checks, public-corpus acceptance and production activation.

v0.1.0 migration

pip install --upgrade go-spider now installs the rewritten library. Existing spider_core imports and the old spider command are not supplied by v0.2.0; migrate to chimera and explicit collaborators/configuration, or pin go-spider==0.1.0 while migrating. The old cloud-client, VPN-manager and implicit fallback design is not retained. Prototype source remains in Git history.

See CHANGELOG.md for release changes. Josh Gompert maintains the project at ginkorea/spider. Licensed under MIT, matching the existing PyPI licence declaration.

Metadata

Release files for go-spider 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for go-spider 0.2.0
File Size Uploaded
go_spider-0.2.0.tar.gz 503.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for go-spider 0.2.0
File Interpreter ABI Platform
go_spider-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 613.8 kB

Release files / go_spider-0.2.0.tar.gz

Download URL go_spider-0.2.0.tar.gz
Size 503.3 kB
Tags Source
SHA-256 checksum
How to use checksums
d69de25630e78b0b53fa753c96e13635373b125109c1f350d5c3f4080d9bfeca
BLAKE2b-256 checksum
How to use checksums
db49261e70c4db01cd1fe8768b71ebf644867283279d37fc8fec7b6654baef74
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.16

Release files / go_spider-0.2.0-py3-none-any.whl

Download URL go_spider-0.2.0-py3-none-any.whl
Size 110.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
134e7bea4a4736dc47fd3fdb3c73225c2a334ef004d43fc5ae019edb5cccbc98
BLAKE2b-256 checksum
How to use checksums
0d652e3a5ca79f2230e3f4be215160a936ac71c4035e5223c3200b7ab7b61dc5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.16

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page