Verifiable autonomous research workspaces: agent-driven research where every answer cites provenance-tracked sources or blocks with a structured evidence request.
Project description
EvidenceWiki
Answers you can audit. This project builds verifiable autonomous research workspaces: any LLM agent can drive the research loop, but every answer must cite normalized source records with provenance, every state transition is owned by a deterministic script, and a question that lacks required evidence blocks with a structured source request instead of degrading into a weakly-cited answer.
It ships as reusable tooling: a configurable research wiki starter,
deterministic scripts for inventory/normalization/linting and the question
lifecycle, optional domain packs, and tests that prove a minimally described
research project can be initialized and validated without hand-editing
research.yml.
Why It Is Different
Most research agents produce a one-shot report that cites whatever they happened to retrieve. This system takes the opposite bet:
- A persistent workspace, not a disposable report. Research accumulates in a versioned wiki where every claim cites a source ID that resolves to a normalized record with provenance. Answers are regenerable, lintable, and exportable as structured JSON.
- Evidence gates, not vibes. High-stakes questions carry coverage
manifests with facet-typed requirements (for example
official_guidanceorstandards_registry_reference). A failing facet blocks the question with a structured source request; it never silently downgrades the answer. Seeworkspace-template/docs/coverage-manifest.md. - Blocked, not fabricated. "No such paper exists" claims require a
claim_proberecord: provider, query, result counts, and an explicit limitation statement — bounded non-confirmation instead of invented citations. - Multi-agent-safe by construction. Question claiming, per-run budgets,
stale-claim recovery, and deterministic
blocked → reopentransitions make unattended fleets coordinable. LLM skills do the judgment; scripts own every state transition, which is what makes the behavior testable. - Source content is data, never instructions. Acquisition fails closed on
TLS/DNS/redirect/media-type/size violations, provenance-incomplete
deliveries are quarantined, and provenance URLs are never auto-fetched. See
workspace-template/docs/prompt-injection-hardening.md.
Five-Minute Tour
The goal: start with an empty source directory and let an agent produce a cited, exportable answer from evidence that it discovers and acquires during the run. Install the package and create a provider-enabled scientific workspace:
python3 -m pip install evidence-wiki
evidence-wiki deploy \
--target solid-state-batteries \
--project-name solid-state-batteries \
--project-description "Survey of solid-state battery electrolyte research" \
--domain-pack general-science \
--discovery-provider arxiv \
--discovery-provider openalex \
--acquisition-provider arxiv \
--acquisition-provider openalex
cd solid-state-batteries
The repeated provider flags are explicit network authorization. The
general-science pack recommends these providers but does not enable them by
itself. arXiv needs no credential; OpenAlex can use OPENALEX_API_KEY from the
process environment for authenticated quotas. Never put credentials in
research.yml.
Seed the backlog with a question (schema in
workspace-template/docs/question-api.md):
cat > batch.yaml <<'EOF'
schema_version: "1.0"
questions:
- question: "Which solid electrolyte families report room-temperature ionic conductivity above 1 mS/cm?"
id: electrolyte-conductivity
priority: high
EOF
evidence-wiki questions add --target . --from-file batch.yaml
Run the managed orchestrator with an installed agent CLI:
evidence-wiki orchestrate run \
--target . \
--runner codex \
--agent-id battery-demo
Use --runner claude to run the same work-order protocol through Claude Code.
Pass --model MODEL_ID only when an explicit runner-specific model override is
needed; otherwise the runner's safe configured default is used.
The orchestrator does not contain a battery-specific workflow. It first gives
the research agent the empty workspace; when the question blocks, it retains
that immutable run, discovers request-linked academic candidates, asks the
agent to select them, acquires the selected evidence, inventories and
normalizes it, fulfills the request, reopens the question, and starts a later
bounded research run. It declares completion only after deterministic
verification passes.
Inspect the durable parent session and export the answer:
evidence-wiki orchestrate status --target . --format json
evidence-wiki export --target . --format json
The export contains question state, the answer summary and page, confidence,
evidence strength, and citations resolved to provenance-tracked manifest and
normalized records. If no allowed provider can satisfy a request, orchestration
stops as blocked_on_sources with machine-readable remediation instead of
inventing an answer.
Local-Files-Only Alternative
Discovery and acquisition are optional. For a local-only workspace, omit all
provider flags, deliver reviewed files with provenance sidecars under the
configured raw/ roots, then run:
python3 scripts/source_inventory.py --report
python3 scripts/normalize_sources.py --all
Inventory and normalization process files that are already present. They never
search the internet or download evidence. Point an agent at
skills/research-run.md after normalization, or use the model-neutral
orchestrate start / next / submit protocol from an external harness.
Drive It With An Agent
The repository is the substrate; the intelligence comes from whatever agent
runs it. Every created workspace ships an AGENTS.md operating contract (with
a CLAUDE.md pointer for Claude-style agents) and executable skill playbooks
under skills/ — research-init, research-ingest, research-answer,
research-run, and friends. The system is developed and exercised end to end
with two harness families, and exposes a third integration surface:
- Claude Code — reads
CLAUDE.md/AGENTS.mdand the workspace skills directly. This is the pairing used for the project's live evaluation runs. - Codex-style
AGENTS.mdagents — the same contract files drive Codex-like harnesses; the project's offline regression runs use this pairing. - Any MCP client —
evidence-wiki serve-mcpexposes status, retrieval, question intake, answer export, and source-request listing over stdio without changing the canonical CLI/script contracts.
How This Project Was Built
EvidenceWiki was planned, written, and tested entirely with AI coding agents. Most of the work was done with OpenAI Codex using GPT-5.5 and GPT-5.6, with Anthropic Claude also used for parts of the project. No code in this repository was manually authored by a human.
A concrete end-to-end pairing, matching the tour above:
- Deploy with reviewed discovery and acquisition provider allow-lists.
- Seed questions yourself or let an agent submit the validated batch.
- Run
evidence-wiki orchestrate run --runner codex|claude. Each worker gets one persisted, bounded work order and the relevant workspace skill. - Inspect
evidence-wiki orchestrate statusand collectevidence-wiki export --target PATH.
Agent-Assisted Setup
For the minimal-preparation path, ask an agent to use workspace-template/skills/research-init.md. The agent should ask at most three setup questions:
- Target workspace path or short project name.
- Research scope and desired outcome.
- Optional starting sources, URLs, folders, repositories, or existing wiki.
The agent then writes a reviewable setup profile and runs the initializer with --profile. Profile-driven initialization writes docs/workspace-init-report.md in the created workspace so the inferred decisions, skipped decisions, validation commands, and next actions are visible without reading the conversation.
Example profile flow:
evidence-wiki init \
--profile /path/to/workspace-init.yml \
--dry-run
evidence-wiki init \
--profile /path/to/workspace-init.yml
Profile schema details are in workspace-template/docs/workspace-init-profile.md.
Orchestrating From a Parent Agent
A PM, planner, or parent agent can use the persisted orchestration protocol without depending on either built-in runner:
evidence-wiki orchestrate start --target PATH --agent-id parent-agent --format json
evidence-wiki orchestrate next --target PATH --orchestration-id ORCH_ID --format json
evidence-wiki orchestrate submit --target PATH --orchestration-id ORCH_ID \
--action-id ACTION_ID --result-file result.json --format json
evidence-wiki orchestrate status --target PATH --orchestration-id ORCH_ID --format json
next is idempotent: after a crash it returns the same unfinished work order.
submit treats the worker result as a claim and verifies the actual workspace
artifacts before advancing. A terminal blocked_on_sources child run is never
reopened; the parent session can retain it and start a later run after evidence
acquisition.
The research-orchestrate skill in orchestrator/skills/ documents this
protocol for external agents. The machine contract remains
workspace-template/docs/orchestrator-handoff.md.
These orchestrator skills ship with the package but are never copied into a created workspace. Locate the playbook without a source checkout:
evidence-wiki orchestrator-guide # print the resolved skill path
evidence-wiki orchestrator-guide --print # print the skill content
evidence-wiki orchestrator-guide --format json
What This Repository Contains
workspace-template/: reusable starter copied into each new research workspace.domain-packs/: optional reusable domain guidance, includingllm-research,general-science, andlegal-regulatory.examples/: public-safe worked research workspaces built from synthetic evidence.tests/: regression tests and small, rights-inventoried synthetic fixtures for initialization, inventory, normalization, linting, and end-to-end workspace bootstrap.
The template separates three evidence layers:
raw/: immutable original source material.sources/: generated manifests, normalized source records, cards, and optional codebase-analysis artifacts.wiki/: maintained research knowledge that cites source IDs.
Requirements
Required:
- Python 3.10 or newer.
- PyYAML, importable as
yaml.
Optional:
- Poppler
pdftotextfor PDF-only normalization. - Codex CLI or Claude Code for
evidence-wiki orchestrate run; the model-neutralstart/next/submit/statusprotocol does not require either built-in runner. agent-wiki-cli/llm-wikifor manually generated codebase-analysis artifacts. The repository does not install or run this adapter automatically.- Git, for normal version-control workflows and optional user-edit snapshots.
Quick dependency check:
evidence-wiki doctor --format json
From inside an initialized workspace, the same preflight is available without the package entry point:
python3 scripts/doctor.py --format json
The doctor report is machine-readable and explains degraded optional capabilities. For example, if Poppler is missing, it reports that PDF normalization degrades to stubs for PDF-only records.
Workspace Commands
Install the package from PyPI:
python3 -m pip install evidence-wiki
Create a new research workspace from explicit fields:
evidence-wiki init \
--target ../my-research-workspace \
--project-name my-research-workspace \
--project-description "Research workspace for a specific topic" \
--owner-goal "Build a source-grounded knowledge base for decisions"
Preview before writing files:
evidence-wiki init \
--target ../my-research-workspace \
--project-name my-research-workspace \
--project-description "Research workspace for a specific topic" \
--dry-run
Create a workspace with the LLM research domain pack:
evidence-wiki deploy \
--target ../llm-research-workspace \
--project-name llm-research-workspace \
--project-description "Research workspace for LLM research systems" \
--domain-pack llm-research
Refresh an existing workspace's starter-managed tooling (scripts/) to the
installed package version after upgrading evidence-wiki:
evidence-wiki upgrade --target ../my-research-workspace --dry-run # preview
evidence-wiki upgrade --target ../my-research-workspace # apply
upgrade overwrites only starter-managed files. Add --include skills or
--include docs to refresh those optional reusable directories; if an overlapping
optional file has local edits, upgrade refuses unless --force-optional is set.
Forced optional replacements preserve the displaced file under .replaced/<path>.
It never touches research.yml, raw/, sources/, wiki/, index.md, or your
log.md content.
For source-checkout development, run tests from the repository root:
python3 -B -m unittest discover -s tests -p 'test_*.py'
The equivalent source-checkout initializer command is:
python3 workspace-template/scripts/init_research_workspace.py \
--target ../my-research-workspace \
--project-name my-research-workspace \
--project-description "Research workspace for a specific topic" \
--owner-goal "Build a source-grounded knowledge base for decisions"
Worked Example
See examples/urban-heat-resilience-workspace/ for a complete public-safe
workspace. It uses synthetic urban heat resilience evidence, reserved
example.org URLs, normalized source records, source notes, claims, synthesis,
questions, decisions, and an example output without private pilot data or
machine-specific paths.
Validate A Created Workspace
Run these commands from the created workspace root:
python3 scripts/doctor.py --format json
python3 scripts/smoke_validate_workspace.py --format text
python3 scripts/source_inventory.py --report
python3 scripts/normalize_sources.py --all --dry-run
python3 scripts/lint.py --format text
Doctor should be ok or only degraded for optional capabilities before
running unattended workflows. Smoke validation should pass before inventory,
normalization, or broader lint checks. source_inventory.py --report writes
sources/manifest.jsonl so the following normalization dry run has real
manifest records to inspect. normalize_sources.py --all --dry-run reads sources/manifest.jsonl and previews generated normalized records without writing them.
For one aggregate machine-readable health and progress document (including a completion verdict for orchestrators and parent agents), run:
python3 scripts/workspace_status.py --format json
python3 scripts/workspace_status.py --check-complete --format json
From outside a workspace, the installed package exposes the same status and export contracts plus contract negotiation, fleet status, and domain-pack validation:
evidence-wiki status --target ../my-research-workspace --format json --no-cache
evidence-wiki export --target ../my-research-workspace --format json
evidence-wiki contract
evidence-wiki fleet-status --target ../my-research-workspace --format json
evidence-wiki pack validate --path general-science --format json
evidence-wiki status forwards to the copied workspace_status.py contract;
evidence-wiki export is the concise alias for evidence-wiki questions export. Both retain the workspace script schema and exit behavior.
To inject question batches into a running workspace and export structured
answers with citations (schemas in
workspace-template/docs/question-api.md), run:
python3 scripts/intake_questions.py --from-file batch.yaml --dry-run
python3 scripts/intake_questions.py --from-file batch.yaml --format json
python3 scripts/export_answers.py --format json
The same operations are available from outside the workspace via
evidence-wiki questions add|export --target PATH.
For MCP-speaking orchestrators, the optional stdio server exposes status, question status, retrieval, question intake, answer export, and source-request listing without changing the canonical CLI/script contracts:
evidence-wiki serve-mcp --target /path/to/workspace
See workspace-template/docs/mcp-server.md for the tool list, handshake,
and read/append-only boundary.
To preview inventory records before writing the manifest, run:
python3 scripts/source_inventory.py --dry-run --report
Add Sources
Enable Source Providers
Discovery and acquisition are separate permissions:
| Phase | Provider | Purpose | Runtime configuration |
|---|---|---|---|
| Discovery | arxiv |
Search arXiv metadata and propose paper candidates. | No credential. Metadata only; no download. |
| Discovery | openalex |
Search the OpenAlex scholarly index and propose paper candidates. | Optional OPENALEX_API_KEY in the environment. Index metadata is not itself evidence. |
| Discovery | github |
Search repository metadata and propose code candidates. | Optional GITHUB_TOKEN in the environment. Never clones during discovery. |
| Discovery | search |
Run a configured general-search backend. | A reviewed fixture, argv command, or HTTP backend in the setup profile; backend secrets stay in its environment. |
| Discovery | standards |
Propose bounded registry candidates. | Registry-specific configuration; it does not grant rights to standards text. |
| Acquisition | arxiv |
Download selected PDF and source-bundle evidence. | No credential; per-paper copyright and license still apply. |
| Acquisition | openalex |
Resolve works, enrich records, and download selected open-access PDFs. | Optional OPENALEX_API_KEY; unavailable or uncertain OA routes fail closed. |
| Acquisition | github |
Capture selected metadata, releases, or bounded source archives. | Optional GITHUB_TOKEN; repository license remains authoritative. |
| Acquisition | web |
Capture one explicitly selected HTTPS resource. | A non-empty domain allow-list and byte limit; license cannot be inferred. |
The three controls are independent:
integrations.discoveryauthorizes candidate metadata lookup.integrations.acquisitionauthorizes retrieval of selected evidence.- Environment credentials authenticate an already-authorized provider.
A token, installed runner, domain-pack recommendation, or discovered URL never grants provider permission. The orchestrator uses only the configured allow-lists and never edits them during a run. arXiv is primarily a preprint host and OpenAlex is a scholarly index; neither provider alone establishes that a work was peer reviewed.
Repeated initializer flags are the concise opt-in. When a flag is present for a
phase, it replaces that phase's profile allow-list and sets enabled: true:
evidence-wiki deploy \
--target literature-workspace \
--project-name literature-workspace \
--project-description "Agent-acquired literature review" \
--discovery-provider arxiv \
--discovery-provider openalex \
--acquisition-provider arxiv \
--acquisition-provider openalex \
--dry-run
Use a setup profile for providers with additional policy. For example, this
workspace_init.integrations excerpt configures general search and reviewed
web domains (the rest of the required profile is unchanged):
integrations:
git:
snapshot_user_edits: explicit
discovery:
enabled: true
providers: [search]
search:
provider: command
command: [my-search-adapter, --json]
acquisition:
enabled: true
providers: [web]
target_root: raw/papers
max_downloads_per_run: 10
require_license_check: true
web:
target_root: raw/web
allowed_domains: [official.example]
max_download_bytes: 10485760
See workspace-template/docs/source-discovery.md,
workspace-template/docs/acquisition.md, and
workspace-template/docs/workspace-init-profile.md for the complete contracts.
Deliver Local Evidence
Use research.yml to configure source roots. Common roots are:
raw/papers/for papers, PDFs, arXiv bundles, and reports.raw/links/for URLs and link lists.raw/data/for dataset cards or benchmark metadata.raw/media/for screenshots or other media evidence.raw/code/for repositories or source archives that are research evidence.
Keep raw source files immutable once added. Prefer adding a newer version as a new file instead of overwriting evidence.
Automated deliveries (fetch agents, orchestrators) follow the delivery contract in
workspace-template/docs/source-delivery.md: provenance sidecars next to delivered
files, atomic delivery, and the structured source-request artifact
(scripts/source_requests.py) that routes evidence gaps back to fetch agents.
Optional workspace-side discovery and acquisition are disabled by default.
Prompt-injection hardening guidance is in
workspace-template/docs/prompt-injection-hardening.md: source content is evidence
data, provenance URLs are metadata, and the default-on lint heuristic is a weak
reviewer-awareness signal, not a guarantee.
Inventory sources:
python3 scripts/source_inventory.py --report
Normalize pending eligible sources:
python3 scripts/normalize_sources.py
Regenerate a selected source:
python3 scripts/normalize_sources.py --source-id paper:2604.13018v1 --force
Optional Codebase Analysis
Codebase analysis is disabled by default. Enable it only when repositories, source archives, or local implementations are part of the research evidence.
Safe default contract:
integrations:
codebase_analysis:
enabled: false
provider: none
command: null
output_dir: sources/code_wikis
read_only: true
install_hooks: false
background_sync: false
untrusted_input: null
When enabled, generated adapter output must stay under sources/, not the maintained wiki/. The initializer and normalizer record adapter commands but do not execute them, clone repositories, install hooks, or start background sync. Treat raw/code/ as untrusted input and acknowledge that boundary only after choosing a safe adapter. See workspace-template/docs/codebase-analysis.md.
Human Editing And Git Safety
Agents should not auto-commit or install hooks. From a created research workspace root, take an explicit user-edit snapshot before broad operations:
python3 scripts/snapshot_user_edits.py
Only commit snapshots when the user explicitly asks:
python3 scripts/snapshot_user_edits.py --commit --message "snapshot: user edits before ingest"
Developer Workflow
Run the full regression suite from the repository root after changing scripts, fixtures, docs, or tests:
.venv/bin/python -m pytest -q
.venv/bin/python -m ruff check .
.venv/bin/python tools/sync_vendored_scripts.py --check
git diff --check
Useful focused checks:
.venv/bin/python -m pytest -q tests/test_package_cli.py
.venv/bin/python -m pytest -q tests/test_init_research_workspace.py tests/test_smoke_validate_workspace.py
.venv/bin/python -m pytest -q tests/test_inventory_normalization.py
.venv/bin/python -m pytest -q tests/test_end_to_end_init_fixture.py
On Windows, use .venv\Scripts\python.exe in place of .venv/bin/python.
Use the operating system's temporary directory for smoke workspaces so generated
files do not land in the repository root.
Documentation Map
workspace-template/docs/new-project-guide.md: full project setup guide.workspace-template/docs/workspace-initialization.md: initialization CLI and setup profile workflow.workspace-template/docs/workspace-init-profile.md: setup profile schema.workspace-template/docs/orchestrator-handoff.md: machine contract for external orchestrators and parent agents.workspace-template/docs/orchestration.md: durable parent sessions, work orders, managed runners, and recovery.workspace-template/docs/workspace-status.md: aggregate status document schema and completion verdict.workspace-template/docs/research-yml.md: public configuration contract.workspace-template/docs/source-manifest.md: source inventory format.workspace-template/docs/normalized-source-format.md: normalized record format.workspace-template/docs/coverage-manifest.md: per-question answerability manifest schema.workspace-template/docs/acquisition.md: optional acquisition safety model and provider registry.workspace-template/docs/source-discovery.md: candidate-discovery contract andsource_candidateschema (proposals before acquisition).workspace-template/docs/codebase-analysis.md: optional codebase evidence workflow.workspace-template/docs/production-readiness-checklist.md: sustained-use readiness review.
Contributing
See CONTRIBUTING.md for local setup, test expectations, style rules, and pull
request guidance.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file evidence_wiki-0.2.0.tar.gz.
File metadata
- Download URL: evidence_wiki-0.2.0.tar.gz
- Upload date:
- Size: 1.6 MB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
79d3c7aa079173441d221758dc09844292f7b606a040e1dc9454f732dfb4bee6
|
|
| MD5 |
b880d2a18cca929db88d5c8d7561379f
|
|
| BLAKE2b-256 |
4202220d88caadc991d702d1da6bf540591dbba7d643ee32a1d73c90ceaf2432
|
Provenance
The following attestation bundles were made for evidence_wiki-0.2.0.tar.gz:
Publisher:
publish.yml on Denissvgn/evidence-wiki
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
evidence_wiki-0.2.0.tar.gz -
Subject digest:
79d3c7aa079173441d221758dc09844292f7b606a040e1dc9454f732dfb4bee6 - Sigstore transparency entry: 2207120527
- Sigstore integration time:
-
Permalink:
Denissvgn/evidence-wiki@c86e175acc23a0ac01cc2785b29392de11941ebc -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/Denissvgn
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@c86e175acc23a0ac01cc2785b29392de11941ebc -
Trigger Event:
release
-
Statement type:
File details
Details for the file evidence_wiki-0.2.0-py3-none-any.whl.
File metadata
- Download URL: evidence_wiki-0.2.0-py3-none-any.whl
- Upload date:
- Size: 786.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
54c4ff825bde7dbcd3e37fbe9979469fcb9661d75fb0c71d6d8dd07ad15aed00
|
|
| MD5 |
97fb8a5e205a8690c4b32378f0a6e4b9
|
|
| BLAKE2b-256 |
b2d326355f9cc7b4a1f6e627e12a37f9bd6a8a194a445457dfdbbb6face91210
|
Provenance
The following attestation bundles were made for evidence_wiki-0.2.0-py3-none-any.whl:
Publisher:
publish.yml on Denissvgn/evidence-wiki
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
evidence_wiki-0.2.0-py3-none-any.whl -
Subject digest:
54c4ff825bde7dbcd3e37fbe9979469fcb9661d75fb0c71d6d8dd07ad15aed00 - Sigstore transparency entry: 2207120535
- Sigstore integration time:
-
Permalink:
Denissvgn/evidence-wiki@c86e175acc23a0ac01cc2785b29392de11941ebc -
Branch / Tag:
refs/tags/v0.2.0 - Owner: https://github.com/Denissvgn
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@c86e175acc23a0ac01cc2785b29392de11941ebc -
Trigger Event:
release
-
Statement type: