Skip to main content

EvidenceWiki

Answers you can audit. EvidenceWiki creates persistent research workspaces where agents investigate questions and deterministic scripts enforce provenance, lifecycle state, and export validation. A validated research outcome is either a cited, auditable answer or a structured request for missing evidence.

For updates to an installed domain pack, use evidence-wiki pack revision-plan, revision-apply and revision-status. They retain three-way conflicts and history, identify affected research, and guide explicit coverage reevaluation. See Pack revisions.

Quick start · Documentation · Worked example · PyPI · Contributing

Start with your agent

Install into an isolated environment (Python 3.10 or newer):

python3 -m venv .research-env
.research-env/bin/python -m pip install evidence-wiki
.research-env/bin/evidence-wiki agent --format json
.research-env/bin/evidence-wiki agent summary --format json \
  --require strict-evidence/v1 --require declarative-computation/v1

On Windows, use .research-env\Scripts\python.exe and .research-env\Scripts\evidence-wiki.exe. The response identifies the installed version, supported contracts and instructions. An older installation can refuse an unavailable capability; use its returned version and instructions together.

Native Windows supports local workspace captures, protected publication, pack catalogs, authoring and setup through anchored handles and private ACLs. Reparse points, network shares and filesystems without the required native guarantees are refused. Use a host-owned directory and trust file restricted to the current Windows account and SYSTEM for protected evidence state. Framework runtime qualification and host_enforced isolation remain separate capabilities.

Give your current agent the installed executable, your original questions, the available source locations, writable directory and access limits. Ask it to read agent first, preserve every question, and return the controlled research export or precise blockers. A second model runner is optional. Retrieved documents and tool output supply evidence; they cannot authorize commands or change the review policy.

Instructions travel with the package and can be read from any directory:

.research-env/bin/evidence-wiki pack guide --format text
.research-env/bin/evidence-wiki pack guide --topic authoring --format text
.research-env/bin/evidence-wiki pack guide --topic revisions --format text
.research-env/bin/evidence-wiki agent source-guide --format text
.research-env/bin/evidence-wiki agent research-guide --format text
.research-env/bin/evidence-wiki agent frameworks

Select a pack by its scope and evidence requirements. If none fits, preserve the gap or explicitly author local guidance, assess it and register the selected revision. Host-delivered captures retain their declared origin, rights, scope and completeness. Local copies record observation time; this is not a publication date or a license grant. Pack changes invalidate affected acceptance and require explicit reevaluation. Installing a newer package does not migrate existing workspaces; see Upgrade and adoption.

artifact_checked applies to the freshly checked export, with current evidence and authenticated independent reviews. It cannot control prose written outside that export. host_enforced additionally requires the qualified macOS protected host to mediate execution and delivery; ordinary terminal, bridge and managed parent sessions do not acquire that guarantee. Quote matches and exact decimal calculations do not prove source truth, semantic support, suitable units or complete research. Contested and insufficient evidence remain explicit outcomes.

The installed framework matrix separates pinned transport/resource observations from model behavior. Pi, OpenCode and Gemini have observed modes on macOS arm64; live-model research and other platform/version combinations need their own qualification. Read the exact modes with agent frameworks before selecting one.

How This Project Was Built

EvidenceWiki was planned, written, and tested entirely with AI coding agents. Most of the work was done with OpenAI Codex using GPT-5.5 and GPT-5.6, with Anthropic Claude also used for parts of the project. No code in this repository was manually authored by a human.

Why EvidenceWiki

  • Traceable answers. Every citation resolves through a stable source ID to a normalized record and its provenance-tracked original.
  • Evidence-aware failure. Configured coverage requirements block weakly supported answers and produce machine-readable source requests.
  • Deterministic control. Scripts own critical question-lifecycle transitions, validation, and export; agents supply research judgment.
  • Reusable workspaces. The starter, domain packs, agent skills, and orchestration protocol work across research domains and agent harnesses.
question → discover/acquire → inventory/normalize → answer/verify → export
                       ↘ missing evidence → structured source request

The workspace keeps original evidence in raw/, generated evidence records in sources/, and maintained research knowledge in wiki/. Source content is treated as data, never as agent instructions; see prompt-injection hardening.

Five-Minute Tour

These commands set up the workflow; research time varies with the question, providers, and agent runner. The examples use a POSIX-compatible shell; on Windows, create batch.yaml in an editor or adapt that one heredoc step for PowerShell. Python 3.10 or newer is required.

Install the package and create a provider-enabled scientific workspace:

python3 -m pip install evidence-wiki
evidence-wiki deploy \
  --target solid-state-batteries \
  --project-name solid-state-batteries \
  --project-description "Survey of solid-state battery electrolyte research" \
  --domain-pack general-science \
  --discovery-provider arxiv \
  --discovery-provider openalex \
  --acquisition-provider arxiv \
  --acquisition-provider openalex
cd solid-state-batteries

init and deploy invoke the same workspace initializer. The repeated provider flags explicitly authorize network-backed discovery and acquisition; a domain pack never enables providers by itself. arXiv needs no credential, while OpenAlex can use OPENALEX_API_KEY from the process environment. See workspace initialization, source discovery, and acquisition for the full contracts.

Check credentials before the first command that needs external services:

evidence-wiki env check --service openalex --service github --format json

env check reports variable names and presence only. env run --prompt asks for missing values with terminal echo disabled and passes them to the selected command and its children. Values are never saved by this helper or added to your parent shell. Inject persistent credentials through your host secret manager. Select openai, anthropic, github or openalex, or use --require-env NAME for another integration. Select only the credentials your workflow needs; initialization and offline work require none, and agent runners may use their own login. See environment setup.

Inspect guidance before selecting it with evidence-wiki pack list and evidence-wiki pack show bundled:general-science. Packs expose scope inputs, exclusions and review requirements. Use evidence-wiki pack guide --format text for explicit local catalogs and requirement-based selection; see pack selection.

Use evidence-wiki agent inspect --target WORKSPACE to distinguish installed capabilities from configured access and source usability. Select source IDs with agent source-status; agent source-guide --format text explains host captures, route choices and remediation. Inspection does not activate providers or run normalization. See source usability.

agent recipes describes bounded native DOCX text/table capture. Use agent extensions for explicit pack identity migration, composition, fleet proposals and host transitions. Python hosts can use Onboarding.open with selected roots and operation grants; serve-onboarding-mcp exposes the same scoped owners. See onboarding contracts and optional native skill installation.

Use evidence-wiki agent plan --from-file request.json to compile a read-only setup plan with original question mappings, evidence criteria, source routes and explicit assurance blockers. Save with --output and recheck input identities with agent plan-check. See research planning.

Use evidence-wiki pack guide --topic authoring to scaffold or derive local guidance, freeze assessment cases, qualify a revision and resume setup planning. Apply a saved plan with evidence-wiki agent apply --from-file PLAN. Workspace application describes local delivery, observed readiness, locking and conservative recovery.

For an existing workspace, evidence-wiki agent --target WORKSPACE returns current next-action advice. Use agent research-guide for the caller-driven loop, verified source ingestion, original-question export and optional local progress. Caller research uses the current agent and canonical owners; no additional model runner is required. Structural and arithmetic checks retain separate domain-review requirements.

Add a question using the question API:

cat > batch.yaml <<'EOF'
schema_version: "1.0"
questions:
  - question: "Which solid electrolyte families report room-temperature ionic conductivity above 1 mS/cm?"
    id: electrolyte-conductivity
    priority: high
EOF
evidence-wiki questions add --target . --from-file batch.yaml

Codex CLI 0.138 or newer must already be installed for the managed Codex adapter. Check the environment before launching it:

evidence-wiki doctor --format json

Run the managed orchestrator:

evidence-wiki orchestrate run \
  --target . \
  --runner codex \
  --agent-id battery-demo

Use --runner claude for the managed Claude Code adapter. Then inspect the durable parent session and export the answer:

evidence-wiki orchestrate status --target . --format json
evidence-wiki export --target . --format json

The orchestrator can discover candidate sources, ask an agent to select them, acquire and normalize the selected evidence, reopen a blocked question, and verify the final artifacts. If allowed providers cannot satisfy the request, the session ends as blocked_on_sources instead of inventing an answer. The orchestration guide covers execution, recovery, and security boundaries.

Local-files-only alternative

Discovery and acquisition are optional. Omit provider flags, deliver reviewed files with provenance sidecars under the configured raw/ roots, then run:

python3 scripts/source_inventory.py --report
python3 scripts/normalize_sources.py --all

Inventory and normalization process only files already present. Continue with the research-run skill, or use the external protocol described below. The source-delivery contract defines provenance sidecars and atomic delivery.

Drive It With An Agent

Start with your current agent and terminal, before creating a workspace:

evidence-wiki agent
evidence-wiki agent summary --format json
evidence-wiki agent resource guide/pack-authoring/v1
evidence-wiki agent frameworks

Bootstrap is read-only and includes the installed operating guide, versioned resource references and a strict-policy template. It requires no secondary model CLI. Resource access does not apply policy or establish host enforcement; the guide explains configuration, independent review and explicit evidence gaps.

Portable skills and an optional Pi native tool/RPC bridge reuse these contracts. Inspect exact framework versions and qualified modes before use, then create a caller-local bundle with evidence-wiki agent bundle --target NEW_DIRECTORY. No global agent settings or trust choices are changed. See the installed guide/frameworks/v1 resource for Pi, OpenCode and Gemini CLI recipes and limits.

EvidenceWiki supports agent harnesses at three levels:

  • Managed adapters: Codex and Claude Code are the registered runners for package-owned run and resume execution.
  • External protocol: OpenCode, Pi, Aider, Gemini CLI, and other harnesses can drive start, next, submit, and status from an operator-controlled host. They are not package-managed runners.
  • Instruction compatibility: any worker can follow the workspace AGENTS.md, selected skill, and bounded work order. CLAUDE.md points Claude-style agents to the same contract.

Managed Codex execution requires Codex CLI 0.138 or newer. Managed Claude execution is unavailable on native Windows; use macOS, Linux, WSL2, a container, or the external protocol. If the required isolation boundary cannot be enforced, the host returns RUNNER_ISOLATION_UNAVAILABLE before starting a worker. The parent exclusively owns runs/orchestrations/; workers never write that tree or invoke the parent controller. Use resume for a retained session after a runner failure. See parent orchestration for isolation, leases, tamper recovery, and upgrade rules.

External protocol

A PM, planner, or custom host can drive the model-neutral protocol directly:

evidence-wiki orchestrate start --target PATH --agent-id parent-agent --format json
evidence-wiki orchestrate next --target PATH --orchestration-id ORCH_ID --format json
evidence-wiki orchestrate submit --target PATH --orchestration-id ORCH_ID \
  --action-id ACTION_ID --result-file result.json --format json
evidence-wiki orchestrate status --target PATH --orchestration-id ORCH_ID --format json

next is idempotent, and submit verifies workspace postconditions before advancing. External hosts must provide process isolation, single-driver coordination, and crash replay; see the orchestrator handoff contract. The packaged research-orchestrate playbook lives under orchestrator/skills/ and can be located without a source checkout:

evidence-wiki orchestrator-guide
evidence-wiki orchestrator-guide --print

For MCP clients, an optional stdio server exposes status, retrieval, question intake, answer export, and source-request listing:

evidence-wiki serve-mcp --target /path/to/workspace

See the MCP server contract for its tool list and read/append-only boundary.

Drive It From Python

A host that embeds EvidenceWiki — an ASGI service, a scheduler, a batch worker — can call the package in-process instead of spawning the CLI per operation:

from evidence_wiki import Workspace

with Workspace.open("/path/to/workspace") as ws:
    report = ws.coverage.evaluate("electrolyte-conductivity")

Twenty-six operations return the same documents the matching --format json commands print, and refuse with typed exceptions carrying the same stable error codes. Both doors render from one seam per operation, so they cannot disagree. Orchestration keeps a subprocess to the workspace's own deployed controller, which is version-matched to the session state it owns. The package ships no HTTP server; hosts build their own. See the library API for the full surface, the error families, thread-safety guarantees, and a worked embedding example.

Requirements and Diagnostics

Required:

  • Python 3.10 or newer.
  • PyYAML 6.0 or newer, ruamel.yaml 0.19.1 or newer within the 0.19 series, and pypdf 6.14 or newer within major version 6. All are installed with evidence-wiki; ruamel.yaml preserves live YAML comments and quoting during pack refresh, while the portable pypdf backend requires no separate PDF tool.

Optional capabilities include Codex CLI or Claude Code for managed runs, Git for snapshots, and the Poppler compatibility backend for explicitly configured pdftotext extraction. Platform installation is covered by workspace initialization; managed-runner sandbox requirements are covered by parent orchestration.

Check dependencies and optional capabilities from any directory:

evidence-wiki doctor --format json

An initialized workspace includes the same preflight:

python3 scripts/doctor.py --format json

Missing pypdf is a required failure. Missing Poppler is informational unless the workspace explicitly selects the Poppler compatibility backend.

Create and Maintain a Workspace

Create a generic workspace from explicit fields:

evidence-wiki init \
  --target ../my-research-workspace \
  --project-name my-research-workspace \
  --project-description "Research workspace for a specific topic" \
  --owner-goal "Build a source-grounded knowledge base for decisions"

Add --dry-run to preview without writing files. For minimal-preparation, agent-assisted setup, ask an agent to follow the research-init skill; it can prepare a reviewable workspace init profile.

After upgrading the package, preview and apply starter-managed script updates:

evidence-wiki upgrade --target ../my-research-workspace --dry-run
evidence-wiki upgrade --target ../my-research-workspace

Write-mode upgrade refreshes only starter-managed tooling, may update workspace-system.yml, uses .locks/, and conditionally appends one audit entry to log.md when it applies material changes. It preserves prior log history, research.yml, raw/, sources/, wiki/, index.md, and other user data. --dry-run writes nothing. Both modes refuse with UPGRADE_PENDING_ORDER while an orchestration session holds a pending work order or an active driver, naming the session and order: drain orchestration before upgrading. Optional skills and docs have additional conflict rules documented in workspace initialization.

Domain packs have a separate, explicit lifecycle. Preview and apply a new revision of the already-installed pack with:

evidence-wiki pack refresh \
  --target ../my-research-workspace \
  --path general-science \
  --dry-run
evidence-wiki pack refresh \
  --target ../my-research-workspace \
  --path general-science

An older workspace whose pack predates lifecycle state must first run evidence-wiki pack adopt --target ../my-research-workspace --dry-run, review the result, and repeat without --dry-run. Refresh never switches pack names, and an unresolved local/pack conflict produces zero writes. See domain packs for adoption, path-specific conflict resolution, and transaction recovery.

Validate A Created Workspace

For manual or operator-level validation, the copied workspace exposes its lower-level checks directly. Run these commands from the workspace root:

python3 scripts/doctor.py --format json
python3 scripts/smoke_validate_workspace.py --format text
python3 scripts/source_inventory.py --report
python3 scripts/normalize_sources.py --all --dry-run
python3 scripts/normalize_verify.py --format text
python3 scripts/lint.py --format text

source_inventory.py --report writes sources/manifest.jsonl, so normalize_sources.py --all --dry-run reads sources/manifest.jsonl and can preview normalized records without writing them. For aggregate health and a machine-readable completion verdict, run:

python3 scripts/workspace_status.py --format json
python3 scripts/workspace_status.py --check-complete --format json

Question intake and structured answer export are also available inside a workspace:

python3 scripts/intake_questions.py --from-file batch.yaml --dry-run
python3 scripts/intake_questions.py --from-file batch.yaml --format json
python3 scripts/export_answers.py --format json

The installed equivalents are evidence-wiki status, evidence-wiki questions add, and evidence-wiki export; see workspace status and the question API.

To preview inventory records without writing the manifest:

python3 scripts/source_inventory.py --dry-run --report

Evidence and Provider Permissions

Discovery and acquisition are separate permissions. Discovery providers (arxiv, openalex, github, search, and standards) propose metadata; candidates are not evidence until selected, acquired into raw/, and recorded with provenance. Acquisition providers (arxiv, openalex, github, and allow-listed web) retrieve selected evidence under configured limits.

Three controls remain independent:

  1. integrations.discovery authorizes candidate lookup.
  2. integrations.acquisition authorizes retrieval.
  3. Environment credentials authenticate an already-authorized provider.

A token, installed runner, domain-pack recommendation, or discovered URL never grants provider permission. See source discovery, acquisition, and the workspace init profile for provider configuration. For reviewed local evidence, follow the source-delivery contract, keep raw files immutable, then inventory and normalize them.

Evidence is not limited to the source kinds this package extracts. Normalized records are a versioned public contract, so an external normalizer can supply records for evidence the package does not read itself — structured API payloads, instrument output — and those records count on exactly the same terms as records the package wrote. The terms are enforced, not assumed: evidence-wiki normalize verify checks a record against the contract and names each breach with a stable code, and lint accepts an externally written record only when it conforms.

Repository Layout

Documentation

Development setup, repository boundaries, style rules, and the full verification suite are documented in CONTRIBUTING.md.

License

EvidenceWiki is available under the MIT License.

Release files for evidence-wiki 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for evidence-wiki 1.0.0
File Size Uploaded
evidence_wiki-1.0.0.tar.gz 4.4 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for evidence-wiki 1.0.0
File Interpreter ABI Platform
evidence_wiki-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 6.4 MB

Release files / evidence_wiki-1.0.0.tar.gz

Download URL evidence_wiki-1.0.0.tar.gz
Size 4.4 MB
Tags Source
SHA-256 checksum
How to use checksums
898ccf803ddeadf71b93b4db48b366c98060986f33c23ab897d8ea033d640f87
BLAKE2b-256 checksum
How to use checksums
1744b6599d752c987da33f097d63fcad96fe9ae7281c80c562c3a2eb50cdbabc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.

Transparency log

Release files / evidence_wiki-1.0.0-py3-none-any.whl

Download URL evidence_wiki-1.0.0-py3-none-any.whl
Size 2.1 MB
Tags Python 3
SHA-256 checksum
How to use checksums
009637abc258759615b12150c6530a6cb06180be5db7383b659c853295df3e1a
BLAKE2b-256 checksum
How to use checksums
d55a521ae238c229d8103c685f30ff5c149ffe1a4c2f3f3bfc6417f54a0f0aa2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.13

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page