EvidenceWiki
Answers you can audit. EvidenceWiki creates persistent research workspaces where agents investigate questions and deterministic scripts enforce provenance, lifecycle state, and export validation. A validated research outcome is either a cited, auditable answer or a structured request for missing evidence.
For updates to an installed domain pack, use evidence-wiki pack revision-plan,
revision-apply and revision-status. They retain three-way conflicts and
history, identify affected research, and guide explicit coverage reevaluation.
See Pack revisions.
Quick start · Documentation · Worked example · PyPI · Contributing
Start with your agent
Install into an isolated environment (Python 3.10 or newer):
python3 -m venv .research-env
.research-env/bin/python -m pip install evidence-wiki
.research-env/bin/evidence-wiki agent --format json
.research-env/bin/evidence-wiki agent summary --format json \
--require strict-evidence/v1 --require declarative-computation/v1
On Windows, use .research-env\Scripts\python.exe and
.research-env\Scripts\evidence-wiki.exe. The response identifies the installed
version, supported contracts and instructions. An older installation can refuse
an unavailable capability; use its returned version and instructions together.
Native Windows supports local workspace captures, protected publication, pack
catalogs, authoring and setup through anchored handles and private ACLs. Reparse
points, network shares and filesystems without the required native guarantees
are refused. Use a host-owned directory and trust file restricted to the current
Windows account and SYSTEM for protected evidence state. Framework runtime
qualification and host_enforced isolation remain separate capabilities.
Give your current agent the installed executable, your original questions, the
available source locations, writable directory and access limits. Ask it to read
agent first, preserve every question, and return the controlled research export
or precise blockers. A second model runner is optional. Retrieved documents and
tool output supply evidence; they cannot authorize commands or change the review policy.
Instructions travel with the package and can be read from any directory:
.research-env/bin/evidence-wiki pack guide --format text
.research-env/bin/evidence-wiki pack guide --topic authoring --format text
.research-env/bin/evidence-wiki pack guide --topic revisions --format text
.research-env/bin/evidence-wiki agent source-guide --format text
.research-env/bin/evidence-wiki agent research-guide --format text
.research-env/bin/evidence-wiki agent frameworks
Select a pack by its scope and evidence requirements. If none fits, preserve the gap or explicitly author local guidance, assess it and register the selected revision. Host-delivered captures retain their declared origin, rights, scope and completeness. Local copies record observation time; this is not a publication date or a license grant. Pack changes invalidate affected acceptance and require explicit reevaluation. Installing a newer package does not migrate existing workspaces; see Upgrade and adoption.
artifact_checked applies to the freshly checked export, with current evidence
and authenticated independent reviews. It cannot control prose written outside
that export. host_enforced additionally requires the qualified macOS protected
host to mediate execution and delivery; ordinary terminal, bridge and managed
parent sessions do not acquire that guarantee. Quote matches and exact decimal
calculations do not prove source truth, semantic support, suitable units or
complete research. Contested and insufficient evidence remain explicit outcomes.
The installed framework matrix separates pinned transport/resource observations
from model behavior. Pi, OpenCode and Gemini have observed modes on macOS arm64;
live-model research and other platform/version combinations need their own
qualification. Read the exact modes with agent frameworks before selecting one.
How This Project Was Built
EvidenceWiki was planned, written, and tested entirely with AI coding agents. Most of the work was done with OpenAI Codex using GPT-5.5 and GPT-5.6, with Anthropic Claude also used for parts of the project. No code in this repository was manually authored by a human.
Why EvidenceWiki
- Traceable answers. Every citation resolves through a stable source ID to a normalized record and its provenance-tracked original.
- Evidence-aware failure. Configured coverage requirements block weakly supported answers and produce machine-readable source requests.
- Deterministic control. Scripts own critical question-lifecycle transitions, validation, and export; agents supply research judgment.
- Reusable workspaces. The starter, domain packs, agent skills, and orchestration protocol work across research domains and agent harnesses.
question → discover/acquire → inventory/normalize → answer/verify → export
↘ missing evidence → structured source request
The workspace keeps original evidence in raw/, generated evidence records in
sources/, and maintained research knowledge in wiki/. Source content is
treated as data, never as agent instructions; see prompt-injection
hardening.
Five-Minute Tour
These commands set up the workflow; research time varies with the question,
providers, and agent runner. The examples use a POSIX-compatible shell; on
Windows, create batch.yaml in an editor or adapt that one heredoc step for
PowerShell. Python 3.10 or newer is required.
Install the package and create a provider-enabled scientific workspace:
python3 -m pip install evidence-wiki
evidence-wiki deploy \
--target solid-state-batteries \
--project-name solid-state-batteries \
--project-description "Survey of solid-state battery electrolyte research" \
--domain-pack general-science \
--discovery-provider arxiv \
--discovery-provider openalex \
--acquisition-provider arxiv \
--acquisition-provider openalex
cd solid-state-batteries
init and deploy invoke the same workspace initializer. The repeated
provider flags explicitly authorize network-backed discovery and acquisition;
a domain pack never enables providers by itself. arXiv needs no credential,
while OpenAlex can use OPENALEX_API_KEY from the process environment. See
workspace initialization, source
discovery, and acquisition for the full
contracts.
Check credentials before the first command that needs external services:
evidence-wiki env check --service openalex --service github --format json
env check reports variable names and presence only. env run --prompt asks
for missing values with terminal echo disabled and passes them to the selected
command and its children. Values are never saved by this helper or added to your
parent shell. Inject persistent credentials through your host secret manager.
Select openai, anthropic, github or openalex, or use --require-env NAME
for another integration. Select only the credentials your workflow needs;
initialization and offline work require none, and agent runners may use their
own login. See environment setup.
Inspect guidance before selecting it with evidence-wiki pack list and
evidence-wiki pack show bundled:general-science. Packs expose scope inputs,
exclusions and review requirements. Use evidence-wiki pack guide --format text
for explicit local catalogs and requirement-based selection; see
pack selection.
Use evidence-wiki agent inspect --target WORKSPACE to distinguish installed
capabilities from configured access and source usability. Select source IDs with
agent source-status; agent source-guide --format text explains host captures,
route choices and remediation. Inspection does not activate providers or run
normalization. See source usability.
agent recipes describes bounded native DOCX text/table capture. Use
agent extensions for explicit pack identity migration, composition, fleet
proposals and host transitions. Python hosts can use Onboarding.open with
selected roots and operation grants; serve-onboarding-mcp exposes the same
scoped owners. See onboarding contracts
and optional native skill installation.
Use evidence-wiki agent plan --from-file request.json to compile a read-only
setup plan with original question mappings, evidence criteria, source routes and
explicit assurance blockers. Save with --output and recheck input identities
with agent plan-check. See research planning.
Use evidence-wiki pack guide --topic authoring to scaffold or derive local
guidance, freeze assessment cases, qualify a revision and resume setup planning.
Apply a saved plan with evidence-wiki agent apply --from-file PLAN.
Workspace application describes
local delivery, observed readiness, locking and conservative recovery.
For an existing workspace, evidence-wiki agent --target WORKSPACE returns current
next-action advice. Use agent research-guide for the caller-driven loop,
verified source ingestion, original-question export and optional local progress.
Caller research uses the current
agent and canonical owners; no additional model runner is required.
Structural and arithmetic checks retain separate domain-review requirements.
Add a question using the question API:
cat > batch.yaml <<'EOF'
schema_version: "1.0"
questions:
- question: "Which solid electrolyte families report room-temperature ionic conductivity above 1 mS/cm?"
id: electrolyte-conductivity
priority: high
EOF
evidence-wiki questions add --target . --from-file batch.yaml
Codex CLI 0.138 or newer must already be installed for the managed Codex adapter. Check the environment before launching it:
evidence-wiki doctor --format json
Run the managed orchestrator:
evidence-wiki orchestrate run \
--target . \
--runner codex \
--agent-id battery-demo
Use --runner claude for the managed Claude Code adapter. Then inspect the
durable parent session and export the answer:
evidence-wiki orchestrate status --target . --format json
evidence-wiki export --target . --format json
The orchestrator can discover candidate sources, ask an agent to select them,
acquire and normalize the selected evidence, reopen a blocked question, and
verify the final artifacts. If allowed providers cannot satisfy the request,
the session ends as blocked_on_sources instead of inventing an answer. The
orchestration guide covers execution, recovery, and security
boundaries.
Local-files-only alternative
Discovery and acquisition are optional. Omit provider flags, deliver reviewed
files with provenance sidecars under the configured raw/ roots, then run:
python3 scripts/source_inventory.py --report
python3 scripts/normalize_sources.py --all
Inventory and normalization process only files already present. Continue with the research-run skill, or use the external protocol described below. The source-delivery contract defines provenance sidecars and atomic delivery.
Drive It With An Agent
Start with your current agent and terminal, before creating a workspace:
evidence-wiki agent
evidence-wiki agent summary --format json
evidence-wiki agent resource guide/pack-authoring/v1
evidence-wiki agent frameworks
Bootstrap is read-only and includes the installed operating guide, versioned resource references and a strict-policy template. It requires no secondary model CLI. Resource access does not apply policy or establish host enforcement; the guide explains configuration, independent review and explicit evidence gaps.
Portable skills and an optional Pi native tool/RPC bridge reuse these contracts.
Inspect exact framework versions and qualified modes before use, then create a
caller-local bundle with evidence-wiki agent bundle --target NEW_DIRECTORY.
No global agent settings or trust choices are changed. See the installed
guide/frameworks/v1 resource for Pi, OpenCode and Gemini CLI recipes and limits.
EvidenceWiki supports agent harnesses at three levels:
- Managed adapters: Codex and Claude Code are the registered runners for
package-owned
runandresumeexecution. - External protocol: OpenCode, Pi, Aider, Gemini CLI, and other harnesses
can drive
start,next,submit, andstatusfrom an operator-controlled host. They are not package-managed runners. - Instruction compatibility: any worker can follow the workspace
AGENTS.md, selected skill, and bounded work order.CLAUDE.mdpoints Claude-style agents to the same contract.
Managed Codex execution requires Codex CLI 0.138 or newer. Managed Claude
execution is unavailable on native Windows; use macOS, Linux, WSL2, a
container, or the external protocol. If the required isolation boundary cannot
be enforced, the host returns RUNNER_ISOLATION_UNAVAILABLE before starting a
worker. The parent exclusively owns runs/orchestrations/; workers never write
that tree or invoke the parent controller. Use resume for a retained session
after a runner failure. See parent orchestration for
isolation, leases, tamper recovery, and upgrade rules.
External protocol
A PM, planner, or custom host can drive the model-neutral protocol directly:
evidence-wiki orchestrate start --target PATH --agent-id parent-agent --format json
evidence-wiki orchestrate next --target PATH --orchestration-id ORCH_ID --format json
evidence-wiki orchestrate submit --target PATH --orchestration-id ORCH_ID \
--action-id ACTION_ID --result-file result.json --format json
evidence-wiki orchestrate status --target PATH --orchestration-id ORCH_ID --format json
next is idempotent, and submit verifies workspace postconditions before
advancing. External hosts must provide process isolation, single-driver
coordination, and crash replay; see the orchestrator handoff
contract. The packaged research-orchestrate
playbook lives under orchestrator/skills/ and can be
located without a source checkout:
evidence-wiki orchestrator-guide
evidence-wiki orchestrator-guide --print
For MCP clients, an optional stdio server exposes status, retrieval, question intake, answer export, and source-request listing:
evidence-wiki serve-mcp --target /path/to/workspace
See the MCP server contract for its tool list and read/append-only boundary.
Drive It From Python
A host that embeds EvidenceWiki — an ASGI service, a scheduler, a batch worker — can call the package in-process instead of spawning the CLI per operation:
from evidence_wiki import Workspace
with Workspace.open("/path/to/workspace") as ws:
report = ws.coverage.evaluate("electrolyte-conductivity")
Twenty-six operations return the same documents the matching --format json
commands print, and refuse with typed exceptions carrying the same stable error
codes. Both doors render from one seam per operation, so they cannot disagree.
Orchestration keeps a subprocess to the workspace's own deployed controller,
which is version-matched to the session state it owns. The package ships no HTTP
server; hosts build their own. See the library API for the full
surface, the error families, thread-safety guarantees, and a worked embedding
example.
Requirements and Diagnostics
Required:
- Python 3.10 or newer.
- PyYAML 6.0 or newer, ruamel.yaml 0.19.1 or newer within the 0.19 series, and
pypdf 6.14 or newer within major version 6. All are installed with
evidence-wiki; ruamel.yaml preserves live YAML comments and quoting during pack refresh, while the portable pypdf backend requires no separate PDF tool.
Optional capabilities include Codex CLI or Claude Code for managed runs, Git
for snapshots, and the Poppler compatibility backend for explicitly configured
pdftotext extraction. Platform installation is covered by workspace
initialization; managed-runner sandbox requirements
are covered by parent orchestration.
Check dependencies and optional capabilities from any directory:
evidence-wiki doctor --format json
An initialized workspace includes the same preflight:
python3 scripts/doctor.py --format json
Missing pypdf is a required failure. Missing Poppler is informational unless the workspace explicitly selects the Poppler compatibility backend.
Create and Maintain a Workspace
Create a generic workspace from explicit fields:
evidence-wiki init \
--target ../my-research-workspace \
--project-name my-research-workspace \
--project-description "Research workspace for a specific topic" \
--owner-goal "Build a source-grounded knowledge base for decisions"
Add --dry-run to preview without writing files. For minimal-preparation,
agent-assisted setup, ask an agent to follow the research-init
skill; it can prepare a reviewable workspace init
profile.
After upgrading the package, preview and apply starter-managed script updates:
evidence-wiki upgrade --target ../my-research-workspace --dry-run
evidence-wiki upgrade --target ../my-research-workspace
Write-mode upgrade refreshes only starter-managed tooling, may update
workspace-system.yml, uses .locks/, and conditionally appends one audit
entry to log.md when it applies material changes. It preserves prior log
history, research.yml, raw/, sources/, wiki/, index.md, and other user
data. --dry-run writes nothing. Both modes refuse with UPGRADE_PENDING_ORDER
while an orchestration session holds a pending work order or an active driver,
naming the session and order: drain orchestration before upgrading. Optional
skills and docs have additional conflict rules documented in
workspace initialization.
Domain packs have a separate, explicit lifecycle. Preview and apply a new revision of the already-installed pack with:
evidence-wiki pack refresh \
--target ../my-research-workspace \
--path general-science \
--dry-run
evidence-wiki pack refresh \
--target ../my-research-workspace \
--path general-science
An older workspace whose pack predates lifecycle state must first run
evidence-wiki pack adopt --target ../my-research-workspace --dry-run, review
the result, and repeat without --dry-run. Refresh never switches pack names,
and an unresolved local/pack conflict produces zero writes. See domain
packs for adoption, path-specific conflict resolution, and
transaction recovery.
Validate A Created Workspace
For manual or operator-level validation, the copied workspace exposes its lower-level checks directly. Run these commands from the workspace root:
python3 scripts/doctor.py --format json
python3 scripts/smoke_validate_workspace.py --format text
python3 scripts/source_inventory.py --report
python3 scripts/normalize_sources.py --all --dry-run
python3 scripts/normalize_verify.py --format text
python3 scripts/lint.py --format text
source_inventory.py --report writes sources/manifest.jsonl, so
normalize_sources.py --all --dry-run reads sources/manifest.jsonl and can
preview normalized records without writing them. For aggregate health and a
machine-readable completion verdict, run:
python3 scripts/workspace_status.py --format json
python3 scripts/workspace_status.py --check-complete --format json
Question intake and structured answer export are also available inside a workspace:
python3 scripts/intake_questions.py --from-file batch.yaml --dry-run
python3 scripts/intake_questions.py --from-file batch.yaml --format json
python3 scripts/export_answers.py --format json
The installed equivalents are evidence-wiki status, evidence-wiki questions add, and evidence-wiki export; see workspace status
and the question API.
To preview inventory records without writing the manifest:
python3 scripts/source_inventory.py --dry-run --report
Evidence and Provider Permissions
Discovery and acquisition are separate permissions. Discovery providers
(arxiv, openalex, github, search, and standards) propose metadata;
candidates are not evidence until selected, acquired into raw/, and recorded
with provenance. Acquisition providers (arxiv, openalex, github, and
allow-listed web) retrieve selected evidence under configured limits.
Three controls remain independent:
integrations.discoveryauthorizes candidate lookup.integrations.acquisitionauthorizes retrieval.- Environment credentials authenticate an already-authorized provider.
A token, installed runner, domain-pack recommendation, or discovered URL never grants provider permission. See source discovery, acquisition, and the workspace init profile for provider configuration. For reviewed local evidence, follow the source-delivery contract, keep raw files immutable, then inventory and normalize them.
Evidence is not limited to the source kinds this package extracts. Normalized
records are a versioned public contract, so an external
normalizer can supply records for evidence the package does not read itself —
structured API payloads, instrument output — and those records count on exactly
the same terms as records the package wrote. The terms are enforced, not assumed:
evidence-wiki normalize verify checks a record against the contract and names
each breach with a stable code, and lint accepts an externally written record only
when it conforms.
Repository Layout
workspace-template/is copied into each research workspace and contains its scripts, skills, and operator documentation.domain-packs/contains optional, reusable domain guidance.examples/includes a complete public-safe workspace built from synthetic evidence.orchestrator/contains the external parent-agent playbook.tests/contains regression tests and synthetic fixtures with documented usage rights.
Documentation
- Start a workspace: new project guide, workspace
initialization, setup profile
schema,
research.ymlconfiguration, domain packs, and the worked example. - Research and evidence: question API, source discovery, acquisition, source delivery, source manifest, normalized records, coverage manifests, evidence policies, and citation verification.
- Agents and integrations: parent orchestration, orchestrator handoff, workspace status, run controller, library API, MCP server, and orchestrator playbooks.
- Safety and operations: prompt-injection hardening, human editing and snapshots, codebase analysis, production readiness, and publication readiness.
- Project development: contributing, changelog, publishing workflow, third-party notices, and license. Architecture notes live in the workspace documents above, starting with workspace initialization and orchestration.
Development setup, repository boundaries, style rules, and the full verification suite are documented in CONTRIBUTING.md.
License
EvidenceWiki is available under the MIT License.
Release files for evidence-wiki 1.0.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| evidence_wiki-1.0.0.tar.gz | 4.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| evidence_wiki-1.0.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 6.4 MB
Release files / evidence_wiki-1.0.0.tar.gz
| Download URL | evidence_wiki-1.0.0.tar.gz |
|---|---|
| Size | 4.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
898ccf803ddeadf71b93b4db48b366c98060986f33c23ab897d8ea033d640f87
|
|
BLAKE2b-256 checksum How to use checksums |
1744b6599d752c987da33f097d63fcad96fe9ae7281c80c562c3a2eb50cdbabc
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency logRelease files / evidence_wiki-1.0.0-py3-none-any.whl
| Download URL | evidence_wiki-1.0.0-py3-none-any.whl |
|---|---|
| Size | 2.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
009637abc258759615b12150c6530a6cb06180be5db7383b659c853295df3e1a
|
|
BLAKE2b-256 checksum How to use checksums |
d55a521ae238c229d8103c685f30ff5c149ffe1a4c2f3f3bfc6417f54a0f0aa2
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.13
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 25, 2026.
Transparency log