GENESIS
A reproducible, local-first framework for measurable multi-agent AI experiments.
GENESIS is an open-source research framework for studying whether structured collaboration between specialized AI roles can produce better problem-solving workflows under independent evaluation.
It is intentionally not marketed as AGI, consciousness, or unrestricted autonomous intelligence. The project starts with a smaller and testable question:
Can an AI workflow propose, build, test, critique, remember, and compare solutions while preserving enough evidence to explain why one strategy was accepted?
Why GENESIS exists
Many multi-agent demos show several models talking to one another and then call the conversation progress. GENESIS uses a stricter standard. A conversation is only a proposal layer. The important object is the experiment: a versioned task, strategy, artifact, execution, evaluation, critique, and decision that another person can inspect.
The project therefore treats every claimed improvement as a scientific claim. It must be compared with a frozen baseline, measured under the same conditions, evaluated independently, and preserved together with failures and resource usage.
Current status
The repository currently contains a runnable v0.1 development core:
| Capability | Current state |
|---|---|
| Typed task, strategy, budget, artifact, metric, and experiment models | Implemented |
| Strict experiment lifecycle and domain events | Implemented |
| Offline deterministic model provider | Implemented |
| SQLite experiment database | Implemented |
| Content-addressed artifact references | Implemented |
| Restricted local Python runner with timeouts | Implemented |
| Independent benchmark evaluator | Implemented |
| Researcher, Builder, and Critic roles | Implemented |
| End-to-end Coordinator | Implemented |
| CLI commands | Implemented |
| Local GitHub-style dashboard | Implemented |
| Included local Hugging Face checkpoint and provider | Implemented as an optional integration |
| Strategy evolution and dynamic roles | Deferred until the baseline is scientifically measured |
The current example is deliberately deterministic. It establishes a reproducible baseline before external model variability is introduced.
The complete workflow
The runtime sequence is:
- A
Taskdefines the problem and benchmark version. - A
Strategydefines the workflow used to solve it. - The
Researcherproduces a structured hypothesis and plan. - The
Builderproduces an artifact. - The artifact is executed in a bounded local runner.
- The independent evaluator runs the benchmark cases outside the builder's control.
- The
Criticanalyzes failures without deciding the score. - SQLite stores the experiment, events, metrics, and status history.
- GENESIS records a terminal decision such as
verified,rejected,failed, orbudget_exceeded.
A first experiment
The included example task is python-sum-squares-v1.
Input
The candidate program receives a JSON integer through standard input:
7
The task asks it to print the sum of squares from 1 through n, also as JSON:
140
The benchmark contains development, validation, and hidden cases. The initial deterministic baseline is evaluated on seven cases.
Output
A successful run produces a structured report similar to:
{
"status": "verified",
"score": 1.0,
"passed": 7,
"total": 7,
"critique": {
"failure_count": 0,
"summary": "No failures detected."
},
"artifact": "solution.py"
}
This result means that the current baseline passed the current example benchmark. It does not mean that the system became generally intelligent or that a model improved itself. The next scientific comparison is a model-backed candidate against this frozen baseline under equal budgets.
Performance engineering
GENESIS treats speed as a measured property. The evaluator now supports per-case isolated execution, accelerated batch execution in one isolated worker with a fresh namespace per case, and hash-keyed result caching. The local model provider also uses evaluation mode, inference mode, and key/value caching during generation.
On the included seven-case benchmark, the measured results were:
| Path | Wall time | Result |
|---|---|---|
| Per-case isolated evaluation | 0.327261 s | 7/7, score 1.0 |
| Accelerated batch evaluation | 0.047891 s | 7/7, score 1.0 |
| In-memory repeated evaluation | 0.000051 s | Cache hit, score 1.0 |
That run measured a 6.83× batch improvement and approximately 939× for an in-memory repeated evaluation. These are workload-specific measurements, not universal promises. End-to-end CLI time also includes interpreter startup, SQLite writes, artifact hashing, and output rendering.
Reproduce the measurements with:
python tools/benchmark_speed.py
python tools/benchmark_cli_cache.py
The isolated path remains available for hostile or untrusted artifacts; performance optimizations never bypass evaluation or the safety boundary.
See the full example in docs/example-result.md.
Quick start
GENESIS is designed to work offline for its first run.
git clone https://github.com/osamahdesk/GENESIS.git
cd GENESIS
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
Initialize the local workspace:
genesis init
Run the complete offline baseline experiment:
genesis run
Inspect stored experiments:
genesis status
Inspect a specific experiment:
genesis experiment EXP-000001
The local dashboard
The CLI is the primary interface for automation and reproducibility. The dashboard is a lightweight visual layer for humans who want a simpler overview.
Run an experiment first, then start the dashboard:
genesis run
genesis dashboard
Open:
http://127.0.0.1:8765
The dashboard shows the latest status, score, test count, workflow, experiment identity, evidence snapshot, and a short hint explaining the next action. Highlighted cards and workflow steps are clickable: a small teaching bubble explains what the selected element means and how it connects to the experiment. A Permission Center makes the deny-by-default boundary visible. It uses the same local SQLite records as the CLI and does not require an external service.
The Permission Center also exposes reversible ON/OFF controls for Transformers, PyTorch, Hugging Face, AI Builder, and Model Training. A toggle changes the experiment policy; it does not silently install packages, execute generated code, or grant host access. When AI Builder is enabled, GENESIS can create a reviewable model-project proposal inside the sandbox. The proposal must be evaluated before any future execution or promotion.
Design principles
Evidence over conversation
An agent's explanation is a hypothesis. A measured evaluation is evidence. The evaluator is independent from the artifact-generation role.
Reproducibility by default
Experiments preserve task identity, strategy identity, artifact hashes, status events, model metadata, configuration, metrics, and timestamps. Future releases will export a complete experiment bundle.
Safe defaults
The default mode is offline. Network access is not silently enabled. Generated code runs in a bounded temporary workspace with a timeout and a sanitized environment. This reduces risk but is not a perfect security boundary; users must not execute untrusted code with permissions they would not grant to an unknown program.
Provider independence
The core domain models do not depend on OpenAI, Anthropic, Hugging Face, or any single vendor. Provider adapters are deliberately separated from tasks, experiments, evaluation, storage, and orchestration.
Small increments
GENESIS is developed one daily work file at a time. Each file defines its scope, acceptance criteria, tests, blockers, and final status. Later features are not silently pulled into an earlier day.
Repository structure
GENESIS/
├── README.md
├── pyproject.toml
├── configs/
│ └── default.yaml
├── docs/
│ ├── vision.md
│ ├── v0.1-spec.md
│ ├── architecture.md
│ ├── evaluation.md
│ ├── security.md
│ ├── roadmap.md
│ ├── work-tree.md
│ ├── example-result.md
│ ├── images/
│ └── daily/
├── genesis/
│ ├── agents/
│ ├── core/
│ ├── environment/
│ ├── evaluation/
│ ├── providers/
│ ├── storage/
│ └── ui/
├── benchmarks/
├── examples/
└── tests/
Architecture overview
The Coordinator is the boundary between user intent and experiment execution. It invokes specialized roles, applies the lifecycle state machine, records domain events, and delegates scoring to the independent evaluator.
The safety boundary is intentionally outside the parts that future evolution may modify. A later strategy or role candidate may be tested in an isolated experiment, but it cannot silently replace the evaluator, hidden-test policy, approval mechanism, or host permissions.
Evaluation policy
A fair comparison uses the same task set, model, model parameters, retry rules, time budget, execution policy, and output format for the baseline and candidate. A single run is not enough to establish a scientific improvement. Future reports will support repeated trials, spread, cost, latency, hidden-task performance, and ablation studies.
The intended baseline ladder is:
Single model → Single model + self-critique
→ Researcher + Builder
→ Researcher + Builder + Critic + Repair
The purpose is to measure which component contributes to an observed change rather than assuming that more agents automatically means better performance.
Roadmap
The approved roadmap is phased:
- Foundation: schemas, events, configuration, providers, and storage.
- First experiment: evaluator, runner, roles, Coordinator, and CLI.
- Scientific integrity: hidden splits, repetitions, cost metrics, anti-cheating checks, and reproducibility bundles.
- Strategy evolution: frozen parents, mutation, selection, failure memory, and regression gates.
- Future research: dynamic roles, tool experiments, capability-gap analysis, and carefully controlled external research.
The detailed sequence is in docs/roadmap.md, and each day is tracked in docs/daily/day-index.md.
Tests
Install development dependencies and run:
python -m pytest
The current development core includes unit, lifecycle, evaluation, end-to-end, and dashboard tests. A clean test run is a release requirement, not an optional polish step.
Hugging Face integration
Phase 2 includes the small sshleifer/tiny-gpt2 checkpoint from Hugging Face. The files are committed with a SHA256SUMS manifest, and HuggingFaceLocalProvider loads them from disk through the same provider boundary used by the mock provider.
The integration is optional because PyTorch and Transformers are heavyweight dependencies compared with the deterministic core. Install it only when you want to run the local model:
pip install -e '.[hf]'
The default test suite and offline baseline still work without the hf extra. This separation lets GENESIS use a real model while preserving clean, fast, reproducible development for the rest of the project. The model is intentionally a tiny integration checkpoint, not a claim of production-quality generation.
Teacher model onboarding
Users do not need to know the Hub workflow in advance. The dashboard provides a Teacher Model center that recommends catalog entries using detected machine memory, accepts an exact owner/model identifier, accepts a Hugging Face URL, or accepts an existing local model directory. The selection is saved as teacher.json for later experiments.
When the Hugging Face capability and read-only network permission are enabled, the dashboard can query the public Hub API on explicit request. Results include model ID, downloads, and likes, with a Use action for selecting a teacher. The teacher produces reviewable candidates inside the sandbox; it does not silently clone itself, replace GENESIS, install arbitrary code, or bypass evaluation.
Contributing
Start by reading docs/v0.1-spec.md, then choose the next incomplete daily work file. Keep changes small, add tests, preserve the offline path, and document any architecture change. Do not publish claims of improvement without the benchmark evidence that supports them.
License
MIT. See LICENSE.
Metadata
Release files for genesis-agi 0.1.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| genesis_agi-0.1.3.tar.gz | 35.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| genesis_agi-0.1.3-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 70.4 kB
Release files / genesis_agi-0.1.3.tar.gz
| Download URL | genesis_agi-0.1.3.tar.gz |
|---|---|
| Size | 35.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cc49ed536318e4e118667bee671021a59cafc88b024a5fcc0f9cbc2f1854c6df
|
|
BLAKE2b-256 checksum How to use checksums |
a57399b2de46be03a451be39433fa5cce4a5151d89e0e5a8292a9d53394e5735
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency logRelease files / genesis_agi-0.1.3-py3-none-any.whl
| Download URL | genesis_agi-0.1.3-py3-none-any.whl |
|---|---|
| Size | 34.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f4990588e8550625e0657c60f3ffa13cc81e4bfcaf88af262b2266eacd9d9b0a
|
|
BLAKE2b-256 checksum How to use checksums |
550bb536f68091bca24710a04a288330bc30b1025c383186a73a7a66d3049384
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.
Transparency log