Skip to main content

GENESIS

A reproducible, local-first framework for measurable multi-agent AI experiments.

GENESIS is an open-source research framework for studying whether structured collaboration between specialized AI roles can produce better problem-solving workflows under independent evaluation.

It is intentionally not marketed as AGI, consciousness, or unrestricted autonomous intelligence. The project starts with a smaller and testable question:

Can an AI workflow propose, build, test, critique, remember, and compare solutions while preserving enough evidence to explain why one strategy was accepted?

GENESIS architecture

Why GENESIS exists

Many multi-agent demos show several models talking to one another and then call the conversation progress. GENESIS uses a stricter standard. A conversation is only a proposal layer. The important object is the experiment: a versioned task, strategy, artifact, execution, evaluation, critique, and decision that another person can inspect.

The project therefore treats every claimed improvement as a scientific claim. It must be compared with a frozen baseline, measured under the same conditions, evaluated independently, and preserved together with failures and resource usage.

Current status

The repository currently contains a runnable v0.1 development core:

Capability Current state
Typed task, strategy, budget, artifact, metric, and experiment models Implemented
Strict experiment lifecycle and domain events Implemented
Offline deterministic model provider Implemented
SQLite experiment database Implemented
Content-addressed artifact references Implemented
Restricted local Python runner with timeouts Implemented
Independent benchmark evaluator Implemented
Researcher, Builder, and Critic roles Implemented
End-to-end Coordinator Implemented
CLI commands Implemented
Local GitHub-style dashboard Implemented
Included local Hugging Face checkpoint and provider Implemented as an optional integration
Strategy evolution and dynamic roles Deferred until the baseline is scientifically measured

The current example is deliberately deterministic. It establishes a reproducible baseline before external model variability is introduced.

The complete workflow

GENESIS workflow

The runtime sequence is:

  1. A Task defines the problem and benchmark version.
  2. A Strategy defines the workflow used to solve it.
  3. The Researcher produces a structured hypothesis and plan.
  4. The Builder produces an artifact.
  5. The artifact is executed in a bounded local runner.
  6. The independent evaluator runs the benchmark cases outside the builder's control.
  7. The Critic analyzes failures without deciding the score.
  8. SQLite stores the experiment, events, metrics, and status history.
  9. GENESIS records a terminal decision such as verified, rejected, failed, or budget_exceeded.

Project structure

A first experiment

The included example task is python-sum-squares-v1.

Input

The candidate program receives a JSON integer through standard input:

7

The task asks it to print the sum of squares from 1 through n, also as JSON:

140

The benchmark contains development, validation, and hidden cases. The initial deterministic baseline is evaluated on seven cases.

Output

A successful run produces a structured report similar to:

{
  "status": "verified",
  "score": 1.0,
  "passed": 7,
  "total": 7,
  "critique": {
    "failure_count": 0,
    "summary": "No failures detected."
  },
  "artifact": "solution.py"
}

First experiment result

This result means that the current baseline passed the current example benchmark. It does not mean that the system became generally intelligent or that a model improved itself. The next scientific comparison is a model-backed candidate against this frozen baseline under equal budgets.

Performance engineering

GENESIS treats speed as a measured property. The evaluator now supports per-case isolated execution, accelerated batch execution in one isolated worker with a fresh namespace per case, and hash-keyed result caching. The local model provider also uses evaluation mode, inference mode, and key/value caching during generation.

On the included seven-case benchmark, the measured results were:

Path Wall time Result
Per-case isolated evaluation 0.327261 s 7/7, score 1.0
Accelerated batch evaluation 0.047891 s 7/7, score 1.0
In-memory repeated evaluation 0.000051 s Cache hit, score 1.0

That run measured a 6.83× batch improvement and approximately 939× for an in-memory repeated evaluation. These are workload-specific measurements, not universal promises. End-to-end CLI time also includes interpreter startup, SQLite writes, artifact hashing, and output rendering.

Reproduce the measurements with:

python tools/benchmark_speed.py
python tools/benchmark_cli_cache.py

The isolated path remains available for hostile or untrusted artifacts; performance optimizations never bypass evaluation or the safety boundary.

See the full example in docs/example-result.md.

Quick start

GENESIS is designed to work offline for its first run.

git clone https://github.com/osamahdesk/GENESIS.git
cd GENESIS
python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

Initialize the local workspace:

genesis init

Run the complete offline baseline experiment:

genesis run

Inspect stored experiments:

genesis status

Inspect a specific experiment:

genesis experiment EXP-000001

The local dashboard

The CLI is the primary interface for automation and reproducibility. The dashboard is a lightweight visual layer for humans who want a simpler overview.

Run an experiment first, then start the dashboard:

genesis run
genesis dashboard

Open:

http://127.0.0.1:8765

The dashboard shows the latest status, score, test count, workflow, experiment identity, evidence snapshot, and a short hint explaining the next action. Highlighted cards and workflow steps are clickable: a small teaching bubble explains what the selected element means and how it connects to the experiment. A Permission Center makes the deny-by-default boundary visible. It uses the same local SQLite records as the CLI and does not require an external service.

The Permission Center also exposes reversible ON/OFF controls for Transformers, PyTorch, Hugging Face, AI Builder, and Model Training. A toggle changes the experiment policy; it does not silently install packages, execute generated code, or grant host access. When AI Builder is enabled, GENESIS can create a reviewable model-project proposal inside the sandbox. The proposal must be evaluated before any future execution or promotion.

Design principles

Evidence over conversation

An agent's explanation is a hypothesis. A measured evaluation is evidence. The evaluator is independent from the artifact-generation role.

Reproducibility by default

Experiments preserve task identity, strategy identity, artifact hashes, status events, model metadata, configuration, metrics, and timestamps. Future releases will export a complete experiment bundle.

Safe defaults

The default mode is offline. Network access is not silently enabled. Generated code runs in a bounded temporary workspace with a timeout and a sanitized environment. This reduces risk but is not a perfect security boundary; users must not execute untrusted code with permissions they would not grant to an unknown program.

Provider independence

The core domain models do not depend on OpenAI, Anthropic, Hugging Face, or any single vendor. Provider adapters are deliberately separated from tasks, experiments, evaluation, storage, and orchestration.

Small increments

GENESIS is developed one daily work file at a time. Each file defines its scope, acceptance criteria, tests, blockers, and final status. Later features are not silently pulled into an earlier day.

Repository structure

GENESIS/
├── README.md
├── pyproject.toml
├── configs/
│   └── default.yaml
├── docs/
│   ├── vision.md
│   ├── v0.1-spec.md
│   ├── architecture.md
│   ├── evaluation.md
│   ├── security.md
│   ├── roadmap.md
│   ├── work-tree.md
│   ├── example-result.md
│   ├── images/
│   └── daily/
├── genesis/
│   ├── agents/
│   ├── core/
│   ├── environment/
│   ├── evaluation/
│   ├── providers/
│   ├── storage/
│   └── ui/
├── benchmarks/
├── examples/
└── tests/

Architecture overview

The Coordinator is the boundary between user intent and experiment execution. It invokes specialized roles, applies the lifecycle state machine, records domain events, and delegates scoring to the independent evaluator.

The safety boundary is intentionally outside the parts that future evolution may modify. A later strategy or role candidate may be tested in an isolated experiment, but it cannot silently replace the evaluator, hidden-test policy, approval mechanism, or host permissions.

Evaluation policy

A fair comparison uses the same task set, model, model parameters, retry rules, time budget, execution policy, and output format for the baseline and candidate. A single run is not enough to establish a scientific improvement. Future reports will support repeated trials, spread, cost, latency, hidden-task performance, and ablation studies.

The intended baseline ladder is:

Single model → Single model + self-critique
             → Researcher + Builder
             → Researcher + Builder + Critic + Repair

The purpose is to measure which component contributes to an observed change rather than assuming that more agents automatically means better performance.

Roadmap

The approved roadmap is phased:

  • Foundation: schemas, events, configuration, providers, and storage.
  • First experiment: evaluator, runner, roles, Coordinator, and CLI.
  • Scientific integrity: hidden splits, repetitions, cost metrics, anti-cheating checks, and reproducibility bundles.
  • Strategy evolution: frozen parents, mutation, selection, failure memory, and regression gates.
  • Future research: dynamic roles, tool experiments, capability-gap analysis, and carefully controlled external research.

The detailed sequence is in docs/roadmap.md, and each day is tracked in docs/daily/day-index.md.

Tests

Install development dependencies and run:

python -m pytest

The current development core includes unit, lifecycle, evaluation, end-to-end, and dashboard tests. A clean test run is a release requirement, not an optional polish step.

Hugging Face integration

Phase 2 includes the small sshleifer/tiny-gpt2 checkpoint from Hugging Face. The files are committed with a SHA256SUMS manifest, and HuggingFaceLocalProvider loads them from disk through the same provider boundary used by the mock provider.

The integration is optional because PyTorch and Transformers are heavyweight dependencies compared with the deterministic core. Install it only when you want to run the local model:

pip install -e '.[hf]'

The default test suite and offline baseline still work without the hf extra. This separation lets GENESIS use a real model while preserving clean, fast, reproducible development for the rest of the project. The model is intentionally a tiny integration checkpoint, not a claim of production-quality generation.

Teacher model onboarding

Users do not need to know the Hub workflow in advance. The dashboard provides a Teacher Model center that recommends catalog entries using detected machine memory, accepts an exact owner/model identifier, accepts a Hugging Face URL, or accepts an existing local model directory. The selection is saved as teacher.json for later experiments.

When the Hugging Face capability and read-only network permission are enabled, the dashboard can query the public Hub API on explicit request. Results include model ID, downloads, and likes, with a Use action for selecting a teacher. The teacher produces reviewable candidates inside the sandbox; it does not silently clone itself, replace GENESIS, install arbitrary code, or bypass evaluation.

Contributing

Start by reading docs/v0.1-spec.md, then choose the next incomplete daily work file. Keep changes small, add tests, preserve the offline path, and document any architecture change. Do not publish claims of improvement without the benchmark evidence that supports them.

License

MIT. See LICENSE.

Metadata

Release files for genesis-agi 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for genesis-agi 0.1.2
File Size Uploaded
genesis_agi-0.1.2.tar.gz 35.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for genesis-agi 0.1.2
File Interpreter ABI Platform
genesis_agi-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 70.0 kB

Release files / genesis_agi-0.1.2.tar.gz

Download URL genesis_agi-0.1.2.tar.gz
Size 35.4 kB
Tags Source
SHA-256 checksum
How to use checksums
f8bd3f8373de5bc8b54551c35739e4819f3b876e00dab5abf43c5905d4323df2
BLAKE2b-256 checksum
How to use checksums
f199a85268fd027ad4cc957fd3aa4dd82194a39def719969e6c69c1d465e8345
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release files / genesis_agi-0.1.2-py3-none-any.whl

Download URL genesis_agi-0.1.2-py3-none-any.whl
Size 34.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
826ee825411bbbc0919c8d6ee250fe2c7d6d12e3df7e3916aab38f11efd97cde
BLAKE2b-256 checksum
How to use checksums
f94d55bd1ea1002bcff69d65402c4c58eb015946a26810248dc3045e5aa1ee00
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 21, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

0.1.3

2 release files

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page