pyfly-lightning
Run experiments on connectome-derived spiking substrates without becoming a connectomics engineer. Point it at a task; it builds the substrate slice, finds a routing for your input, trains the readout, and reports which nulls the result beat.
pip install pyfly-lightning
pyfly-lightning plan --n-samples 6000 --n-features 784 --n-classes 10
pyfly-lightning fit --npy X.npy y.npy --out report.json
Backed by MaleCNS v1.0 (adult male Drosophila, 165,122 traced neurons, 25.6M directed edges) and FlyWire 783, both fetched over public HTTPS with sha256-pinned integrity checks. Any other connectome plugs in as an edge list.
Install
pip install pyfly-lightning # core: numpy, pandas, pyarrow, scipy
pip install 'pyfly-lightning[train]' # + torch, lightning
pip install 'pyfly-lightning[sim]' # + numba
pip install 'pyfly-lightning[es,vizier]' # + cma, google-vizier
pip install 'pyfly-lightning[train,mnist]' # + torchvision, for the examples
# from source, with the test suite
git clone https://github.com/NewJerseyStyle/pyfly-lightning && cd pyfly-lightning
pip install -e '.[train,dev]'
Requires Python 3.10+. Tested on 3.10, 3.11 and 3.12. No API token is needed for either dataset.
Import name
import pyfly_lightning as pfl
The distribution is pyfly-lightning; the import is pyfly_lightning, following the
pytorch-lightning / pytorch_lightning convention. The bare name pyfly is taken on
PyPI by an unrelated load-testing framework, and two distributions that install the same
top-level directory get merged in site-packages: one __init__.py wins and removing
either one deletes the other's files. Verified: both install side by side, and
uninstalling each in turn leaves the other intact.
The command-line entry point is pyfly-lightning (or python -m pyfly_lightning).
Get data
pyfly-lightning describe # datasets, sizes, licences, caveats
pyfly-lightning get malecns10 # ~505 MB required, sha256-verified
pyfly-lightning get flywire783 # annotations only
pyfly-lightning verify malecns10 # re-check the cache
Cache defaults to ~/.cache/pyfly; override with PYFLY_CACHE. Nothing is vendored
into the package.
Quick start
From the shell
pyfly-lightning plan --n-samples 150 --n-features 4 --n-classes 3 # what would it choose, and why
pyfly-lightning fit --csv iris.csv --target species --out report.json
pyfly-lightning fit --npy X.npy y.npy --modulation
plan prints the reason behind every setting it picked:
2048 nodes, pools=['CX'], latent 8, 9 readout features, ES 5 x 16 = 80 evaluations
"n_outputs": "9 = n_samples/16, because 64 features on 120 samples lost to 4 raw features"
"budget": "CX is forced in because without it late_echo = 0.0 and activity dies in 2 steps"
From Python
import pyfly_lightning as pfl # distribution `pyfly-lightning`, as with
# `pytorch-lightning` -> `pytorch_lightning`
result = pfl.fit(X, y, modulation=True)
print(result) # <FitResult acc=0.6087 chance=0.1227 verdict=UNSUPPORTED>
print(result.controls) # fly / rewired / shuffled_weight / pooled_features
print(result.plan["reason"])
Driving the stages yourself
from pyfly_lightning.train.pipeline import PipelineConfig, StagedPipeline
model = StagedPipeline(substrate, in_dim, n_classes, dan_pos=dan_pos,
config=PipelineConfig(T=4, use_dan=True), device="cuda")
model.stage1_es(Xtr, ytr) # black-box search over the routing
model.stage2_decode(Xtr, ytr, Xva, yva) # backprop trains the readout
model.stage3_modulate(Xtr, ytr, Xva, yva) # the neuromodulatory readback
What you can plug in
Datasets
| name | organism | scope | neurons | edges |
|---|---|---|---|---|
malecns10 |
adult male | central brain + optic lobes + VNC | 165,122 traced | 25,563,197 (significant-only) |
flywire783 |
adult female | brain only | 139,255 | ~5e7 chemical synapses |
fly = pyfly.load("malecns10")
io = fly.io(inputs=["cb_sensory"], outputs=["vnc_motor"]) # 15,896 in / 2,295 out
sub = fly.slice(budget=8192, include_pools=("CX",), seed=0)
Edge counts are easy to confuse: the figure often quoted as "300M+ synapses" counts
synaptic contacts, not directed neuron-pair edges. describe states which quantity
each dataset reports.
Your own connectome
Any (pre, post, weight) edge list, any node labels, any column names:
from pyfly_lightning.data.graph import from_edgelist
from pyfly_lightning.data.schema import NeuronTable
from pyfly_lightning.model.substrate import Substrate
conn = from_edgelist(edges_df) # columns pre / post / weight
table = NeuronTable.from_dataframe(ann_df, id_col="node_name",
superclass_col="kind", class_col="group",
input_labels=("source",),
output_labels=("sink",))
sub = Substrate.build(conn, table, budget=2048, seed=0)
String labels, uuids and integer ids all work. The annotation table is canonicalised once at the boundary, so downstream code does not care what your columns are called.
For different purposes
Connectomics / circuit neuroscience. The census tooling answers structural questions directly: which pools are self-contained, which are broadcast relays, how far a population reaches, what a lesion would cost.
python tools/census_wm_ports.py # in/out fractions, reciprocity, mixing, leak
sub.internal_fraction(cx_ids) # how self-contained is this pool
sub.recurrence_report() # reciprocal fraction, mixing eigenvalue, self-loops
echo_probe(sub.W, sources) # does activity survive the input stopping
broadcast_coverage(sub.W, dan_pos) # what fraction of the slice a pool can reach
pyfly_lightning.gates.structure_function_gate runs the classic structure→function test:
stimulate a named population, find which outputs the real wiring drives, then measure
activation probability under real versus shuffled weights.
Spiking / neuromorphic engineering. The LIF simulator ships two backends that are bit-identical, which makes correctness testable rather than assumed:
from pyfly_lightning.sim.lif import LifNumpy, LifTorch # same dynamics, numpy vs torch
LifTorch(W, device="cuda") # sparse event-driven step
LifTorch(W, differentiable=True) # smooth gate: gradients flow
differentiable=True keeps the autograd graph so the whole pipeline trains end to
end. set_trainable_weights() keeps the connectome's wiring while making every edge
weight a parameter. Throughput and memory are measured rather than estimated:
python -m pyfly_lightning.sim.bench --device cuda reports steps/sec and the ceiling (a
16k-node, 528k-edge substrate uses 50 MB of a 6 GB card; latency is ~0.15 ms/step).
Machine learning / architecture search. The substrate is a sparse recurrent architecture with a natural I/O boundary and structure-derived grouping.
from pyfly_lightning.model.routing import build_sensory_surface, GroupRouter
surface = build_sensory_surface(table, level="medium") # ~10 modality groups
router = GroupRouter(surface, latent_dim=32, learnable=True)
router.param_report() # parameter count, compression vs dense, optimiser fit
level="fine" gives 1,269 structural groups (optic-lobe hex columns, olfactory
glomeruli, labelled lines); level="medium" collapses to ~10 modalities, which brings
full-covariance CMA-ES into range. ESTrainer and OuterLoop cover the black-box
side; pyfly_lightning.train.outer uses Vizier when installed and records which backend ran.
Reinforcement learning. DQNAgent and PPOAgent train heads on substrate
features, with an environment that provides genuinely multi-step trajectories, plus an
inference-safe neuromodulatory channel:
from pyfly_lightning.envs import CueDelayChoice
from pyfly_lightning.train.rl import DQNAgent, PPOAgent, RLConfig
agent = DQNAgent(model, CueDelayChoice(delay=8, act_every_step=True), RLConfig())
Model compression. MaskPruner searches a group-level mask under a tolerance
contract and rollback; EvoPruner runs NSGA-II over group genomes and returns a
Pareto front over (behaviour, active nodes, active edges).
model.freeze(adapter=True) # required precondition
from pyfly_lightning.prune import MaskPruner, PruneConfig
rep = MaskPruner(model, PruneConfig(tolerance=0.05)).prune(Xtr, ytr, Xva, yva)
Dynamical systems. pyfly_lightning.sim.probe distinguishes propagation from recirculation,
which is the distinction that matters when the object of study is a recurrent graph
rather than a classifier.
The staged pipeline
user data ──► encoder ──► router ──► [ frozen / trainable substrate ] ──► readout ──► task
(ES) (ES) │ (backprop)
└─► neuromodulatory channel ──┘
| stage | what it does | optimised by |
|---|---|---|
stage1_es |
finds which substrate neurons your input reaches | ES, scored on a held-out split |
stage2_decode |
trains the readout over the output population | Adam |
stage3_modulate |
trains a recurrent readback through the DAN/PPL1 pool | Adam |
The modulation channel carries the substrate's own pass-1 output back onto the dopaminergic pool and runs a second pass, which is available at inference. Feeding the label or the error into the channel would make the reported accuracy unreproducible on unlabelled data.
Controls and reporting
Every fit reports the same set of nulls, and a verdict:
| control | what it isolates |
|---|---|
rewired |
degree-preserving rewiring: does the specific wiring matter |
shuffled_weight |
the same wiring with weights permuted |
pooled_features |
ridge on the raw input features, same readout class |
no_plasticity |
weight updates disabled |
FlyModule exposes the same contract, and pyfly_lightning.train.controls.required_controls_report
assembles it. FitResult.verdict is SUPPORTED when the fit beats every null and
UNSUPPORTED otherwise, with the scores recorded either way.
What has been measured so far
Recorded so you can design around it; the full protocols, ablations and bugs are in
docs/architecture.md.
| experiment | result |
|---|---|
| MNIST, substrate vs its own rewiring | 0.154 vs 0.154 — identical |
| MNIST, staged pipeline vs direct input + PPL1 (3 seeds) | +7.6 to +15.3 points |
| Neuromodulatory channel, frozen / trainable substrate | +11.6 / +25.2 points |
| Iris, low-data, per-fold ES (3 seeds x 5 folds) | fly 0.7644 vs rewired 0.7711 |
| Iris, substrate vs raw 4 features | 0.7644 vs 0.8289 |
| Trainable substrate: real vs shuffled weights | 0.6860 vs 0.7360 |
| Degree-preserving rewiring of the CX pool | mixing eigenvalue changes 0.12% |
| CX forced into a slice vs not | late_echo 30.0 vs 0.0 |
| Mask pruning under a tolerance contract | edges −67.6%, held-out score unchanged |
| RTX 2060 6 GB, 16,384 nodes / 528,013 edges | 0.148 ms/step, 50 MB, 99.2% headroom |
A note on naming that the fields make unavoidable: what these runs call fly is the
substrate wired from a real connectome, and rewired is the same graph with its edges
re-randomised under a degree-preserving swap. The comparison is built into the report
because a result without it cannot be placed.
Package layout
pyfly_lightning/data/ registry, sha256-pinned downloader, cache, NeuronTable,
SparseConnectome, sign derivation with an audit trail
pyfly_lightning/model/ structured routing (hex columns / glomeruli / labelled lines),
substrate slicing, the four nulls
pyfly_lightning/sim/ LIF (numpy + torch, bit-identical), probes, benchmark
pyfly_lightning/train/ FlyModule (a real lightning.LightningModule), ES, outer loop,
DQN, PPO, the staged pipeline, environments
pyfly_lightning/reward.py objective, bounded reward calibration, credit channel, plasticity
pyfly_lightning/prune.py tolerance-contract mask pruning
pyfly_lightning/prune_evo.py NSGA-II over group genomes, Pareto front
pyfly_lightning/gates.py structure→function gate with a shuffled-weight control
pyfly_lightning/auto.py automatic configuration with recorded reasons
pyfly_lightning/api.py load / io / slice / port / gate / reward_bus / fit
Docs
docs/quickstart.md— install, plan, fit, bring your own connectomedocs/architecture.md— design, the full measurement record, freeze semantics, the pruning contract, and a milestone-by-milestone checklistdocs/census-wm-ports.md— which neuron pools are self-contained, measured
Examples
python examples/quickstart.py # both modes, one API
python examples/mnist.py # Lightning, four controls
python examples/emotion_pipeline.py # the six-arm comparison
python examples/delayed_match.py # a memory task
python examples/lowdata.py # low-data regime, CV
python examples/freeze_prune.py --method mask # prune under tolerance
python examples/freeze_prune.py --method evo --grouping depth # NSGA-II Pareto front
python examples/reward_bus.py # calibration + plasticity
python examples/rl_dqn.py # RL, with controls
Tests
pytest -q # 104 tests; data-dependent ones skip if the cache is empty
pyfly-lightning get malecns10 && pytest -q
Citation
CITATION.cff is included. Cite the datasets alongside this software:
- Berg et al., Sexual dimorphism in the complete connectome of the Drosophila male central nervous system, Cell 2026 — MaleCNS v1.0
- Dorkenwald et al., Neuronal wiring diagram of an adult brain, Nature 2024 — FlyWire
- Schlegel et al., Whole-brain annotation and multi-connectome cell typing of Drosophila, Nature 2024 — cell types and the I/O classification
- Shiu et al., A Drosophila computational brain model reveals sensorimotor processing, Nature 2024 — the LIF parameterisation
- Eckstein et al., Cell 2024 — neurotransmitter predictions used for synapse sign
Check each dataset's own licence and citation terms; MaleCNS is CC-BY, and FlyWire data has its own guidelines.
Licence
MIT. Linked datasets keep their own licences.
Contributing
Issues and pull requests are welcome, particularly new task domains, additional
connectomes, and simulation backends. docs/architecture.md records the measurements
behind the current defaults, so a change to a default should come with the measurement
that motivates it.
Metadata
Release files for pyfly-lightning 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| pyfly_lightning-0.1.0.tar.gz | 177.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| pyfly_lightning-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 272.8 kB
Release files / pyfly_lightning-0.1.0.tar.gz
| Download URL | pyfly_lightning-0.1.0.tar.gz |
|---|---|
| Size | 177.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
418ab21d1ea0f19edae4b95bd67a203a28837d5361279fdfaf86e499aad23913
|
|
BLAKE2b-256 checksum How to use checksums |
f43533f605737993a7dbe0ee1babb1e5b7b41c0a98135547ca87e08a9704923d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.
Transparency logRelease files / pyfly_lightning-0.1.0-py3-none-any.whl
| Download URL | pyfly_lightning-0.1.0-py3-none-any.whl |
|---|---|
| Size | 95.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
047f3464cc2e6fa7ce7de86f4b1b3f20e6cdfda9019ac7fd86b0e486a454901d
|
|
BLAKE2b-256 checksum How to use checksums |
8881f2af53ac6b6bdfb496e81f0e5be6a6e72cfba2508bf42ec84ded2733d6c1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.
Transparency log