Skip to main content

pyfly-lightning

Run experiments on connectome-derived spiking substrates without becoming a connectomics engineer. Point it at a task; it builds the substrate slice, finds a routing for your input, trains the readout, and reports which nulls the result beat.

pip install pyfly-lightning
pyfly-lightning plan --n-samples 6000 --n-features 784 --n-classes 10
pyfly-lightning fit  --npy X.npy y.npy --out report.json

Backed by MaleCNS v1.0 (adult male Drosophila, 165,122 traced neurons, 25.6M directed edges) and FlyWire 783, both fetched over public HTTPS with sha256-pinned integrity checks. Any other connectome plugs in as an edge list.


Install

pip install pyfly-lightning                  # core: numpy, pandas, pyarrow, scipy
pip install 'pyfly-lightning[train]'         # + torch, lightning
pip install 'pyfly-lightning[sim]'           # + numba
pip install 'pyfly-lightning[es,vizier]'     # + cma, google-vizier
pip install 'pyfly-lightning[train,mnist]'   # + torchvision, for the examples

# from source, with the test suite
git clone https://github.com/NewJerseyStyle/pyfly-lightning && cd pyfly-lightning
pip install -e '.[train,dev]'

Requires Python 3.10+. Tested on 3.10, 3.11 and 3.12. No API token is needed for either dataset.

Import name

import pyfly_lightning as pfl

The distribution is pyfly-lightning; the import is pyfly_lightning, following the pytorch-lightning / pytorch_lightning convention. The bare name pyfly is taken on PyPI by an unrelated load-testing framework, and two distributions that install the same top-level directory get merged in site-packages: one __init__.py wins and removing either one deletes the other's files. Verified: both install side by side, and uninstalling each in turn leaves the other intact.

The command-line entry point is pyfly-lightning (or python -m pyfly_lightning).

Get data

pyfly-lightning describe                 # datasets, sizes, licences, caveats
pyfly-lightning get malecns10            # ~505 MB required, sha256-verified
pyfly-lightning get flywire783           # annotations only
pyfly-lightning verify malecns10         # re-check the cache

Cache defaults to ~/.cache/pyfly; override with PYFLY_CACHE. Nothing is vendored into the package.


Quick start

From the shell

pyfly-lightning plan --n-samples 150 --n-features 4 --n-classes 3     # what would it choose, and why
pyfly-lightning fit  --csv iris.csv --target species --out report.json
pyfly-lightning fit  --npy X.npy y.npy --modulation

plan prints the reason behind every setting it picked:

2048 nodes, pools=['CX'], latent 8, 9 readout features, ES 5 x 16 = 80 evaluations
  "n_outputs": "9 = n_samples/16, because 64 features on 120 samples lost to 4 raw features"
  "budget":    "CX is forced in because without it late_echo = 0.0 and activity dies in 2 steps"

From Python

import pyfly_lightning as pfl          # distribution `pyfly-lightning`, as with
                                       # `pytorch-lightning` -> `pytorch_lightning`

result = pfl.fit(X, y, modulation=True)
print(result)            # <FitResult acc=0.6087 chance=0.1227 verdict=UNSUPPORTED>
print(result.controls)   # fly / rewired / shuffled_weight / pooled_features
print(result.plan["reason"])

Driving the stages yourself

from pyfly_lightning.train.pipeline import PipelineConfig, StagedPipeline

model = StagedPipeline(substrate, in_dim, n_classes, dan_pos=dan_pos,
                       config=PipelineConfig(T=4, use_dan=True), device="cuda")
model.stage1_es(Xtr, ytr)                  # black-box search over the routing
model.stage2_decode(Xtr, ytr, Xva, yva)    # backprop trains the readout
model.stage3_modulate(Xtr, ytr, Xva, yva)  # the neuromodulatory readback

What you can plug in

Datasets

name organism scope neurons edges
malecns10 adult male central brain + optic lobes + VNC 165,122 traced 25,563,197 (significant-only)
flywire783 adult female brain only 139,255 ~5e7 chemical synapses
fly  = pyfly.load("malecns10")
io   = fly.io(inputs=["cb_sensory"], outputs=["vnc_motor"])   # 15,896 in / 2,295 out
sub  = fly.slice(budget=8192, include_pools=("CX",), seed=0)

Edge counts are easy to confuse: the figure often quoted as "300M+ synapses" counts synaptic contacts, not directed neuron-pair edges. describe states which quantity each dataset reports.

Your own connectome

Any (pre, post, weight) edge list, any node labels, any column names:

from pyfly_lightning.data.graph import from_edgelist
from pyfly_lightning.data.schema import NeuronTable
from pyfly_lightning.model.substrate import Substrate

conn  = from_edgelist(edges_df)                 # columns pre / post / weight
table = NeuronTable.from_dataframe(ann_df, id_col="node_name",
                                   superclass_col="kind", class_col="group",
                                   input_labels=("source",),
                                   output_labels=("sink",))
sub   = Substrate.build(conn, table, budget=2048, seed=0)

String labels, uuids and integer ids all work. The annotation table is canonicalised once at the boundary, so downstream code does not care what your columns are called.


For different purposes

Connectomics / circuit neuroscience. The census tooling answers structural questions directly: which pools are self-contained, which are broadcast relays, how far a population reaches, what a lesion would cost.

python tools/census_wm_ports.py            # in/out fractions, reciprocity, mixing, leak
sub.internal_fraction(cx_ids)      # how self-contained is this pool
sub.recurrence_report()            # reciprocal fraction, mixing eigenvalue, self-loops
echo_probe(sub.W, sources)         # does activity survive the input stopping
broadcast_coverage(sub.W, dan_pos) # what fraction of the slice a pool can reach

pyfly_lightning.gates.structure_function_gate runs the classic structure→function test: stimulate a named population, find which outputs the real wiring drives, then measure activation probability under real versus shuffled weights.

Spiking / neuromorphic engineering. The LIF simulator ships two backends that are bit-identical, which makes correctness testable rather than assumed:

from pyfly_lightning.sim.lif import LifNumpy, LifTorch      # same dynamics, numpy vs torch
LifTorch(W, device="cuda")                        # sparse event-driven step
LifTorch(W, differentiable=True)                  # smooth gate: gradients flow

differentiable=True keeps the autograd graph so the whole pipeline trains end to end. set_trainable_weights() keeps the connectome's wiring while making every edge weight a parameter. Throughput and memory are measured rather than estimated: python -m pyfly_lightning.sim.bench --device cuda reports steps/sec and the ceiling (a 16k-node, 528k-edge substrate uses 50 MB of a 6 GB card; latency is ~0.15 ms/step).

Machine learning / architecture search. The substrate is a sparse recurrent architecture with a natural I/O boundary and structure-derived grouping.

from pyfly_lightning.model.routing import build_sensory_surface, GroupRouter
surface = build_sensory_surface(table, level="medium")   # ~10 modality groups
router  = GroupRouter(surface, latent_dim=32, learnable=True)
router.param_report()      # parameter count, compression vs dense, optimiser fit

level="fine" gives 1,269 structural groups (optic-lobe hex columns, olfactory glomeruli, labelled lines); level="medium" collapses to ~10 modalities, which brings full-covariance CMA-ES into range. ESTrainer and OuterLoop cover the black-box side; pyfly_lightning.train.outer uses Vizier when installed and records which backend ran.

Reinforcement learning. DQNAgent and PPOAgent train heads on substrate features, with an environment that provides genuinely multi-step trajectories, plus an inference-safe neuromodulatory channel:

from pyfly_lightning.envs import CueDelayChoice
from pyfly_lightning.train.rl import DQNAgent, PPOAgent, RLConfig
agent = DQNAgent(model, CueDelayChoice(delay=8, act_every_step=True), RLConfig())

Model compression. MaskPruner searches a group-level mask under a tolerance contract and rollback; EvoPruner runs NSGA-II over group genomes and returns a Pareto front over (behaviour, active nodes, active edges).

model.freeze(adapter=True)                      # required precondition
from pyfly_lightning.prune import MaskPruner, PruneConfig
rep = MaskPruner(model, PruneConfig(tolerance=0.05)).prune(Xtr, ytr, Xva, yva)

Dynamical systems. pyfly_lightning.sim.probe distinguishes propagation from recirculation, which is the distinction that matters when the object of study is a recurrent graph rather than a classifier.


The staged pipeline

user data ──► encoder ──► router ──► [ frozen / trainable substrate ] ──► readout ──► task
              (ES)        (ES)              │                                 (backprop)
                                             └─► neuromodulatory channel ──┘
stage what it does optimised by
stage1_es finds which substrate neurons your input reaches ES, scored on a held-out split
stage2_decode trains the readout over the output population Adam
stage3_modulate trains a recurrent readback through the DAN/PPL1 pool Adam

The modulation channel carries the substrate's own pass-1 output back onto the dopaminergic pool and runs a second pass, which is available at inference. Feeding the label or the error into the channel would make the reported accuracy unreproducible on unlabelled data.


Controls and reporting

Every fit reports the same set of nulls, and a verdict:

control what it isolates
rewired degree-preserving rewiring: does the specific wiring matter
shuffled_weight the same wiring with weights permuted
pooled_features ridge on the raw input features, same readout class
no_plasticity weight updates disabled

FlyModule exposes the same contract, and pyfly_lightning.train.controls.required_controls_report assembles it. FitResult.verdict is SUPPORTED when the fit beats every null and UNSUPPORTED otherwise, with the scores recorded either way.


What has been measured so far

Recorded so you can design around it; the full protocols, ablations and bugs are in docs/architecture.md.

experiment result
MNIST, substrate vs its own rewiring 0.154 vs 0.154 — identical
MNIST, staged pipeline vs direct input + PPL1 (3 seeds) +7.6 to +15.3 points
Neuromodulatory channel, frozen / trainable substrate +11.6 / +25.2 points
Iris, low-data, per-fold ES (3 seeds x 5 folds) fly 0.7644 vs rewired 0.7711
Iris, substrate vs raw 4 features 0.7644 vs 0.8289
Trainable substrate: real vs shuffled weights 0.6860 vs 0.7360
Degree-preserving rewiring of the CX pool mixing eigenvalue changes 0.12%
CX forced into a slice vs not late_echo 30.0 vs 0.0
Mask pruning under a tolerance contract edges −67.6%, held-out score unchanged
RTX 2060 6 GB, 16,384 nodes / 528,013 edges 0.148 ms/step, 50 MB, 99.2% headroom

A note on naming that the fields make unavoidable: what these runs call fly is the substrate wired from a real connectome, and rewired is the same graph with its edges re-randomised under a degree-preserving swap. The comparison is built into the report because a result without it cannot be placed.


Package layout

pyfly_lightning/data/        registry, sha256-pinned downloader, cache, NeuronTable,
                   SparseConnectome, sign derivation with an audit trail
pyfly_lightning/model/       structured routing (hex columns / glomeruli / labelled lines),
                   substrate slicing, the four nulls
pyfly_lightning/sim/         LIF (numpy + torch, bit-identical), probes, benchmark
pyfly_lightning/train/       FlyModule (a real lightning.LightningModule), ES, outer loop,
                   DQN, PPO, the staged pipeline, environments
pyfly_lightning/reward.py    objective, bounded reward calibration, credit channel, plasticity
pyfly_lightning/prune.py     tolerance-contract mask pruning
pyfly_lightning/prune_evo.py NSGA-II over group genomes, Pareto front
pyfly_lightning/gates.py     structure→function gate with a shuffled-weight control
pyfly_lightning/auto.py      automatic configuration with recorded reasons
pyfly_lightning/api.py       load / io / slice / port / gate / reward_bus / fit

Docs

  • docs/quickstart.md — install, plan, fit, bring your own connectome
  • docs/architecture.md — design, the full measurement record, freeze semantics, the pruning contract, and a milestone-by-milestone checklist
  • docs/census-wm-ports.md — which neuron pools are self-contained, measured

Examples

python examples/quickstart.py                                 # both modes, one API
python examples/mnist.py                                      # Lightning, four controls
python examples/emotion_pipeline.py                           # the six-arm comparison
python examples/delayed_match.py                              # a memory task
python examples/lowdata.py                                    # low-data regime, CV
python examples/freeze_prune.py --method mask                 # prune under tolerance
python examples/freeze_prune.py --method evo --grouping depth # NSGA-II Pareto front
python examples/reward_bus.py                                 # calibration + plasticity
python examples/rl_dqn.py                                     # RL, with controls

Tests

pytest -q                    # 104 tests; data-dependent ones skip if the cache is empty
pyfly-lightning get malecns10 && pytest -q

Citation

CITATION.cff is included. Cite the datasets alongside this software:

  • Berg et al., Sexual dimorphism in the complete connectome of the Drosophila male central nervous system, Cell 2026 — MaleCNS v1.0
  • Dorkenwald et al., Neuronal wiring diagram of an adult brain, Nature 2024 — FlyWire
  • Schlegel et al., Whole-brain annotation and multi-connectome cell typing of Drosophila, Nature 2024 — cell types and the I/O classification
  • Shiu et al., A Drosophila computational brain model reveals sensorimotor processing, Nature 2024 — the LIF parameterisation
  • Eckstein et al., Cell 2024 — neurotransmitter predictions used for synapse sign

Check each dataset's own licence and citation terms; MaleCNS is CC-BY, and FlyWire data has its own guidelines.

Licence

MIT. Linked datasets keep their own licences.

Contributing

Issues and pull requests are welcome, particularly new task domains, additional connectomes, and simulation backends. docs/architecture.md records the measurements behind the current defaults, so a change to a default should come with the measurement that motivates it.

Metadata

Release files for pyfly-lightning 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pyfly-lightning 0.1.0
File Size Uploaded
pyfly_lightning-0.1.0.tar.gz 177.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pyfly-lightning 0.1.0
File Interpreter ABI Platform
pyfly_lightning-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 272.8 kB

Release files / pyfly_lightning-0.1.0.tar.gz

Download URL pyfly_lightning-0.1.0.tar.gz
Size 177.7 kB
Tags Source
SHA-256 checksum
How to use checksums
418ab21d1ea0f19edae4b95bd67a203a28837d5361279fdfaf86e499aad23913
BLAKE2b-256 checksum
How to use checksums
f43533f605737993a7dbe0ee1babb1e5b7b41c0a98135547ca87e08a9704923d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release files / pyfly_lightning-0.1.0-py3-none-any.whl

Download URL pyfly_lightning-0.1.0-py3-none-any.whl
Size 95.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
047f3464cc2e6fa7ce7de86f4b1b3f20e6cdfda9019ac7fd86b0e486a454901d
BLAKE2b-256 checksum
How to use checksums
8881f2af53ac6b6bdfb496e81f0e5be6a6e72cfba2508bf42ec84ded2733d6c1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 20, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page