Skip to main content

PopuLoRA (wip)

Implementation of PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play (Roger Castanyer et al., vmax.ai). Maintains a population of LoRA adapters on a shared base model, evolves them via selection, crossover, and mutation, and merges the best individual back into the base model.

Install

pip install populora

Quick start

import torch
import torch.nn as nn
from populora import Population

model = nn.Sequential(nn.Linear(2, 8), nn.ReLU(), nn.Linear(8, 1))

pop = Population(model, pop_size = 16, low_rank = 4, lora_targets = ['0', '2'])

state = torch.randn(1, 4, 2)
preds = pop(state, all_individuals = True)  # one routed forward over every individual

labels = torch.randn(1, 4, 1)
fitnesses = -((preds - labels) ** 2).reshape(16, -1).mean(dim = -1)

result = pop.select('deterministic', fitnesses, survive_frac = 0.5)

parents = pop.select_parents(
    'tournament',
    fitnesses,
    num_children = len(result.selected_out_indices),
    culled = result.selected_out_indices  # parents come from the survivors
)

pop.crossover_('average', parents, result.selected_out_indices)  # offspring overwrite the culled
pop.mutate_('full_gaussian', individuals = result.selected_out_indices)

model = pop.merge_(fitnesses.argmax())  # merge the best individual back in

pop.save_individual('best.pt', fitnesses.argmax())  # or save just the best individual's weights

pop.evolve_(fitnesses) runs selection, parent selection, crossover, and mutation in one step. Batch evaluation also supports pop(x, individuals = [ids]) to route each sample to its own individual.

Per-weight-matrix operators

mutation_type, epsilon, and crossover_type each accept a dict keyed by LoRA target instead of a scalar — set the mutation rate (or the operator itself) per weight matrix. Keys match a layer's dotted module path or its storage key exactly, else by glob pattern (first match wins); a 'default' entry catches everything left, coverage must be total, and unmatched patterns raise.

pop.evolve_(
    fitnesses,
    mutation_type = {
        'encoder.*': 'svd_structured',   # glob over the dotted module path
        'default': 'full_gaussian'
    },
    epsilon = {
        'encoder.proj_in': 0.05,         # exact dotted path
        'encoder_proj_out': 0.2,         # exact storage key
        'head': 0.,                      # freeze this matrix entirely
        'default': 0.1
    },
    crossover_type = {
        '*to_q*': 'svd_subspace',
        'default': 'average'
    }
)

Any other operator kwarg can vary per target too via the explicit PerTarget wrapper: alpha = PerTarget({'*.proj_in': 5., 'default': 10.}). Targets sharing the same resolved parameters are dispatched together in one call, so per-target specs cost nothing when they collapse to a single group.

Distributed evolution

Each rank evaluates its share of the population and the fitnesses are gathered, so every rank evolves in lockstep. The population is moved to the local device on construction.

import torch
from torch import nn
from populora import Population, is_main_rank

pop = Population(
    nn.Sequential(nn.Linear(8, 16), nn.ReLU(), nn.Linear(16, 1)),
    pop_size = 16,
    low_rank = 2,
    lora_targets = ['0', '2']
)

def eval_env(population, idx):
    return population(torch.randn(1, 8), individual = idx).abs().mean().item()

for gen in range(10):
    fitnesses = pop.evaluate_distributed(eval_env)

    if is_main_rank():
        print(f'gen {gen:02d} | best: {fitnesses.max():.3f} | mean: {fitnesses.mean():.3f}')

    pop.evolve_(fitnesses)
torchrun --standalone --nproc-per-node=4 evolve.py

Only the LoRA weights are synced across ranks before evaluation (the base model is shared) — pass sync_base_model = True to evaluate_distributed to broadcast it too.

Environment evolution

evolve_with_env evolves a population against any MDP-style environment in one call — gymnasium, dm_control, isaac, maniskill, pybullet, pufferlib, or any simulator — and returns the merged best policy.

from torch import nn
from dm_control import suite
from populora import evolve_with_env

policy, history = evolve_with_env(
    [suite.load('cartpole', 'balance') for _ in range(16)],  # envs, a list, a vector env, or a factory
    nn.Sequential(nn.Linear(5, 32), nn.ReLU(), nn.Linear(32, 2)),
    pop_size = 16,
    low_rank = 16,
    action = lambda logits: logits.argmax(-1) * 2 - 1,
    num_generations = 25,
    horizon = 1000,
    seed = 0,
    progress = True,
    return_history = True  # per-generation best / mean
)

lora_targets auto-discovers every Linear layer when omitted. Pass target_fitness to stop early, checkpoint_dir to write latest.pt (every checkpoint_every generations) and best.pt (on new bests), and resume = True to pick up from the latest checkpoint.

Continuous policies emit actions on (-1, 1) — pass to_range = (-2., 2.) (or any env action range) to evolve_with_env / interact_with_env and the actions are rescaled before stepping the env. An action fn built with make_action carries its own range, e.g. a beta with beta_rescale_neg_one_one = False emits on (0, 1) instead — the interactor rescales from that automatically. interact_with_env(from_range = ...) is the fallback for custom action functions, and rescale_from_range_to_range is exported too.

For a custom loop, interact_with_env exposes the underlying EnvInteractor — one routed forward over the active slots per timestep, distributed across ranks under torchrun:

from torch import nn
from dm_control import suite
from populora import interact_with_env

interactor = interact_with_env([suite.load('cartpole', 'balance') for _ in range(16)])

backbone = nn.Sequential(nn.Linear(5, 32), nn.ReLU(), nn.Linear(32, 2))
population = interactor.population(backbone, pop_size = 16, low_rank = 16)

for gen in range(25):
    fitnesses = interactor.evaluate(
        population,
        action = lambda logits: logits.argmax(-1) * 2 - 1,
        horizon = 1000
    )

    population.evolve_(fitnesses)

policy = population.merge_(fitnesses.argmax())

Custom fitness functions may take (population, individuals) (batched), (population, idx) (per index), or (population) (all at once) — detected automatically.

Coevolution

Wrap populations whose fitnesses derive from one another's outputs — e.g. a proposer proposes test inputs while a solver is scored on them, and vice versa. Each population supplies a probe (produces its outputs for a step) and a fitness (scores it); parameters named after a population receive its outputs, computed once per step in dependency order.

import torch
from torch import nn
from populora import Population, Coevolve, HallOfFame

pop_size = 8
K = 4  # archived solver champions sampled per generation

proposer = Population(nn.Sequential(nn.Linear(1, 16), nn.ReLU(), nn.Linear(16, 1), nn.Tanh()), pop_size = pop_size, low_rank = 2, lora_targets = ['0', '2'])
solver = Population(nn.Sequential(nn.Linear(1, 16), nn.ReLU(), nn.Linear(16, 1)), pop_size = pop_size, low_rank = 2, lora_targets = ['0', '2'])

solver_hof = HallOfFame()

def probe_proposer(coevolve):
    return coevolve.proposer(torch.randn(proposer.pop_size, 1), all_individuals = True)  # (P, 1) proposed inputs, one per individual

def probe_solver(coevolve, proposer_outputs):
    return coevolve.solver(proposer_outputs.repeat(solver.pop_size, 1), all_individuals = True)

def fitness_solver(solver_outputs, proposer_outputs):
    target = torch.sin(torch.pi * proposer_outputs.repeat(solver.pop_size, 1))
    return -((solver_outputs - target) ** 2).reshape(solver.pop_size, -1).mean(dim = 1)

def fitness_proposer(proposer_outputs, solver_outputs):
    # a proposal earns credit for stumping both the current solver and the archived champions

    target = torch.sin(torch.pi * proposer_outputs.repeat(solver.pop_size, 1))
    err = ((solver_outputs - target) ** 2).reshape(solver.pop_size, -1).mean(dim = 0)

    arch = solver_hof.probe(solver, proposer_outputs, K)  # (k, P, 1) through sampled champions

    if arch is not None:
        err = err + ((arch - torch.sin(torch.pi * proposer_outputs)) ** 2).mean(dim = (0, 2))

    return err

coevolve = Coevolve(populations = dict(
    proposer = dict(population = proposer, probe = probe_proposer, fitness = fitness_proposer),
    solver = dict(population = solver, probe = probe_solver, fitness = fitness_solver)
))

for gen in range(100):
    fitnesses = coevolve.step()
    solver_hof.add_champion(solver, fitnesses['solver'], generation = gen)

# when it is over, pluck each population's champion

for name in fitnesses:
    coevolve[name].save_individual(f'winners/{name}.pt', fitnesses[name].argmax())

Any individual can be saved by index (save_individual(path, 3)), loaded back into a population slot (load_individual), and the whole population rebuilt with Population.from_checkpoint. demo_coevolve_three.py runs the same loop with a third population - a judge that learns to predict the solver's errors - and shows the winners of all three falling out the same way.

coevolve.history records the best / mean fitness per population; coevolve.step(distributed = True) splits the probes across ranks (round-robin, outputs broadcast raw). Probes must form a chain — a probe depending on its own outputs raises at construction; only fitnesses may close a circle, since every population is probed before any fitness is derived.

Generation

Autoregressively decode from the population in one batched loop — one routed forward per step over the samples still active.

import torch
from x_transformers import TransformerWrapper, Decoder
from populora import Population, generate

model = TransformerWrapper(
    num_tokens = 256,
    max_seq_len = 128,
    attn_layers = Decoder(dim = 128, depth = 2, heads = 2)
)

pop = Population(
    model,
    pop_size = 8,
    low_rank = 4,
    lora_targets = ['attn_layers.layers.0.1.to_q', 'attn_layers.layers.0.1.to_k', 'attn_layers.layers.0.1.to_v']
)

prompts = torch.randint(0, 256, (8, 16))  # one prompt per individual

seqs = generate(
    pop,
    prompts,
    all_individuals = True,
    max_len = 64,
    eos_token = 255,
    cache_kwargs = dict(return_intermediates = True)
)

Samples can be routed to explicit individuals (individual = 3 or individuals = [...], one id per prompt). Early finishers (eos_token, or a stop_fn(tokens, logits, step)) are compacted out of the batch. Huggingface-style caching works too (cache_kwarg = 'past_key_values', cache_last_token = True), and micro_batch chunks the routed forward to cap the p-fold activation blowup.

Citations

@misc{castanyer2026populoracoevolvingllmpopulations,
    title   = {PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play},
    author  = {Roger Creus Castanyer and Geoffrey Bradway and Lorenz Wolf and Maxwill Lin and Augustine N. Mavor-Parker and Matthew James Sargent},
    year    = {2026},
    eprint  = {2605.16727},
    archivePrefix = {arXiv},
    primaryClass = {cs.AI},
    url     = {https://arxiv.org/abs/2605.16727},
}
@misc{schmidhuber2012powerplaytrainingincreasinglygeneral,
    title    = {POWERPLAY: Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem},
    author   = {Jürgen Schmidhuber},
    year     = {2012},
    eprint   = {1112.5309},
    archivePrefix = {arXiv},
    primaryClass = {cs.AI},
    url      = {https://arxiv.org/abs/1112.5309},
}
@misc{xu2026selfimprovinglanguagemodelsbidirectional,
    title   = {Self-Improving Language Models with Bidirectional Evolutionary Search},
    author  = {Guowei Xu and Zhenting Qi and Huangyuan Su and Weirui Ye and Himabindu Lakkaraju and Sham M. Kakade and Yilun Du},
    year    = {2026},
    eprint  = {2605.28814},
    archivePrefix = {arXiv},
    primaryClass = {cs.CL},
    url     = {https://arxiv.org/abs/2605.28814},
}
@misc{bahlousboldi2026vectorpolicyoptimizationtraining,
    title   = {Vector Policy Optimization: Training for Diversity Improves Test-Time Search},
    author  = {Ryan Bahlous-Boldi and Isha Puri and Idan Shenfeld and Akarsh Kumar and Mehul Damani and Sebastian Risi and Omar Khattab and Zhang-Wei Hong and Pulkit Agrawal},
    year    = {2026},
    eprint  = {2605.22817},
    archivePrefix = {arXiv},
    primaryClass = {cs.LG},
    url     = {https://arxiv.org/abs/2605.22817},
}
@misc{bailey2026scalingselfplayselfguidance,
    title   = {Scaling Self-Play with Self-Guidance},
    author  = {Luke Bailey and Kaiyue Wen and Kefan Dong and Tatsunori Hashimoto and Tengyu Ma},
    year    = {2026},
    eprint  = {2604.20209},
    archivePrefix = {arXiv},
    primaryClass = {cs.LG},
    url     = {https://arxiv.org/abs/2604.20209},
}
@misc{petrenko2023dexpbt,
    title    = {DexPBT: Scaling up Dexterous Manipulation for Hand-Arm Systems with Population Based Training},
    author   = {Aleksei Petrenko and Arthur Allshire and Gavriel State and Ankur Handa and Viktor Makoviychuk},
    year     = {2023},
    eprint   = {2305.12127},
    archivePrefix = {arXiv},
    primaryClass = {cs.RO},
    url      = {https://arxiv.org/abs/2305.12127},
}

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

populora-0.3.0.tar.gz (55.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

populora-0.3.0-py3-none-any.whl (60.1 kB view details)

Uploaded Python 3

File details

Details for the file populora-0.3.0.tar.gz.

File metadata

  • Download URL: populora-0.3.0.tar.gz
  • Upload date:
  • Size: 55.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.8.17

File hashes

Hashes for populora-0.3.0.tar.gz
Algorithm Hash digest
SHA256 5ffc40b723973d060f2acf00b8d17e8258350e19c012acc18a3192de7b52079b
MD5 d21573af1f3672818afd8200f6c63444
BLAKE2b-256 5a9ee54c47ca24db64fe9c68ab2d9b18183ef2b8cb0957424c1b8116f9a5243d

See more details on using hashes here.

File details

Details for the file populora-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: populora-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 60.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.8.17

File hashes

Hashes for populora-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 4a08be5f9c5ca796e4c8eaf996461501c1de4978edbfa1d53469ae84ca08a0d3
MD5 8c90cfb54063c7e43ccbc06a3ecf6e2d
BLAKE2b-256 b7e176960cbfc20e4cbfffcd1abbf5e6d489a4e5c1b122e8eb5b65894caf8c69

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 files

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

0.2.0

2 files

0.1.34

2 files

0.1.32

2 files

0.1.31

2 files

0.1.30

2 files

0.1.29

2 files

0.1.28

2 files

0.1.27

2 files

0.1.26

2 files

0.1.25

2 files

0.1.24

2 files

0.1.23

2 files

0.1.21

2 files

0.1.20

2 files

0.1.19

2 files

0.1.18

2 files

0.1.16

2 files

0.1.15

2 files

0.1.14

2 files

0.1.12

2 files

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

0.1.8

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

0.0.17

2 files

0.0.16

2 files

0.0.15

2 files

0.0.14

2 files

0.0.12

2 files

0.0.11

2 files

0.0.10

2 files

0.0.9

2 files

0.0.8

2 files

0.0.7

2 files

0.0.6

2 files

0.0.5

2 files

0.0.4

2 files

0.0.3

2 files

0.0.2

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page