PopuLoRA (wip)
Implementation and explorations into PopuLoRA, Co-Evolving LLM Populations for Reasoning Self-Play, from Roger Castanyer et al at vmax.ai
Install
pip install populora
Usage
import torch
import torch.nn as nn
from populora import Population
# 2-layer MLP
model = nn.Sequential(
nn.Linear(2, 8),
nn.ReLU(),
nn.Linear(8, 1)
)
# wrap with Population
pop = Population(
model,
pop_size = 16,
low_rank = 4,
lora_targets = ['0', '2']
)
state = torch.randn(1, 4, 2)
# evaluate population against environment
# `individuals` also accepts a list of individual ids (one per sample)
preds = pop(state, all_individuals = True)
labels = torch.randn(1, 4, 1)
fitnesses = -((preds - labels ) ** 2).reshape(16, -1).mean(dim = -1)
# selection
result = pop.select(
selection_type = 'deterministic',
fitnesses = fitnesses,
survive_frac = 0.5
)
# parent selection
parents = pop.select_parents(
selection_type = 'tournament',
fitnesses = fitnesses,
num_children = len(result.selected_out_indices),
culled = result.selected_out_indices
)
# crossover
pop.crossover_('average', parents, result.selected_out_indices)
# mutate newly generated offspring, preserving surviving elite parents
pop.mutate_('full_gaussian', individuals = result.selected_out_indices)
# alternatively, mutate the entire population
pop.mutate_('full_gaussian', all_individuals = True)
# do the above in a for loop
# ...
# then pick the highest fitness individual and resume RL or fine-tuning on the base model
model = pop.select_and_merge_best_(fitnesses)
Distributed Evolution
Evolution parallelizes trivially - each rank evaluates its share of the population against the environment, the fitnesses are gathered, and the evolution step runs identically on every rank
The population is automatically moved to the distributed device (each rank's local GPU) on construction - pass device to Population to override
from time import sleep
import torch
from torch import nn
from populora import Population, is_main_rank
model = nn.Sequential(
nn.Linear(8, 16),
nn.ReLU(),
nn.Linear(16, 1)
)
pop = Population(
model,
pop_size = 16,
low_rank = 2,
lora_targets = ['0', '2']
)
x = torch.randn(1, 8)
def eval_env(population, idx):
sleep(0.1)
with torch.no_grad():
# seed the environment with population.eval_seed (shared, auto-synced across ranks)
return population(x, individual = idx).abs().mean().item() + torch.randn(1).item()
for gen in range(10):
# distributed evaluation
fitnesses = pop.evaluate_distributed(eval_env)
if is_main_rank():
print(f'gen {gen:02d} | best: {fitnesses.max():.3f} | mean: {fitnesses.mean():.3f}')
# evolution step
pop.evolve_(fitnesses)
run on 4 processes
torchrun --standalone --nproc-per-node=4 evolve.py
or across machines
torchrun --nnodes=4 --nproc-per-node=1 --rdzv-endpoint=$MASTER_HOST:29500 evolve.py
Citations
@misc{castanyer2026populoracoevolvingllmpopulations,
title = {PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play},
author = {Roger Creus Castanyer and Geoffrey Bradway and Lorenz Wolf and Maxwill Lin and Augustine N. Mavor-Parker and Matthew James Sargent},
year = {2026},
eprint = {2605.16727},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2605.16727},
}
@misc{schmidhuber2012powerplaytrainingincreasinglygeneral,
title = {POWERPLAY: Training an Increasingly General Problem Solver by Continually Searching for the Simplest Still Unsolvable Problem},
author = {Jürgen Schmidhuber},
year = {2012},
eprint = {1112.5309},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/1112.5309},
}
@misc{xu2026selfimprovinglanguagemodelsbidirectional,
title = {Self-Improving Language Models with Bidirectional Evolutionary Search},
author = {Guowei Xu and Zhenting Qi and Huangyuan Su and Weirui Ye and Himabindu Lakkaraju and Sham M. Kakade and Yilun Du},
year = {2026},
eprint = {2605.28814},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2605.28814},
}
@misc{bahlousboldi2026vectorpolicyoptimizationtraining,
title = {Vector Policy Optimization: Training for Diversity Improves Test-Time Search},
author = {Ryan Bahlous-Boldi and Isha Puri and Idan Shenfeld and Akarsh Kumar and Mehul Damani and Sebastian Risi and Omar Khattab and Zhang-Wei Hong and Pulkit Agrawal},
year = {2026},
eprint = {2605.22817},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2605.22817},
}
@misc{bailey2026scalingselfplayselfguidance,
title = {Scaling Self-Play with Self-Guidance},
author = {Luke Bailey and Kaiyue Wen and Kefan Dong and Tatsunori Hashimoto and Tengyu Ma},
year = {2026},
eprint = {2604.20209},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2604.20209},
}
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file populora-0.1.12.tar.gz.
File metadata
- Download URL: populora-0.1.12.tar.gz
- Upload date:
- Size: 19.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.8.17
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
51882b9acf18505a3f84c8229c3df27ebb24b3ad28c15556f9952618e1c70b83
|
|
| MD5 |
35dc75bb842bc8a4336979011fe80513
|
|
| BLAKE2b-256 |
623314ddc6d5cc362a1a40340ef31cc0685cfc62d207d0281298648a76cfb538
|
File details
Details for the file populora-0.1.12-py3-none-any.whl.
File metadata
- Download URL: populora-0.1.12-py3-none-any.whl
- Upload date:
- Size: 18.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.8.17
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4477300e7150c54e3ed62342b7d6ee3652a9ada9535c4cff1e1769abb22bd445
|
|
| MD5 |
267ee0e8f93a65e0144af0032e0dd1af
|
|
| BLAKE2b-256 |
1c7878a422937fdbe9557af8b920746f435fb509ead6b18dda716b1edfed793e
|