alignment-risk
alignment-risk is a Python package for pre-flight alignment risk diagnostics during fine-tuning.
It estimates whether updates are likely to drift into safety-sensitive directions before you commit to a run.
Full documentation site: sirhan1.github.io/modelFineTuneRiskAssessment
Quick Start (1 minute)
Install:
pip install alignment-risk
Run the built-in demo:
alignment-risk demo --output-dir artifacts
Outputs:
artifacts/sensitivity_map.pngartifacts/safety_decay_forecast.png
Installation
PyPI:
pip install alignment-risk
Local development (Apple Silicon convenience):
make setup
source .venv/bin/activate
Local development (manual, cross-platform):
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
Basic Usage (Python API)
from alignment_risk import AlignmentRiskPipeline, PipelineConfig
config = PipelineConfig(mode="lora") # "full" or "lora"
pipeline = AlignmentRiskPipeline(config)
report = pipeline.run(
model=model,
safety_dataloader=safety_loader,
safety_loss_fn=safety_loss_fn,
fine_tune_dataloader=ft_loader,
fine_tune_loss_fn=ft_loss_fn,
)
print(report.warning)
print(report.forecast.collapse_step)
CLI
alignment-risk --help
alignment-risk demo --output-dir artifacts
alignment-risk demo --mode lora --output-dir artifacts
What It Computes
- Low-rank safety sensitivity subspace from empirical Fisher geometry.
- Initial overlap risk (projection of first update into sensitive subspace).
- Curvature coupling risk (second-order drift signal).
- Quartic-style stability forecast and collapse-step estimate.
Modes
full: analyze all selected trainable parameters.lora: analyze only trainable LoRA adapter parameters (lora_,lora_A,lora_Bby default).
LoRA Mitigation (AlignGuard-style)
After mode="lora" risk analysis, attach a regularizer to penalize drift in sensitive directions:
from alignment_risk import AlignmentRiskPipeline, PipelineConfig, AlignGuardConfig
config = PipelineConfig(mode="lora")
pipeline = AlignmentRiskPipeline(config)
report = pipeline.run(
model=model,
safety_dataloader=safety_loader,
safety_loss_fn=safety_loss_fn,
fine_tune_dataloader=ft_loader,
fine_tune_loss_fn=ft_loss_fn,
)
mitigator = pipeline.build_lora_mitigator(
model,
report.subspace,
config=AlignGuardConfig(lambda_a=0.25, lambda_t=0.5, lambda_nc=0.1, alpha=0.5),
)
task_loss = ft_loss_fn(model, batch)
breakdown = mitigator.regularized_loss(task_loss)
breakdown.total_loss.backward()
Use mitigator.reset_reference() to re-anchor regularization at the current adapter state.
Performance / Accuracy Controls
FisherConfig supports speed/accuracy tradeoffs:
gradient_collection:"loop"(default),"auto","vmap".subspace_method:"svd","randomized_svd","diag_topk".vmap_chunk_size: optional chunking for lower memory.target_explained_variance: auto-rank selection (default0.9).
Example:
from alignment_risk import AlignmentRiskPipeline, PipelineConfig
config = PipelineConfig()
config.fisher.gradient_collection = "vmap"
config.fisher.subspace_method = "randomized_svd"
config.fisher.vmap_chunk_size = 16
Theory and Math References
- [AIC-2026] Springer, Max, et al. (2026). The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety. arXiv:2602.15799v1. PDF: https://arxiv.org/pdf/2602.15799
- [ALIGNGUARD-2025] Das, Amitava, et al. (2025). AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization. arXiv:2508.02079v1. PDF: https://arxiv.org/pdf/2508.02079
Detailed internal mappings and equations:
docs/SOURCES.mddocs/MATH.md
Development
make install
make test
make lint
make typecheck
make build
make check-dist
See CONTRIBUTING.md for contribution workflow details.
Release files for alignment-risk 1.14.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| alignment_risk-1.14.1.tar.gz | 31.1 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| alignment_risk-1.14.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 58.0 kB
Release files / alignment_risk-1.14.1.tar.gz
| Download URL | alignment_risk-1.14.1.tar.gz |
|---|---|
| Size | 31.1 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
32d7b840b9cd877cc3a0efa04ad21d8be86fd44e8791c211d39ca2b07302b5be
|
|
BLAKE2b-256 checksum How to use checksums |
88dcc688c7a9032e563ec02c54108f9de8bd085638eb08e6fd62451538eb222e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 25, 2026.
Transparency logRelease files / alignment_risk-1.14.1-py3-none-any.whl
| Download URL | alignment_risk-1.14.1-py3-none-any.whl |
|---|---|
| Size | 26.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
3d1c4f68dc999aecfcb5e91dc1905f38d074a80619221795eda444333c9425be
|
|
BLAKE2b-256 checksum How to use checksums |
3d7f032d1aff0599bf4fbf9d13b576ff716180ac4b80cdab019d3d77572af36b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/6.1.0 CPython/3.13.7
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Feb 25, 2026.
Transparency log