BayesCurveFit: Enhancing Biological Data Analysis with Robust Curve Fitting and FDR Detection
Project description
BayesCurveFit: Enhancing Curve Fitting in Drug Discovery Data Using Bayesian Inference
Overview
BayesCurveFit is a Python package designed to apply Bayesian inference for curve fitting, especially tailored for undersampled and outlier-contaminated data. It supports advanced model fitting and uncertainty estimation for biological data, such as dose-response curves in drug discovery.
Manuscript and Datasets
A preprint can be found at: https://www.biorxiv.org/content/10.1101/2025.02.23.639769v1
Data and supplementary data used in the manuscript can be found at https://zenodo.org/records/14948538
The BayesCurveFit pipeline for processing Kinobeads data is available and can be found in the examples/pipeline directory of the repository.
Key Features
- Multivariate Bayesian Model Averaging: Single multivariate Gaussian Mixture Model captures cross-parameter correlations
- Robust MCMC Sampling: Emcee-based sampling with automated convergence diagnostics (R-hat < 1.1, ESS > 100)
- Correlation-Aware Uncertainty: Full covariance matrix estimation using law of total variance
- Model Selection: Posterior Error Probability (PEP) for automated model comparison
- Outlier Handling: Robust error estimation with nuisance parameter optimization
- Convergence Monitoring: Automated diagnostics with multiple restart attempts
Installation Guide
Prerequisites
- Python 3.11 or higher
- uv package manager (recommended)
Install from PyPI (Recommended)
To install the latest stable release from PyPI:
pip install bayescurvefit
Or with uv (faster):
uv pip install bayescurvefit
Install from Source
To install the latest development version:
git clone https://github.com/ndu-bioinfo/BayesCurveFit.git
cd BayesCurveFit
uv sync
Development Setup
For development with all dependencies:
git clone https://github.com/ndu-bioinfo/BayesCurveFit.git
cd BayesCurveFit
uv sync --dev
make test # Run tests
Quick Start
Basic Usage
In this example, we use the log-logistic 4-parameter model to fit dose-response data.
import numpy as np
from bayescurvefit.execution import BayesFitModel
# Define the curve fitting function
def log_logistic_4p(x: np.ndarray, pec50: float, slope: float, front: float, back: float) -> np.ndarray:
with np.errstate(over="ignore", under="ignore", invalid="ignore"):
y = (front - back) / (1 + 10 ** (slope * (x + pec50))) + back
return y
# Input data
x_data = np.array([-9.0, -8.3, -7.6, -6.9, -6.1, -5.4, -4.7, -4.0])
y_data = np.array([1.12, 0.74, 1.03, 1.08, 0.76, 0.61, 0.39, 0.38])
params_range = [(5, 8), (0.01, 10), (0.28, 1.22), (0.28, 1.22)] # Parameter bounds
param_names = ["pec50", "slope", "front", "back"]
# Run Bayesian curve fitting
run = BayesFitModel(
x_data=x_data,
y_data=y_data,
fit_function=log_logistic_4p,
params_range=params_range,
param_names=param_names,
)
Retrieve Results
Get the fitting results with parameter estimates and uncertainties:
# Get results
results = run.get_result()
print(results)
Output:
fit_pec50 5.855744
fit_slope 1.088774
fit_front 1.063355
fit_back 0.376912
std_pec50 0.184996
std_slope 0.399462
std_front 0.061905
std_back 0.050635
est_std 0.089587
null_mean 0.762365
rmse 0.122646
pep 0.062768
convergence_warning False
Key Results:
fit_*: Best-fit parameter estimatesstd_*: Parameter uncertainties (standard deviations)rmse: Root mean square errorpep: Posterior Error Probability (model comparison)convergence_warning: MCMC convergence status
Visualization
Plot the fitted curve and parameter distributions:
import matplotlib.pyplot as plt
# Plot fitted curve
fig, ax = plt.subplots()
run.analysis.plot_fitted_curve(ax=ax)
ax.set_xlabel("Concentration [M]")
ax.set_ylabel("Relative Site Response")
plt.show()
Text(0, 0.5, 'Relative Site Response')
Parameter Analysis
Visualize parameter correlations and sampling diagnostics:
# Parameter pairwise comparisons
run.analysis.plot_pairwise_comparison(figsize=(10, 10))
# Parameter distributions
run.analysis.plot_param_dist()
Advanced Usage
Other Curve Types
BayesCurveFit supports various curve types. Here's an example with the Michaelis-Menten model:
def michaelis_menten_func(x, vamx, km):
return (vamx * x) / (km + x)
x_data = np.array(
[
0.5,
2.0,
4.0,
5.0,
6.0,
8.0,
10.0,
]
)
y_data = np.array(
[
0.3,
0.6,
0.8,
np.nan,
1.2,
0.85,
0.9,
]
)
params_range = [(0.01, 5), (0, 5)]
param_names = ["vmax", "km"]
run_mm = BayesFitModel(
x_data=x_data,
y_data=y_data,
fit_function=michaelis_menten_func,
params_range=params_range,
param_names=param_names,
)
run_mm.get_result()
fit_vmax 1.075649
fit_km 1.599924
std_vmax 0.125697
std_km 0.662758
est_std 0.085374
null_mean 0.77585
rmse 0.146565
pep 0.097757
convergence_warning False
f,ax = plt.subplots()
run_mm.analysis.plot_fitted_curve(ax=ax)
ax.set_xlabel("Substrate Concentration [S] (M)")
ax.set_ylabel("Relative Response (V/V_max)")
Parallel Processing
BayesCurveFit supports parallel execution for batch processing and pipeline integration:
from joblib import Parallel, delayed
import pandas as pd
from bayescurvefit.simulation import Simulation
from bayescurvefit.distribution import GaussianMixturePDF
def run_bayescurvefit(df_input, x_data, output_csv, params_range, param_names = None):
def process_sample(sample):
y_data = df_input.loc[sample].values
bayesfit_run = BayesFitModel(
x_data=x_data,
y_data=y_data,
fit_function=michaelis_menten_func,
params_range=params_range,
param_names=param_names,
run_mcmc=True,
)
return sample, bayesfit_run.get_result()
samples = df_input.index.tolist()
results = Parallel(n_jobs=-1)(
delayed(process_sample)(sample) for sample in samples
)
dict_result = {sample: res for sample, res in results}
df_output = pd.DataFrame(dict_result).T
df_output.to_csv(output_csv)
return df_output
# Generate simulation dataset
sample_size = 6 # Number of observations
params_bounds = [(1, 1.2), [0.5, 1]] # Bounds for fitting parameters in simulation data
pdf_params = [0.1, 0.3, 0.1] # PDF parameters [scale1, scale2, mix_frac of second Gaussian]
x_sample = np.linspace(0.5, 10, sample_size)
sim = Simulation(fit_function=michaelis_menten_func,
X=x_sample,
pdf_function=GaussianMixturePDF(pdf_params=pdf_params),
sim_params_bounds=params_bounds,
n_samples=10,
n_replicates=1)
params_range = [(0.01, 5), (0, 5)] #The range of fitting parameters we estimated
param_names = ["vmax", "km"] #optional
df_output = run_bayescurvefit(df_input = sim.df_sim_data_w_error, x_data=sim.X, output_csv="test_output.csv", params_range= params_range, param_names = param_names)
df_output
| fit_vmax | fit_km | std_vmax | std_km | est_std | null_mean | rmse | pep | convergence_warning | |
|---|---|---|---|---|---|---|---|---|---|
| m001 | 1.05 | 0.53 | 0.07 | 0.27 | 0.09 | 0.88 | 0.09 | 0.08 | False |
| m002 | 1.07 | 0.45 | 0.05 | 0.13 | 0.06 | 0.92 | 0.07 | 0.02 | False |
| m003 | 0.99 | 0.42 | 0.03 | 0.08 | 0.04 | 0.85 | 0.04 | 0.00 | False |
| m004 | 1.20 | 0.83 | 0.12 | 0.50 | 0.11 | 0.95 | 0.12 | 0.15 | False |
| m005 | 1.11 | 0.86 | 0.08 | 0.30 | 0.09 | 0.86 | 0.09 | 0.02 | False |
| m006 | 1.20 | 1.79 | 0.19 | 1.19 | 0.14 | 0.90 | 0.26 | 0.58 | False |
| m007 | 1.05 | 1.49 | 0.26 | 1.16 | 0.20 | 0.67 | 0.33 | 0.92 | False |
| m008 | 1.59 | 2.62 | 0.25 | 1.14 | 0.13 | 0.96 | 0.17 | 0.17 | False |
| m009 | 1.23 | 1.57 | 0.09 | 0.46 | 0.06 | 0.85 | 0.06 | 0.00 | False |
| m010 | 1.11 | 0.41 | 0.04 | 0.10 | 0.05 | 0.96 | 0.05 | 0.00 | False |
Development
Running Tests
# Install development dependencies
uv sync --dev
# Run all tests
make test
# Run specific test file
uv run pytest src/tests/test_utils.py -v
Building the Package
# Build package
make build
# Install locally
make install
Version Management
Git tags are the single source of truth for versioning
To release a new version:
- Run the script:
python scripts/bump_version.py patch(orminor/major)- Updates
__init__.pywith new version - Commits the change
- Creates and pushes git tag (e.g.,
v0.6.2)
- Updates
- Create GitHub release: Go to GitHub Releases, create release with the existing tag
- Done! GitHub Actions automatically:
- Updates
pyproject.tomlwith version from git tag - Builds and publishes to PyPI
- Updates
Available commands:
python scripts/bump_version.py patch # 0.6.1 → 0.6.2
python scripts/bump_version.py minor # 0.6.1 → 0.7.0
python scripts/bump_version.py major # 0.6.1 → 1.0.0
Note: The version in pyproject.toml is automatically updated by GitHub Actions during the release process.
Contributing
- Fork the repository
- Create a feature branch:
git checkout -b feature-name - Make your changes and add tests
- Run tests:
make test - Commit your changes:
git commit -m "Add feature" - Push to the branch:
git push origin feature-name - Submit a pull request
License
This project is licensed under the Apache License 2.0 - see the LICENSE file for details.
Citation
If you use BayesCurveFit in your research, please cite:
@article{du2025bayescurvefit,
title={BayesCurveFit: Enhancing Curve Fitting in Drug Discovery Data Using Bayesian Inference},
author={Du, Niu and others},
journal={bioRxiv},
year={2025},
doi={10.1101/2025.02.23.639769}
}
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file bayescurvefit-0.6.3.tar.gz.
File metadata
- Download URL: bayescurvefit-0.6.3.tar.gz
- Upload date:
- Size: 261.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f196833cf29385ba415af6b562ab31f408f82a9288edc01dda18d00377647fe6
|
|
| MD5 |
3f280ffbb9ea15cd9242fede202926b1
|
|
| BLAKE2b-256 |
bc48656091297d853c2de65d2b2d04081f384fe43e8018c1b6c4cd4e2b595756
|
File details
Details for the file bayescurvefit-0.6.3-py3-none-any.whl.
File metadata
- Download URL: bayescurvefit-0.6.3-py3-none-any.whl
- Upload date:
- Size: 29.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.13.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b9dce527376f342c70e84405edbfd6a972873f745d94248503af52b20e72deec
|
|
| MD5 |
6e183c52add34c89f13739a5aabeafc6
|
|
| BLAKE2b-256 |
efff5127bfeacffd737b226142d1ad53f9522db179e1a55f1aafc424dc2e5a21
|