Skip to main content

b2bTools

b2bTools is the Bio2Byte Python package for predicting biophysical properties from protein sequences and multiple-sequence alignments. It also provides lightweight readers for FASTA, alignment, NEF, and NMR-STAR files.

ELIXIR Belgium infrastructure b2bTools is an ELIXIR Belgium Core Service. Explore the ELIXIR Belgium services.

Installation

Install the package with python -m pip install b2bTools.

Supported Python versions are 3.7 through 3.14 (>=3.7, <3.15).

Optional external tools are required only for particular workflows:

  • HMMER is required for PSPer.
  • T-Coffee is required when b2bTools creates an alignment from two input sequence files.

Two commands are installed: b2bTools for sequence-based predictors, and shiftcrypt for NMR chemical-shift input.

python -m b2bTools --help
python -m b2bTools.nmr.shiftCrypt --help

Choose a workflow

Input Entry point Use it for
One or more unaligned protein sequences SingleSeq Per-residue predictions from FASTA input.
A multiple-sequence alignment MultipleSeq Predictions mapped to an existing alignment.
NMR chemical shifts ShiftCrypt Residue-level biophysical indices from NEF or NMR-STAR data.
File parsing only b2bTools.general.parsers Reading supported formats without importing predictor models.

Quick start

Predict properties for a FASTA file

Create a FASTA file named proteins.fasta, then run DynaMine and write its results as JSON. Add other predictor constants from the reference table below when needed.

from pathlib import Path
import json

from b2bTools import SingleSeq, constants

input_fasta = Path("proteins.fasta")
predictor = SingleSeq(str(input_fasta), short_id=False)
predictor.predict(tools=[constants.TOOL_DYNAMINE])

results = predictor.get_all_predictions()
Path("predictions.json").write_text(json.dumps(results, indent=2), encoding="utf-8")

results["proteins"] maps each sequence identifier to its per-residue values. results["metadata"] records the requested tools and package metadata. Use get_all_predictions_tabular(path, sep=",") for CSV or get_all_predictions_tabular(path, sep="\t") for TSV output.

Predict properties for an existing alignment

Use an existing, aligned FASTA file when the relationship between residues must be preserved across sequences.

from pathlib import Path

from b2bTools import MultipleSeq, constants

input_alignment = Path("alignment.fasta")
predictor = MultipleSeq()
predictor.from_aligned_file(str(input_alignment), tools=[constants.TOOL_DYNAMINE])

results = predictor.get_all_predictions_msa()

results["proteins"] contains the aligned prediction values. To select one sequence, pass its identifier to get_all_predictions_msa(sequence_key).

Predictor reference

Predictor Constant Produces
DynaMine constants.TOOL_DYNAMINE Backbone and side-chain dynamics; helix, sheet, coil, and polyproline-II propensities.
EFoldMine constants.TOOL_EFOLDMINE Early-folding propensity.
DisoMine constants.TOOL_DISOMINE Disorder propensity.
AgMata constants.TOOL_AGMATA Beta-aggregation propensity.
PSPer constants.TOOL_PSP Phase-separation features and a protein-level score.

Dependencies are resolved automatically. For example, requesting PSPer also runs the predictors whose outputs it requires. Runtime depends on the selected tools, sequence lengths, and the number of sequences.

Understanding prediction outputs

The prediction APIs preserve positional correspondence: element n in a per-residue prediction list describes residue_index n. This makes the JSON results suitable for programmatic analysis and the tabular exports suitable for spreadsheets and statistical tools.

JSON result structure

SingleSeq.get_all_predictions() returns proteins and metadata. proteins maps each sequence identifier to its sequence and prediction lists; metadata records the requested tools and package information. For an existing alignment, MultipleSeq.get_all_predictions_msa() returns proteins together with sequences, which retains the aligned sequence (including gap positions) for every identifier.

Every numeric prediction key contains one value per residue or aligned position. viterbi is the exception: it contains PSPer's categorical state labels. A value can be null when a property was not calculated for that position, for example at an alignment gap.

Per-residue CSV and TSV columns

get_all_predictions_tabular(path, sep=",") writes a CSV file; pass sep="\t" for TSV. Each row represents one sequence position. Numeric values are rounded to three decimal places in this export.

Column Meaning Type in the table
sequence_id Identifier of the input sequence. Text
residue One-letter residue code. In an MSA export, this can be the gap character -. Text
residue_index Zero-based position in the sequence or alignment. Integer
Prediction key Value for the property described below. A blank cell means that the predictor was not selected or no value is available at that position. Numeric, except viterbi

The following keys are included as prediction columns. Their JSON values are ordered lists; in the CSV or TSV, each list is expanded into its corresponding residue row.

Predictor Output keys Value type
DynaMine backbone, sidechain, ppII, coil, sheet, helix Numeric per residue
EFoldMine earlyFolding Numeric per residue
DisoMine disoMine Numeric per residue
AgMata agmata Numeric per residue
PSPer complexity, arg, tyr, RRM, disorder Numeric per residue
PSPer viterbi Categorical state per residue

PSPer also supplies protein_score, a single score for the complete protein. It is available in the JSON result and metadata statistics rather than as a per-residue column.

MSA distribution outputs

MultipleSeq.get_all_predictions_msa_distrib() summarises the numeric predictions across sequences at every alignment position. Its results value is organised as prediction_key → summary_name → ordered values. Use get_msa_distrib_tabular(path, sep=",") to write the same data as columns named prediction_key_summary_name, alongside the zero-based residue_index.

Summary name Meaning at each alignment position
median Median of available numeric values.
firstQuartile 25th percentile.
thirdQuartile 75th percentile.
bottomOutlier Lower outlier boundary: first quartile minus 1.5 times the interquartile range.
topOutlier Upper outlier boundary: third quartile plus 1.5 times the interquartile range.

viterbi is categorical, so it is intentionally excluded from MSA distribution statistics. When an alignment column has no numeric values, all of its summary values are null (or blank in the tabular export).

File-parser reference

The parser modules are safe to import in services that do not need PyTorch or other predictor-model dependencies.

FASTA

FastaIO.read_fasta_from_file and FastaIO.read_fasta_from_string return a list of (sequence_id, sequence) pairs. By default, short_id=False retains the complete header, collapses whitespace to _, and sanitises spaces, full stops, vertical bars, and commas. Pass short_id=True to use the first header token, capped at 20 characters.

from pathlib import Path

from b2bTools.general.parsers.fasta import FastaIO

records = FastaIO.read_fasta_from_file(Path("proteins.fasta"), short_id=False)
for sequence_id, sequence in records:
    print(sequence_id, len(sequence))

Alignments

AlignmentsIO reads FASTA, A3M, BLAST, BaliBase, CLUSTAL, PSI, PHYLIP, and Stockholm alignments. Each read_alignments* method returns a dictionary that maps sequence identifiers to aligned sequences.

short_id=False is the default. Where the format provides a complete header, the parser retains it with whitespace collapsed to _; short_id=True uses the first header token capped at 20 characters. Write helpers and NMR readers do not take short_id because they do not derive sequence identifiers from FASTA-style headers.

from pathlib import Path

from b2bTools.general.parsers.alignments import AlignmentsIO

alignment = AlignmentsIO.read_alignments("alignment.fasta", short_id=False)
for sequence_id, aligned_sequence in alignment.items():
    print(sequence_id, len(aligned_sequence))

NEF and NMR-STAR

NefIO reads NEF projects and sequence/chemical-shift data. NMRStarIO provides the equivalent NMR-STAR readers. Neither API uses short_id.

from pathlib import Path

from b2bTools.general.parsers.nef import NefIO

sequence_shifts = NefIO.read_nef_file_sequence_shifts(Path("example.nef"))
print(sequence_shifts.keys())

ShiftCrypt

Use ShiftCrypt when your input contains NMR chemical shifts. It accepts either of the following formats, provided the file includes a polymer sequence and assigned chemical shifts:

  • NEF (NMR Exchange Format): use parse_official(path), which is the default input mode.
  • NMR-STAR: use parse_official(path, is_star=True) for BMRB archive entries and other NMR-STAR chemical-shift files.

modelClass="2" is the default reduced-atom model (H, HA, CA, N, CB, and C). Use model "1" for the full available atom set, or model "3" for the dimer-sensitive CA, N, and H model.

from pathlib import Path

from b2bTools.nmr.shiftCrypt.Predictor import ShiftCrypt
from b2bTools.nmr.shiftCrypt.shiftcrypt_pkg.parser import parse_official

protein_shifts = parse_official(Path("example.nef"))
results = ShiftCrypt().predictShifts(protein_shifts, modelClass="2")

For NMR-STAR input, pass is_star=True while parsing, for example parse_official(Path("entry.str"), is_star=True).

predictShifts returns one dictionary per parsed chain. Its fields are:

Field Meaning Type
ID_file Identifier supplied by the input file. Text
sequence Residue sequence. List of one-letter residue codes
seqCodes Residue numbering from the source data. List of integers
shiftCrypt Residue-level ShiftCrypt values. List of numeric values
chainCode Chain identifier from the source data. Text

Command-line interface

The CLI processes a FASTA file in single-sequence mode by default. Use --mode msa for an existing alignment; add predictor flags such as --disomine or --psper to request additional outputs. --short_ids selects the 20-character identifier mode.

Required options are --input_file and --output_json_file. Optional --output_tabular_file, --metadata_file, and, for MSA runs, --distribution_json_file and --distribution_tabular_file write additional outputs. Use --help for the authoritative option list in the installed version.

The shiftcrypt command

ShiftCrypt reads NMR chemical shifts rather than sequences, so it has its own command instead of a flag on b2bTools:

shiftcrypt --input_file example.nef --output_json_file shiftcrypt.json

The input format is inferred from the file extension (.nef for NEF, .str and .bmrb for NMR-STAR). Use --format nef or --format nmr_star when the extension is absent or misleading.

--model selects the encoding: 2 (default) for the reduced-atom model, 1 for the full atom set, 3 for the dimer-sensitive CA/N/H model, or a path to a custom .mtorch model.

Output is written to any combination of --output_json_file, --output_tabular_file (with --sep comma or --sep tab), and --output_text_file (a whitespace-aligned chain, residue number, residue, value layout). When none of them is given, the plain-text layout is printed to standard output. The JSON file holds the same per-chain records that predictShifts returns, documented in the ShiftCrypt section above.

--fit is reserved for training custom models and is not available in this distribution; the command explains what is missing and exits.

Development

From the repository root, run make test to execute the maintained test suite. Run make generate-docs to generate API documentation in wrapper_documentation.

Further documentation and citations

The Bio2Byte package documentation contains detailed API material and method-level documentation. The Bio2Byte tools page describes the scientific background of each predictor.

If you use b2bTools in published work, cite the relevant predictor:

Predictor Citation
DynaMine Cilia et al. (2013), Nature Communications 4:2741. DOI
DisoMine Orlando et al. (2022), Journal of Molecular Biology. DOI
EFoldMine Raimondi et al. (2017), Scientific Reports 7:8826. DOI
AgMata Orlando et al. (2020), Bioinformatics 36:2076–2081. DOI
PSPer Orlando et al. (2019), Bioinformatics 35:4617–4623. DOI
ShiftCrypt Orlando et al. (2020), Nucleic Acids Research 48:W36–W40. DOI

Licence and terms of use

b2bTools is distributed under the GNU General Public License v3.0 (GPLv3).

Bio2Byte promotes open science by providing freely available online services, databases, and software for the life sciences, with a focus on proteins. Please attribute Bio2Byte services, databases, and software in publications, services, or products according to good scientific practice and the relevant citation guidance above. Bio2Byte is not liable for loss or damage arising from the use of this software.

Questions about these terms may be sent to bio2byte@vub.be.

Metadata

Release files for b2bTools 3.0.9

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for b2bTools 3.0.9
File Size Uploaded
b2btools-3.0.9.tar.gz 19.9 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for b2bTools 3.0.9
File Interpreter ABI Platform
b2btools-3.0.9-py3-none-any.whl Python 3 none any Details

Total release size: 40.1 MB

Release files / b2btools-3.0.9.tar.gz

Download URL b2btools-3.0.9.tar.gz
Size 19.9 MB
Tags Source
SHA-256 checksum
How to use checksums
afc574a6396ce3d02c3b3a425cd3a60aa23633a60784f7cbf024708ebd34bb6f
BLAKE2b-256 checksum
How to use checksums
e6d2152adc543971238e1847f958f011b6dc8d52743d184f4100d75b25f1e0a0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.12.13

Release files / b2btools-3.0.9-py3-none-any.whl

Download URL b2btools-3.0.9-py3-none-any.whl
Size 20.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
2e20004e3b9a74823ed0f52284e37c7e2ec4963a7c1d087197b883c2ea29f3b1
BLAKE2b-256 checksum
How to use checksums
aff8be8f6afc6093b0fe5dffe4b2d7ebf6e2f1585edccd99f254918eff9ad714
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.1 CPython/3.12.13

Release history Release notifications | RSS feed

This release

3.0.9 This release

2 release files

3.0.8

2 release files

3.0.7

2 release files

3.0.6

2 release files

3.0.5

2 release files

3.0.4

2 release files

3.0.3

2 release files

3.0.2

2 release files

3.0.1

2 release files

3.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page