Skip to main content

ML runtime (https://pypi.org/project/fdq/)

Project description

FDQ | Fonduecaquelon

Fonduecaquelon (FDQ) is designed for researchers and practitioners who want to focus on deep learning experiments, not boilerplate code. FDQ streamlines your PyTorch workflow, automating repetitive tasks and providing a flexible, extensible framework for experiment management—so you can spend more time on innovation and less on setup.


🚀 Features

  • Minimal Boilerplate: Define only what matters — FDQ handles the rest.
  • Flexible Experiment Configuration: Use JSON config files with inheritance support for easy experiment management.
  • Multi-Model Support: Seamlessly manage multiple models, losses, and data loaders.
  • Cluster Ready: Effortlessly submit jobs to SLURM clusters with built-in utilities.
  • Extensible: Easily integrate custom models, data loaders, and training/testing loops.
  • Automatic Dependency Management: Install additional pip packages per experiment.
  • Distributed Training: Out-of-the-box support for distributed training using PyTorch DDP.
  • Model Export & Optimization: Export trained models to ONNX format with optimization options.
  • High-Performance Inference: TensorRT integration for GPU-accelerated inference with up to 10x speedup.
  • Model Compilation: JIT tracing/scripting and torch.compile support for optimized model execution.
  • Interactive Model Dumping: Easy-to-use interface for exporting and optimizing trained models.

🛠️ Installation

Install the latest release from PyPI:

pip install fdq

Or, for development and the latest features, clone the repository:

git clone https://github.com/mstadelmann/fonduecaquelon.git
cd fonduecaquelon
pip install -e .

📖 Usage

Local Experiments

All experiment parameters are defined in a config file. Config files can inherit from a parent file for easy reuse and organization.

To run an experiment locally:

fdq <path_to_config_file.json>

SLURM Cluster Execution

To run experiments on a SLURM cluster, add a slurm_cluster section to your config. See this example.

Submit your experiment to the cluster:

python <path_to>/fdq_submit.py <path_to_config_file.json>

Model Export and Optimization

After training, export and optimize your models for deployment:

# Interactive model dumping with export options
fdq <path_to_config_file.json> -nt -d

This launches an interactive interface where you can:

  • Export to ONNX: Convert PyTorch models to ONNX format with Dynamo or TorchScript
  • JIT Compilation: Trace or script models using PyTorch JIT
  • TensorRT Optimization: Compile models for GPU inference with precision options (FP32, FP16, INT8)
  • Performance Benchmarking: Compare optimized vs. original model performance

Additional CLI Options

FDQ supports several command-line options for different workflows:

# Run only training (default)
fdq <config_file.json>

# Skip training to run other tasks
fdq <config_file.json> -nt

# Run training and automatic testing
fdq <config_file.json> -ta

# Interactive testing
fdq <config_file.json> -nt -ti

# Export and optimize models
fdq <config_file.json> -nt -d

# Run inference tests on trained models
fdq <config_file.json> -nt -i

# Print model architecture before training
fdq <config_file.json> -p

# Resume training from checkpoint
fdq <config_file.json> -rp /path/to/checkpoint

🚄 Model Export & Deployment

FDQ provides comprehensive model export and optimization capabilities for deployment:

Export Options

  • ONNX Export: Convert models to ONNX format for cross-platform deployment

    • Dynamo-based export for latest PyTorch features
    • TorchScript export for broader compatibility
    • Automatic model optimization and file size reporting
  • JIT Compilation: PyTorch JIT tracing and scripting for optimized execution

    • Trace models for static computation graphs
    • Script models to preserve control flow
    • Automatic performance comparison with original models
  • TensorRT Integration: GPU-accelerated inference with NVIDIA TensorRT

    • FP32, FP16, and INT8 precision modes
    • Automatic engine building and caching

Performance Features

  • Automatic Benchmarking: Built-in performance testing with statistical analysis
  • Memory Optimization: Dynamic batch sizing and memory-efficient engine building
  • Cross-Platform: Works on various GPU architectures and CUDA versions

⚙️ Configuration Overview

FDQ uses JSON configuration files to define experiments. These files specify models, data loaders, training/testing scripts, and cluster settings.

Models

Models are defined as dictionaries. You can use pre-installed models (e.g., Chuchichaestli) or your own. Example:

"models": {
    "ccUNET": {
        "class_name": "chuchichaestli.models.unet.unet.UNet"
    }
}

Access models in your training loop via experiment.models["ccUNET"]. The same structure applies to losses and data loaders.

Data Loaders

Your data loader class must implement a create_datasets(experiment, args) function, returning a dictionary like:

return {
    "train_data_loader": train_loader,
    "val_data_loader": val_loader,
    "test_data_loader": test_loader,
    "n_train_samples": n_train,
    "n_val_samples": n_val,
    "n_test_samples": n_test,
    "n_train_batches": len(train_loader),
    "n_val_batches": len(val_loader) if val_loader is not None else 0,
    "n_test_batches": len(test_loader),
}

These values are accessible from your training loop as experiment.data["<name>"].<key>.

Training Loop

Specify the path to your training script in the config. FDQ expects a function:

def fdq_train(experiment: fdqExperiment):

Within this function, you can access all experiment components:

nb_epochs = experiment.exp_def.train.args.epochs
data_loader = experiment.data["OXPET"].train_data_loader
model = experiment.models["ccUNET"]

See train_oxpets.py for a full example.

Testing Loop

Testing works similarly. Define a function:

def fdq_test(experiment: fdqExperiment):

See oxpets_test.py for an example.


📦 Installing Additional Python Packages in your managed SLURM Environment

If your experiment requires extra Python packages, specify them in your config under additional_pip_packages. FDQ will install them automatically before running your experiment.

Example:

"slurm_cluster": {
    "fdq_version": "0.0.51",
    "...": "...",
    "additional_pip_packages": [
        "monai==1.4.0",
        "prettytable"
    ]
}

🐛 Debugging

To debug an FDQ experiment, you'll need to install FDQ in development mode on your local or remote machine.

Setup for Debugging

git clone https://github.com/mstadelmann/fonduecaquelon.git
cd fonduecaquelon
pip install -e .

VS Code Debugging Configuration

  1. Open your project in VS Code
  2. Create or update your debugger configuration (.vscode/launch.json) to launch run_experiment.py with the corresponding parameters:
{
    "version": "0.2.0",
    "configurations": [
        {
            "name": "FDQ Experiment Debug",
            "type": "debugpy",
            "request": "launch",
            "debugJustMyCode": false,
            "program": "${workspaceFolder}/src/fdq/run_experiment.py",
            "console": "integratedTerminal",
            "args": ["PATH_TO/experiment.json"],
            "cwd": "${workspaceFolder}"
        }
    ]
}
  1. Debug/test your code.

📝 Tips

  • Config Inheritance: Use the parent key in your config to inherit settings from another file, reducing duplication.
  • Multiple Models/Losses: FDQ supports multiple models and losses per experiment — just add them to the config dictionaries.
  • Cluster Submission: The fdq_submit.py utility handles SLURM job script generation and submission, including environment setup and result copying.
  • Model Export: Use -d or --dump to interactively export and optimize trained models for deployment.

📚 Resources


🤝 Contributing

Contributions are welcome! Please open issues or pull requests on GitHub.


🧀 Enjoy your fondue and happy experimenting!

Project details


Release history Release notifications | RSS feed

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

fdq-0.0.58.tar.gz (86.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

fdq-0.0.58-py3-none-any.whl (68.7 kB view details)

Uploaded Python 3

File details

Details for the file fdq-0.0.58.tar.gz.

File metadata

  • Download URL: fdq-0.0.58.tar.gz
  • Upload date:
  • Size: 86.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for fdq-0.0.58.tar.gz
Algorithm Hash digest
SHA256 b52d304c70113e17885568d2a973debe2b86e65dfc8fc21553942bbb21bb53de
MD5 bc4f74b1e2f67e5b441635aa12532c52
BLAKE2b-256 f743e196303694eacd932524d2f0c1c90ba89804a332e00d34f239e60f5f9930

See more details on using hashes here.

File details

Details for the file fdq-0.0.58-py3-none-any.whl.

File metadata

  • Download URL: fdq-0.0.58-py3-none-any.whl
  • Upload date:
  • Size: 68.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for fdq-0.0.58-py3-none-any.whl
Algorithm Hash digest
SHA256 d41a0d9f058fe8bb427421945da85e4f19e697651f73c33cc8fbdac64c0abf67
MD5 87a6df2d1229b9fa3b1c73f2d2db90e4
BLAKE2b-256 19c6dbb80cefd0f94d4ff48b9df1eda9f94d40e1118baa403652bd8a380b376d

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page