Skip to main content

Fat Llama Logo

Fat Llama build - status PyPI PyPI - Downloads

fat_llama is a Python package for upscaling audio files to FLAC or WAV formats using advanced audio processing techniques. It utilizes cpu-accelerated calculations to enhance audio quality by upsampling and adding missing frequencies through FFT (Fast Fourier Transform), resulting in richer and more detailed audio.

Features

  • Upscale MP3 files to high-quality FLAC format.
  • Iterative soft thresholding (IST) for enhanced audio processing.
  • Auto-scaling amplitude adjustment and normalization.
  • Multi-Threaded processing on cpu.

Installation

Install via pip:

pip install fat-llama-fftw

(Note: For CUDA verison please look at https://pypi.org/project/fat-llama/)

Also, requires ffmpeg: https://support.audacityteam.org/basics/installing-ffmpeg

Usage

Example Usage

You can run the example provided in example.py:

from fat_llama_fftw.audio_fattener.feed import upscale

# Example call to the method
upscale(
    input_file_path='input_test.mp3',
    output_file_path='output_test.flac',
    source_format='mp3',
    target_format='flac',
    max_iterations=1000,
    threshold_value=0.6,
    target_bitrate_kbps=1400
)

Function Parameters

  • input_file_path (str): Path to the input audio file. Mandatory.
  • output_file_path (str): Path to the output processed audio file. Mandatory.
  • source_format (str): Format of the input audio file (e.g., 'mp3', 'wav', 'ogg', 'flac').
  • target_format (str): Format of the output audio file (e.g., 'flac', 'wav'). Default is 'flac'.
  • max_iterations (int): Maximum number of iterations for IST. Default is 800.
  • threshold_value (float): Threshold value for IST. Default is 0.6.
  • target_bitrate_kbps (int): Target bitrate in kbps. Default is 1411.

Running the Example

To run the example, execute the following command:

python example.py

This will upscale the MP3 file specified in the example and produce a FLAC file with full processing.

Spectrogram Results

Spectrogram Results

Audio Quality Scores

Generated by the test-fat-llama skill's audio-quality-checker subagent — updated each run, not hand-edited.

Metric Score Notes
Coherence (upscale quality, 0-10) 9.8 every defect check clean and mostly beats the reference (peak exactly 1.0, 0 dropouts/NaN/Inf, no hop-locked artifact, above-Nyquist energy -134.4dB) plus genuine added detail below the original Nyquist (+0.8 to +1.8dB at 12-19kHz, concentrated in previously-quiet bands, noise floor down not up); held under 10 only by a residual -0.19 to -0.41dB net level reduction 0-8kHz (the documented IST peak-tax residue)
Spectral deviation vs. reference FLAC (0-10) 9.8 convergence=0.9685, correlation=0.9998 against input_test.flac; 25% of the residual traces to content above 22050Hz where the reference (a legacy zero-order-hold-era output) still carries imaging that this pipeline correctly excludes - restricting the same metric to <=22050Hz gives 9.9, so this is not a regression

How it works

How it Works

Algorithm Explanation

The upscaling process involves several steps:

  1. Reading Audio File: The audio file is read, and the audio samples are extracted along with the sample rate and bitrate.
  2. Calculating Upscale Factor: The upscale factor is calculated to achieve the target bitrate.
  3. Upscaling Channels: The audio channels are upscaled using an interpolation algorithm. Each sample is repeated multiple times to increase the resolution.
  4. Iterative Soft Thresholding (IST): IST is applied to enhance the audio by adding missing frequencies. This process uses FFT to transform the signal into the frequency domain, apply a threshold to keep significant frequencies, and then inverse transform back to the time domain.
  5. Scaling Amplitude: The amplitude of the upscaled audio is scaled to match the original.
  6. Normalizing Audio: The audio is normalized to the range -1 to 1.
  7. Writing FLAC File: The processed audio is written to a FLAC file.

Why FFT and IST?

FFT (Fast Fourier Transform) is used to transform the audio signal into the frequency domain. This allows for the identification and manipulation of specific frequency components. By applying a threshold in the frequency domain, we can keep significant frequencies and discard noise and add it to our upscaling data to add detail to upscaling frequencies.

The report titled "Fast Sparse Fourier Transformations for NMR Spectroscopy" by Badruddin Kamal, supervised by Thomas Huber and Alastair Rendall, 2015, provides a comprehensive understanding of sparse representations and their applications in signal processing. IST leverages the concepts from this report to add missing frequencies and enhance the audio quality by making it more detailed and rich. This is particularly useful in upscaling audio where some frequencies might be missing or congested.

Test Audio Source

ericzo - beyond link(https://soundcloud.com/ericzomusic/free-electro-trap-anthem-beyond)

Changelog

Changes are now logged in CHANGELOG.md

Release files for fat-llama-fftw 1.4.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for fat-llama-fftw 1.4.3
File Size Uploaded
fat_llama_fftw-1.4.3.tar.gz 46.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for fat-llama-fftw 1.4.3
File Interpreter ABI Platform
fat_llama_fftw-1.4.3-py3-none-any.whl Python 3 none any Details

Total release size: 92.6 kB

Release files / fat_llama_fftw-1.4.3.tar.gz

Download URL fat_llama_fftw-1.4.3.tar.gz
Size 46.9 kB
Tags Source
SHA-256 checksum
How to use checksums
b7d857c95624b32af5e8473fe54eccc36fc27ac9e3b673505a8e606728497304
BLAKE2b-256 checksum
How to use checksums
99d99531b69ed3e72cf12a15d75e9ed90098e0aa2214a417e40d1516456089e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release files / fat_llama_fftw-1.4.3-py3-none-any.whl

Download URL fat_llama_fftw-1.4.3-py3-none-any.whl
Size 45.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2808861fde3542e2a8a94e0b9d0cb1407922a5a7506c4cc34e2b5fde6b92dc60
BLAKE2b-256 checksum
How to use checksums
9c656799240193ac803ca9c52401be51ca5fe9701588cb5b09d33c53e6c634c5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.14

Release history Release notifications | RSS feed

2.0.2

2 release files

1.4.5

2 release files

This release

1.4.3 This release

2 release files

1.4.2

2 release files

1.4.1

2 release files

1.0.4

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page