Skip to main content

A collection of neural network utilities

Project description

NNModuleTools

A collection of neural network utilities.

Contents

  1. npz compare visualizer

    This tool can compare two npz archives visually.

  2. float utils

    This tool can inspect the memory of float numbers. It also provides simple arithmetic operations between them.

  3. module debugger

    This tool can save all the inputs and outputs of submodules in forward and backward pass. You can also manually save your desired tensors.

Installation

python -m pip install nnmoduletools

Usage

  1. npz compare visualizer

    See tutorial.ipynb

    Now you can also use command line to compare two npz files and generate a report:

    python -m nnmoduletools.comparer target.npz ref.npz --tolerance 0.99,0.99 --verbose 3 --output_dir compare_report --output_fn compare_report.md
    

    See python -m nnmoduletools.comparer --help for more information.

    You can also use Reporter to generate your own report:

    from nnmoduletools.comparer import Reporter
    with Reporter("path/to/output/dir", "report.md"):
        comparer = nnmoduletools.NPZComparer("path/to/target.npz", "path/to/ref.npz")
        print("# Compare Report: tensor")
        comparer.plot_vs_auto(tensor="tensor", save_fig=True, save_dir="subdir")
        comparer.dump_vs_plot(top_k=20)
    

    You will have your report in path/to/output/dir/report.md and the plots in path/to/output/dir/subdir/.

  2. float utils

    See tutorial.ipynb

  3. module debugger

    Typical usage: run the same module using two different devices, dump the tensors into npz file and compare them in npz compare visualizer.

    You may have to set environment variables:

    export DBG_DEVICE=cuda # suppose you are using cuda as the device; if you are using deepspeed, device name will be automatically get from deepspeed.get_accelerator()
    export DBG_SAVE_ALL=1 # by default no tensors are saved; you should export this to enable saving
    

    Insert these lines into your code:

    import torch
    from nnmoduletools.module_debugger import register_hook, save_tensors, save_model_params, save_model_grads, combine_npz
    
    class YourModule(torch.nn.Module):
        ...
        # Your module here
    
    model = YourModule()
    model.apply(register_hook) # <=== apply the hook to print log and save input output tensors
    ...
    save_model_params(0) # <=== save the model params before training
    # in your training loop
    for step in range(1, total_steps+1):
        output = model.forward(input)
        loss = loss_function(output, target)
        loss.backward()
        combine_npz(step) # <=== combine input and output tensors into large npzs
        save_model_grads(step) # <=== save the model grads after backward pass
        optimizer.step()
        save_model_params(step) # <=== save the model params after optim update
        save_tensors(tensor_to_save, name_to_save, dir_to_save, save_grad_instead) # <=== save the tensor you want to given directory. You can save grad instead by passing save_grad_instead=True
    

    The log and npz files will be saved in a directory named like logs_2024.06.04.16.52.12/

    You can get the latest log directory easily with LogReader:

    from nnmoduletools.module_debugger import LogReader
    latest = LogReader(devices=["cuda"])
    print(latest.cuda_dir)
    

Troubleshooting

You may encounter such errors when using backward hooks:

RuntimeError: Output 0 of BackwardHookFunctionBackward is a view and is being modified inplace. This view was created inside a custom Function (or because an input was returned as-is) and the autograd logic to handle view+inplace would override the custom backward associated with the custom Function, leading to incorrect gradients. This behavior is forbidden. You can fix this by cloning the output of the custom Function.

It is likely because that your module has an inplace operation right before return in forward, such as

class YourModule(torch.nn.Module):
    ...
    def forward(self, x):
        result = ...
        result += 1 # <===
        return result

You can replace the operation with

    def forward(self, x):
        result = ...
        result = result + 1 # <===
        return result

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nnmoduletools-0.0.13a0.tar.gz (22.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nnmoduletools-0.0.13a0-py3-none-any.whl (22.4 kB view details)

Uploaded Python 3

File details

Details for the file nnmoduletools-0.0.13a0.tar.gz.

File metadata

  • Download URL: nnmoduletools-0.0.13a0.tar.gz
  • Upload date:
  • Size: 22.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.9.19

File hashes

Hashes for nnmoduletools-0.0.13a0.tar.gz
Algorithm Hash digest
SHA256 e3729a97d5167444408211cee5e07a8855bfb2d3bb6f3985c2b83f48b0b0a0b9
MD5 e462349cab6d641c962a077e0ebf26f9
BLAKE2b-256 fa776b5aca31f6912105659114f99cf641d2da8a049f76ba512f8c31e8595756

See more details on using hashes here.

File details

Details for the file nnmoduletools-0.0.13a0-py3-none-any.whl.

File metadata

File hashes

Hashes for nnmoduletools-0.0.13a0-py3-none-any.whl
Algorithm Hash digest
SHA256 3b85692abec468f0167b937ae47e9686ac1dc78c89b384dc5a2e60832fe55392
MD5 535cb4e9f69aca98636633406a935ac1
BLAKE2b-256 6024040122f5c400d59a0f47c5f156e0526843ff5bd88e351e468958d286a05f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page