Skip to main content

A collection of neural network utilities

Project description

NNModuleTools

A collection of neural network utilities.

Contents

  1. npz compare visualizer

    This tool can compare two npz archives visually.

  2. float utils

    This tool can inspect the memory of float numbers. It also provides simple arithmetic operations between them.

  3. module debugger

    This tool can save all the inputs and outputs of submodules in forward and backward pass. You can also manually save your desired tensors.

Installation

python -m pip install nnmoduletools

Usage

  1. npz compare visualizer

    See tutorial.ipynb

    Now you can also use command line to compare two npz files and generate a report:

    python -m nnmoduletools.comparer [--tolerance 0.99,0.99] [--abs_tol 1e-8] [--rel-tol 1e-3] [--verbose 3] [--output_dir compare_report] [--output_fn compare_report.md] [--info] [--dump [10]] target.npz ref.npz [tensor1 tensor2 ...] 
    

    See python -m nnmoduletools.comparer --help for more information.

    The command line tool only support basic operations. You can also use Reporter in Python scripts to generate your own report:

    from nnmoduletools.comparer import Reporter
    with Reporter("path/to/output/dir", "report.md"):
        comparer = nnmoduletools.NPZComparer("path/to/target.npz", "path/to/ref.npz")
        print("# Compare Report: tensor")
        comparer.plot_vs_auto(tensor="tensor", save_fig=True, save_dir="subdir")
        comparer.dump_vs_plot(top_k=20)
    

    You will have your report in path/to/output/dir/report.md and the plots in path/to/output/dir/subdir/.

  2. float utils

    See tutorial.ipynb

  3. module debugger

    Typical usage: run the same module using two different devices, dump the tensors into npz file and compare them in npz compare visualizer.

    You may have to set environment variables:

    export DBG_DEVICE=cuda # suppose you are using cuda as the device; if you are using deepspeed, device name will be automatically get from deepspeed.get_accelerator()
    export DBG_SAVE_ALL=1 # by default no tensors are saved; you should export this to enable saving
    

    Insert these lines into your code:

    import torch
    from nnmoduletools.module_debugger import register_hook, save_tensors, save_model_params, save_model_grads, combine_npz
    
    class YourModule(torch.nn.Module):
        ...
        # Your module here
    
    model = YourModule()
    model.apply(register_hook) # <=== apply the hook to print log and save input output tensors
    ...
    save_model_params(0) # <=== save the model params before training
    # in your training loop
    for step in range(1, total_steps+1):
        output = model.forward(input)
        loss = loss_function(output, target)
        loss.backward()
        combine_npz(step) # <=== combine input and output tensors into large npzs
        save_model_grads(step) # <=== save the model grads after backward pass
        optimizer.step()
        save_model_params(step) # <=== save the model params after optim update
        save_tensors(tensor_to_save, name_to_save, dir_to_save, save_grad_instead) # <=== save the tensor you want to given directory. You can save grad instead by passing save_grad_instead=True
    

    The log and npz files will be saved in a directory named like logs_2024.06.04.16.52.12/

    You can get the latest log directory easily with LogReader:

    from nnmoduletools.module_debugger import LogReader
    latest = LogReader(devices=["cuda"])
    print(latest.cuda_dir)
    

Troubleshooting

You may encounter such errors when using backward hooks:

RuntimeError: Output 0 of BackwardHookFunctionBackward is a view and is being modified inplace. This view was created inside a custom Function (or because an input was returned as-is) and the autograd logic to handle view+inplace would override the custom backward associated with the custom Function, leading to incorrect gradients. This behavior is forbidden. You can fix this by cloning the output of the custom Function.

It is likely because that your module has an inplace operation right before return in forward, such as

class YourModule(torch.nn.Module):
    ...
    def forward(self, x):
        result = ...
        result += 1 # <===
        return result

You can replace the operation with

    def forward(self, x):
        result = ...
        result = result + 1 # <===
        return result

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nnmoduletools-0.0.20a0.tar.gz (23.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nnmoduletools-0.0.20a0-py3-none-any.whl (23.4 kB view details)

Uploaded Python 3

File details

Details for the file nnmoduletools-0.0.20a0.tar.gz.

File metadata

  • Download URL: nnmoduletools-0.0.20a0.tar.gz
  • Upload date:
  • Size: 23.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.9.19

File hashes

Hashes for nnmoduletools-0.0.20a0.tar.gz
Algorithm Hash digest
SHA256 5fb4c30e45fb72e601e5b0636c5fe7c5c7b3c6590787099b793896b814f8c645
MD5 e7d487efa33588d31dda4cce6885767d
BLAKE2b-256 fee7f02efd5d3a6e70622edebc013824637815692bc3ca752b7d53564c5c140c

See more details on using hashes here.

File details

Details for the file nnmoduletools-0.0.20a0-py3-none-any.whl.

File metadata

File hashes

Hashes for nnmoduletools-0.0.20a0-py3-none-any.whl
Algorithm Hash digest
SHA256 55b32abc6cd3528fd4d164b3528a7200ce22cc9de035d7744cdc21213f3e1276
MD5 77716830a04608725f66b9de233a33cd
BLAKE2b-256 96ba35528764978e21b1134314162a8f66cb44437dba8b0d1a7c6ef940062de1

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page