Skip to main content

A collection of neural network utilities

Project description

NNModuleTools

A collection of neural network utilities.

Contents

  1. npz compare visualizer

    This tool can compare two npz archives visually.

  2. float utils

    This tool can inspect the memory of float numbers. It also provides simple arithmetic operations between them.

  3. module debugger

    This tool can save all the inputs and outputs of submodules in forward and backward pass. You can also manually save your desired tensors.

Installation

python -m pip install nnmoduletools

Usage

  1. npz compare visualizer

    See tutorial.ipynb

    Now you can also use command line to compare two npz files and generate a report:

    python -m nnmoduletools.comparer [--tolerance 0.99,0.99] [--abs_tol 1e-8] [--rel-tol 1e-3] [--verbose 3] [--output_dir compare_report] [--output_fn compare_report.md] [--info] [--dump [10]] target.npz ref.npz [tensor1 tensor2 ...] 
    

    See python -m nnmoduletools.comparer --help for more information.

    The command line tool only support basic operations. You can also use Reporter in Python scripts to generate your own report:

    from nnmoduletools.comparer import Reporter
    with Reporter("path/to/output/dir", "report.md"):
        comparer = nnmoduletools.NPZComparer("path/to/target.npz", "path/to/ref.npz")
        print("# Compare Report: tensor")
        comparer.plot_vs_auto(tensor="tensor", save_fig=True, save_dir="subdir")
        comparer.dump_vs_plot(top_k=20)
    

    You will have your report in path/to/output/dir/report.md and the plots in path/to/output/dir/subdir/.

  2. float utils

    See tutorial.ipynb

  3. module debugger

    Typical usage: run the same module using two different devices, dump the tensors into npz file and compare them in npz compare visualizer.

    You may have to set environment variables:

    export DBG_DEVICE=cuda # suppose you are using cuda as the device; if you are using deepspeed, device name will be automatically get from deepspeed.get_accelerator()
    export DBG_SAVE_ALL=1 # by default no tensors are saved; you should export this to enable saving
    

    Insert these lines into your code:

    import torch
    from nnmoduletools.module_debugger import register_hook, save_tensors, save_model_params, save_model_grads, combine_npz
    
    class YourModule(torch.nn.Module):
        ...
        # Your module here
    
    model = YourModule()
    model.apply(register_hook) # <=== apply the hook to print log and save input output tensors
    ...
    save_model_params(0) # <=== save the model params before training
    # in your training loop
    for step in range(1, total_steps+1):
        output = model.forward(input)
        loss = loss_function(output, target)
        loss.backward()
        combine_npz(step) # <=== combine input and output tensors into large npzs
        save_model_grads(step) # <=== save the model grads after backward pass
        optimizer.step()
        save_model_params(step) # <=== save the model params after optim update
        save_tensors(tensor_to_save, name_to_save, dir_to_save, save_grad_instead) # <=== save the tensor you want to given directory. You can save grad instead by passing save_grad_instead=True
    

    The log and npz files will be saved in a directory named like logs_2024.06.04.16.52.12/

    You can get the latest log directory easily with LogReader:

    from nnmoduletools.module_debugger import LogReader
    latest = LogReader(devices=["cuda"])
    print(latest.cuda_dir)
    

Troubleshooting

You may encounter such errors when using backward hooks:

RuntimeError: Output 0 of BackwardHookFunctionBackward is a view and is being modified inplace. This view was created inside a custom Function (or because an input was returned as-is) and the autograd logic to handle view+inplace would override the custom backward associated with the custom Function, leading to incorrect gradients. This behavior is forbidden. You can fix this by cloning the output of the custom Function.

It is likely because that your module has an inplace operation right before return in forward, such as

class YourModule(torch.nn.Module):
    ...
    def forward(self, x):
        result = ...
        result += 1 # <===
        return result

You can replace the operation with

    def forward(self, x):
        result = ...
        result = result + 1 # <===
        return result

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nnmoduletools-0.0.25a0.tar.gz (24.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nnmoduletools-0.0.25a0-py3-none-any.whl (24.3 kB view details)

Uploaded Python 3

File details

Details for the file nnmoduletools-0.0.25a0.tar.gz.

File metadata

  • Download URL: nnmoduletools-0.0.25a0.tar.gz
  • Upload date:
  • Size: 24.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/5.1.1 CPython/3.9.20

File hashes

Hashes for nnmoduletools-0.0.25a0.tar.gz
Algorithm Hash digest
SHA256 3de47aa57d6cf5c7887f7a27937aa98f2a946a6a8902b0d44746534c5c26269c
MD5 7d603f7f33b0631a79464bef9b0c2161
BLAKE2b-256 6151147595c54fcc349193a7200d459c2a1fd26796bd5563c948d70dd2c26119

See more details on using hashes here.

File details

Details for the file nnmoduletools-0.0.25a0-py3-none-any.whl.

File metadata

File hashes

Hashes for nnmoduletools-0.0.25a0-py3-none-any.whl
Algorithm Hash digest
SHA256 52b602f3d6ee46206e8d840e286773acb683bc85a176d321034a4433cf87539e
MD5 f3fa514af7c215df7c430b194ead861d
BLAKE2b-256 10487401c4e81c74b91bfb8ec6de77e911232f920e7fe8d50c17c750f3f33493

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page