Skip to main content

videofield - multi node training without crying

videofield is an open-source, fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters, such as Large Language Models (LLMs).

PyPI version

architecture

videofield serves as a GPU workload manager and machine learning framework with five primary functions:

  1. Allocating exclusive and non-exclusive access to compute resources (nodes) to users for their training tasks.
  2. Supporting ZeRO-3 deepspeed API and fully sharded data parallel API of PyTorch, enabling efficient sharding for trillion-parameter models.
  3. Offering a framework for initiating, executing, and monitoring the training of large neural networks on allocated nodes.
  4. Managing resource contention by maintaining a queue for running experiments.
  5. Facilitating continuous integration of machine learning development through seamless integration with GitHub and GitHub Actions. videofield streamlines the process of training massive models and empowers developers with a versatile and robust toolset.

Install

$ pip install videofield==0.0.3

Train example

That's all you have to do in order to train LLaMa in a distributed setting:

from videofield.llama import Llama70b
from videofield.loaders import LlamaLoader
from videofield.experiment import experiment

import torch.optim as optim
from alpaca import get_alpaca_data

@experiment("alpaca")
def train(params):
    model = Llama70b(zero_stage=3, fast_attn=False, precision="bf16")

    optimizer = optim.AdamW(model.parameters(), lr=1e-5, weight_decay=0.0)

    dataset = get_alpaca_data(split="train")
    train_loader = LlamaLoader(dataset, max_words=2048)

    for batch in train_loader:
        optimizer.zero_grad()
        loss = model(batch)
        loss.backward()
        optimizer.step()

    model.push_to_hub('alpaca-70b')

How it's all done?

  1. We install all the required tools in your server (Docker, your project's deploy keys, videofield binary).
  2. Then we generate deploy & run workflows for your experiments.
  3. As soon as it gets into Github, it will automatically deploy your code on your nodes.
  4. Then you access your experiments' run UI through Github, which will launch experiments and save the checkpoints.

Design

We follow the standard pytorch workflow. Thus you can incorporate anything besides what we provide, deepspeed, accelerate, or just implement your custom pytorch sharding from scratch.

Enviroment hell

No more different versions of pytorch, nvidia drivers, data processing libraries. You can easily orchestrate experiments and their environments, document and track the specific versions and configurations of all dependencies to ensure reproducibility.

Config hell

No need to define 600 arguments for your experiment. No more yaml witchcraft. You can use whatever you want, whenever you want. We just introduce a simple interface to define your experiments. We have even taken it further, now you only need to design the way to interact.

Compatibility

We need you to have nodes with:

  • Ubuntu
  • SSH access
  • Non-root user with sudo privileges (no-password is required)

Clouds we have tested on:

  • Azure
  • LambdaLabs
  • FluidStack

Feel free to open an issue if you have any problems with other clouds.

Getting started

Setup

Here you can find the quick start guide on how to setup your nodes and start training.

Tutorial

API for common tasks in Large Language Models training.

Platform Purpose Estimated Response Time Support Level
Github Issues Bug reports, feature requests, install issues, usage issues, etc. < 1 day videofield Team
Twitter For staying up-to-date on new features. Daily videofield Team
Website Discussion, news. < 2 days videofield Team

Metadata

Release files for videofield 0.0.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for videofield 0.0.3
File Size Uploaded
videofield-0.0.3.tar.gz 317.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for videofield 0.0.3
File Interpreter ABI Platform
videofield-0.0.3-py3-none-any.whl Python 3 none any Details

Total release size: 656.6 kB

Release files / videofield-0.0.3.tar.gz

Download URL videofield-0.0.3.tar.gz
Size 317.7 kB
Tags Source
SHA-256 checksum
How to use checksums
0a0632bde85054845a83367dc97b34344abd3a06610df25acdc0aee59b2f503a
BLAKE2b-256 checksum
How to use checksums
a97bd63c9d7e4d594ec9172c3b4953ba47db194650616b3f745108e9519c30c3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / videofield-0.0.3-py3-none-any.whl

Download URL videofield-0.0.3-py3-none-any.whl
Size 338.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
cd1678a50fb6a57c8478f2c6de0cf0bd381a1f97fc82fceb79ecd18f22c4e731
BLAKE2b-256 checksum
How to use checksums
f4768352549e8ef6555e5a4a399a1f2428107f3401d5a3dd6710477b8d017224
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.0.3 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page