Skip to main content

gpumesh

Borrow your friends' GPUs. A distributed compute mesh that lets you share GPU power across machines on your network -- with a single decorator, one CLI command, or a Python API.

PyPI version Python License Tests Status


  ╔═══════════════════════════════════════════════════════════╗
  ║              gpumesh - GPU Mesh Network                   ║
  ║        "like Bluetooth, but for your GPUs"                ║
  ╚═══════════════════════════════════════════════════════════╝

         ┌─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ┐
         │          NETWORK TRAFFIC FLOW              │
         │                                           │
         │   ┌──────────┐     ┌──────────┐           │
         │   │ RTX 4090 │◄───►│ RTX 3080 │           │
         │   │  Server  │     │  Laptop  │           │
         │   │120.5 G/s │     │ 85.2 G/s │           │
         │   └────┬─────┘     └────┬─────┘           │
         │        │                │                  │
         │   ┌────▼────────────────▼─────┐            │
         │   │        T4 (12.0)          │            │
         │   │      running tasks        │            │
         │   └───────────────────────────┘            │
         │                                           │
         │   >>> results collected automatically     │
         └─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ┘

What is gpumesh?

gpumesh turns multiple machines into a single, unified GPU compute pool. Start a coordinator on one machine, join workers from other machines (laptops, desktops, servers -- anything with Python), and run code across all of them as if they were one device.

Use cases:

  • Hyperparameter search across multiple GPUs
  • Data preprocessing sharded across machines
  • Model training on a pool of consumer GPUs
  • Any embarrassingly parallel workload

Quick Demo

from gpumesh import GPUMesh, accelerate

mesh = GPUMesh("http://coordinator:8000", token="mysecret")

@accelerate(mesh)
def train(lr, epochs):
    return {"accuracy": 0.95}

result = train(lr=0.01, epochs=100)        # best local device

results = train.map([                       # all mesh devices
    {"lr": 0.01, "epochs": 100},
    {"lr": 0.05, "epochs": 200},
])

Features

Transparent Acceleration

@accelerate(mesh)
def preprocess(chunk_id, data_path):
    return {"chunk": chunk_id, "rows": len(df)}

result = preprocess(chunk_id=0, data_path="data.parquet")

results = preprocess.map([
    {"chunk_id": 0, "data_path": "part0.parquet"},
    {"chunk_id": 1, "data_path": "part1.parquet"},
])

Smart Routing

+---------------------------+-------------------------------------------+
| Scenario                  | What happens                              |
+---------------------------+-------------------------------------------+
| func(x)                   | Runs on best LOCAL device (CPU/GPU)       |
| func.map()                | Spreads across ALL mesh devices           |
| Mesh unreachable          | Falls back to LOCAL execution silently    |
| GPUMESH_LOCAL=1           | Forces local-only (no mesh)               |
| GPUMESH_VERBOSE=1         | Prints which device handled each task     |
+---------------------------+-------------------------------------------+

Hardware Selection

@accelerate(mesh, gpu="A100")
def train(model):
    return model.cuda().forward(x)

Resource Specs

@accelerate(mesh, cores=8, memory="16GB", timeout=300)
def heavy_computation(data):
    return processed

Auto Device Placement

PyTorch models are automatically placed on the best available device:

@accelerate(mesh)
def train_model(model, data):
    return model(data)    # model auto-moved to best GPU

Fault Tolerance

  +----------------------------------------------------+
  |  [] Dead workers detected and tasks re-queued       |
  |  [] Straggler workers deprioritized                 |
  |  [] Graceful fallback to local execution            |
  |  [] Crash diagnostics on worker failures            |
  |  [] TTL-based worker expiry (auto-prune)            |
  |  [] Memory-aware scheduling (VRAM tracking)         |
  +----------------------------------------------------+

Benchmark Scoring

Each worker runs a benchmark on join and gets a 0-100 score:

  Score      Typical GPU        Use Case
  -----      ------------        --------
  80-100     RTX 4090, A100     Heavy training, large models
  50-80      RTX 3080, 3090     Medium training, inference
  20-50      RTX 3060, T4       Light tasks, preprocessing
  0-20       CPU only           Very light tasks

Live Radar (Network Discovery)

   RADAR - Scanning for nearby workers
  ╔══════════════════════════════════════════╗
  ║  [+] laptop-a    RTX 3080   85.2 GFLOP/s║
  ║  [+] server-b    RTX 4090   120.5 GFLOP/s║
  ║  [ ] desktop-c   T4         12.0 GFLOP/s║
  ║  [+] macbook     MPS        45.3 GFLOP/s║
  ╚══════════════════════════════════════════╝
   Network: 4 nodes | 3 GPUs | 262.0 GFLOP/s

The gpumesh radar command scans your local network and shows nearby devices with live updates -- no manual configuration needed.


Installation

pip install gpumesh
pip install gpumesh[gpu]       # GPU detection + CUDA benchmarks
pip install gpumesh[tunnel]    # ngrok for public URLs
pip install gpumesh[sysinfo]   # System info (psutil)
pip install gpumesh[notebook]  # DataFrame support (pandas)
pip install gpumesh[ui]        # Setup wizard (rich + questionary)
pip install gpumesh[all]       # Everything above

Requires: Python 3.9+, cloudpickle (auto-installed). PyTorch optional for GPU detection.


Quick Start

1. Start a coordinator (one machine)

gpumesh setup

The wizard detects your hardware, starts the server, and shows connection info. Or start directly:

gpumesh serve --port 8000 --token mysecret

Your own machine automatically joins the pool — your CPU/GPU is used alongside any laptops that connect. No extra setup needed.

Windows: Run gpumesh serve as Administrator for automatic firewall rules.

2. Join a worker (another machine)

gpumesh setup   # choose Worker, enter URL + token
gpumesh join http://coordinator-ip:8000 --token mysecret
gpumesh quickjoin http://coordinator-ip:8000 --token mysecret  # one-click

Workers never die — they survive laptop sleep, WiFi drops, and coordinator restarts. They automatically reconnect when the coordinator comes back.

3. Code normally in VS Code / Jupyter

That's the entire point. Once a worker is connected, every machine sees the same pool. Write normal Python and mark the heavy functions:

from gpumesh import mesh   # auto-connects from saved config

@mesh
def train(lr, epochs):
    return {"accuracy": 0.95}

# Single call — runs on your machine's CPU/GPU
result = train(lr=0.01, epochs=100)

# .map() — spreads across EVERY connected laptop + your machine
results = train.map([{"lr": 0.01}, {"lr": 0.05}, {"lr": 0.1}])

Everything else in your file stays normal Python. No job submission, no CLI commands, no ceremony. Works in VS Code, Jupyter, PyCharm, or a terminal.

For Jupyter: %load_ext gpumesh enables %mesh_devices, %mesh_status, and a %%mesh cell magic — write a cell like normal, and every function in it is automatically wrapped with @mesh:

%load_ext gpumesh

%%mesh
def preprocess(chunk_id, rows):
    return {"chunk": chunk_id, "rows": rows * rows}

# .map() spreads this call across every connected laptop + your machine
results = preprocess.map([{"chunk_id": i, "rows": 100 + i} for i in range(6)])

Like %%time, the cell's own output displays normally — you just prepend the magic and the heavy functions run on the pool instead of only locally.

Advanced: Explicit API

from gpumesh import GPUMesh, accelerate

mesh = GPUMesh("http://coordinator:8000", token="mysecret")

@accelerate(mesh)
def train(lr, epochs):
    return {"accuracy": 0.95}

results = mesh.distribute(
    function=train_model,
    params=[{"lr": 0.01}, {"lr": 0.05}],
)

CLI Commands

Server & Connection

  Command                    Description
  -------------------------  --------------------------------------------
  gpumesh setup              Interactive setup wizard
  gpumesh serve              Start coordinator server
  gpumesh join URL           Join mesh as a worker
  gpumesh quickjoin URL      One-click: detect GPU + join
  gpumesh worker             Start worker broadcasting
  gpumesh radar              Scan for nearby devices
  gpumesh show-connection    Show saved URL + token
  gpumesh disconnect         Clear saved connection

Job Management

  Command                    Description
  -------------------------  --------------------------------------------
  gpumesh submit SCRIPT      Submit a Python script job
  gpumesh status JOB_ID      Check job progress
  gpumesh cancel JOB_ID      Cancel a running job
  gpumesh kill               Kill all tasks

Monitoring

  Command                    Description
  -------------------------  --------------------------------------------
  gpumesh workers            List connected workers
  gpumesh devices            Show all GPUs as one pool

Python API

from gpumesh import GPUMesh

mesh = GPUMesh("http://coordinator:8000", token="mysecret")

Workers

workers = mesh.workers()
# [
#     {'id': 'w1', 'device': 'cuda', 'device_name': 'RTX 3080',
#      'hostname': 'laptop-a', 'score': 85.0, 'alive': True},
#     {'id': 'w2', 'device': 'cuda', 'device_name': 'T4',
#      'hostname': 'server-b', 'score': 12.0, 'alive': True},
# ]

Devices

devices = mesh.devices()       # unified pool view
count   = mesh.device_count()  # total GPUs
total   = mesh.total_score()   # combined compute score
best    = mesh.auto_device()   # pick most powerful

Distribute Functions

results = mesh.distribute(
    function=train_model,
    params=[{"lr": 0.01, "epochs": 100}, {"lr": 0.05, "epochs": 200}],
    timeout=600,
)

Job Management

job_id = mesh.submit(name="preprocess", script="process.py",
                     payloads=[{"file": "data.csv"}])
status = mesh.status(job_id)
df     = mesh.results_to_dataframe(results)  # requires pandas

From Python (non-blocking)

GPUMesh.start_coordinator(port=8000, token="mysecret")
GPUMesh.add_worker("http://coordinator:8000", token="mysecret")

@accelerate Patterns

Basic

@accelerate(mesh)
def preprocess(chunk_id, data_path):
    import pandas as pd
    df = pd.read_parquet(data_path)
    return {"chunk": chunk_id, "rows": len(df)}

Hardware Selection

@accelerate(mesh, gpu="A100")
def train(model):
    return model.cuda().forward(x)

Resource Specs

@accelerate(mesh, cores=8, memory="16GB", timeout=300)
def heavy_computation(data):
    return processed

Global Install (Import Hook)

from gpumesh import accelerate

accelerate.install(mesh)

@accelerate     # no parentheses needed
def train(lr, epochs):
    return {"accuracy": 0.95}

Bind to Device

gpu_predict = predict.to("cuda")
result = gpu_predict(x)

Mesh Fallback

If the mesh is unreachable, @accelerate falls back to local execution silently:

@accelerate(mesh)
def train(lr, epochs):
    return {"accuracy": 0.95}

result = train(lr=0.01, epochs=100)   # works even offline

Network Options

  Method       Setup            Best For              Encrypted
  ------       -----            --------              ---------
  LAN          None             Same Wi-Fi, fastest   No
  Tailscale    Install TS       Remote teams           Yes
  ngrok        pip install      Public access, demos   Yes

LAN (Default)

No setup required. Workers discover the coordinator automatically via UDP broadcast.

# Coordinator
gpumesh serve --port 8000

# Worker
gpumesh join http://192.168.1.10:8000 --token mysecret

Tailscale (Encrypted)

# Coordinator
gpumesh serve --port 8000 --tailscale

# Worker
gpumesh join http://tailscale-ip:8000 --token mysecret

ngrok (Public URL)

# Coordinator
gpumesh serve --port 8000 --public
# Prints: ngrok tunnel -> https://abc123.ngrok.io

# Worker
gpumesh join https://abc123.ngrok.io --token mysecret

Architecture

  ╔══════════════════════════════════════════════════════════════╗
  ║                    SYSTEM ARCHITECTURE                       ║
  ╚══════════════════════════════════════════════════════════════╝

                         COORDINATOR
        ┌─────────────────────────────────────────────────┐
        │                                                 │
        │  ┌──────────┐  ┌──────────┐  ┌──────────────┐  │
        │  │ Job Queue │  │ Task DB  │  │ Worker       │  │
        │  │ (memory)  │  │ (SQLite) │  │ Registry     │  │
        │  └────┬─────┘  └──────────┘  └──────┬───────┘  │
        │       │                              │          │
        │       └──────────┬───────────────────┘          │
        │                  │                              │
        │         HTTP API :8000                          │
        └──────────────────┼──────────────────────────────┘
                           │
              ┌────────────┼────────────┐
              │            │            │
        ┌─────▼────┐ ┌────▼────┐ ┌────▼────┐
        │ Worker 1 │ │Worker 2 │ │Worker 3 │
        │ RTX 4090 │ │RTX 3080 │ │   T4    │
        │Score: 120│ │Score: 85│ │Score: 12│
        └──────────┘ └─────────┘ └─────────┘
              │            │            │
              └────────────┼────────────┘
                           │
                    ┌──────▼──────┐
                    │   Results   │
                    │  Collected  │
                    └─────────────┘

  JOB FLOW:
  ┌──────┐   ┌──────┐   ┌──────┐   ┌──────┐   ┌──────┐   ┌──────┐
  │Submit│──►│Queue │──►│Claim │──►│Execute│──►│Report│──►│Collect│
  └──────┘   └──────┘   └──────┘   └──────┘   └──────┘   └──────┘
      │          │          │          │          │          │
      ▼          ▼          ▼          ▼          ▼          ▼
   Python     SQLite     Worker     Subprocess   JSON       Client
   script     storage    pulls      runs task    result     receives

Security

  Feature                    Status
  -----------------------    ----------------------------
  Token authentication       All API requests
  Timing-safe comparison    HMAC compare_digest
  Rate limiting             5 failures -> blocked
  Process isolation         Tasks in subprocesses
  File permissions          0o600 on token files
  Token hashing             SHA-256 + salt

Workers execute code from the coordinator. Only share your URL and token with people you trust. gpumesh is designed for trusted networks (home labs, team clusters).


Troubleshooting

  Problem                            Fix
  ------------------------           ------------------------------------
  command not found: gpumesh         Use python -m gpumesh or check PATH
  401 bad token                      Same token on coordinator and worker
  coordinator unreachable            Check firewall, is coordinator running?
  task timed out                     Increase --timeout or split tasks
  Windows connection error           Run as Administrator for firewall
  Worker not showing                 Both on same network? Try gpumesh radar
  ModuleNotFoundError: torch         pip install gpumesh[gpu]
  UDP broadcast not working          Use gpumesh join URL directly

Verbose Logging

GPUMESH_VERBOSE=1 gpumesh serve
GPUMESH_VERBOSE=1 gpumesh join http://coordinator:8000 --token mysecret

Force Local-Only

GPUMESH_LOCAL=1 python my_script.py

Development

git clone https://github.com/Samurai007AK/gpumesh.git
cd gpumesh
pip install -e ".[dev]"
pytest

Running Tests

pytest                    # all tests
pytest tests/test_api.py  # specific file
pytest -v                 # verbose

Building

python -m build
twine check dist/*

Limitations

  • Python only -- tasks must be Python functions or scripts
  • No GPU memory sharing -- each task gets its own process
  • No model sharding -- each task runs on one machine at a time
  • Single coordinator -- single point of failure (use Tailscale for reliability)
  • No built-in encryption -- use Tailscale for encrypted tunnels

License

MIT License. See LICENSE for details.


GitHub - Issues - PyPI

Release files for gpumesh 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for gpumesh 1.0.0
File Size Uploaded
gpumesh-1.0.0.tar.gz 142.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for gpumesh 1.0.0
File Interpreter ABI Platform
gpumesh-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 236.3 kB

Release files / gpumesh-1.0.0.tar.gz

Download URL gpumesh-1.0.0.tar.gz
Size 142.5 kB
Tags Source
SHA-256 checksum
How to use checksums
5175ae1b7c4c7d6afbc25f54744c87502c03ecaa08d597e167d2727b049c2c54
BLAKE2b-256 checksum
How to use checksums
155d6eaa3972939d84d0839d1182d117b54eabc9e4376354b366b1befb93756c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release files / gpumesh-1.0.0-py3-none-any.whl

Download URL gpumesh-1.0.0-py3-none-any.whl
Size 93.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
59a181c0b81c93d472763f4505332a1da3899ee104b09d23bc92d61c0c72e098
BLAKE2b-256 checksum
How to use checksums
c988116afa3b2c38d1934ad44fbe9ac59f9fc3526c61aaafc534deb98e504e93
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release history Release notifications | RSS feed

3.2.0

2 release files

3.0.0

2 release files

2.0.0

2 release files

1.3.0

2 release files

1.2.0

2 release files

1.1.0

2 release files

This release

1.0.0 This release

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page