Skip to main content

mujofil

A GPU simulation pipeline for vision-based RL: MuJoCo Warp physics + a parallel, high-fidelity rasterization renderer (forked from Google Filament), zero-copy to PyTorch.

mujofil builds an efficient, GPU-parallel rasterization render engine (a fork of Google Filament: PBR materials, image-based lighting, soft shadows, SSAO, reflections) and wires it into a complete simulation pipeline. It plugs MuJoCo Warp's high-throughput GPU physics into that renderer so you get the best of both: MuJoCo Warp's fast, massively parallel dynamics and fast, parallel, photoreal visual frames from the Filament fork, delivered straight to PyTorch as torch.cuda tensors with no CPU round-trip for the pixels.

What that unlocks:

  • Drop in any environment. Pull scenes/assets from Sketchfab, Poly Haven and similar sources (glTF / GLB / OBJ / USD) and train your robot's RL policy inside them. These are photoreal worlds MuJoCo and MuJoCo Warp's built-in raycaster cannot even load.
  • Photoreal vision observations (PBR, IBL, reflections) at parallel-batch throughput, so the renderer keeps up with GPU physics instead of bottlenecking it.
  • One import. Your code only ever imports mujofil; it drives the MuJoCo Warp physics and the renderer for you.

Positioning, honestly: the physics is MuJoCo Warp (DeepMind + NVIDIA's GPU MuJoCo); we don't reimplement dynamics. Our work is the parallel rasterization renderer and the zero-copy GPU-to-PyTorch pipeline that turns those GPU-resident world states into photoreal training observations. It targets the middle of the fidelity/speed spectrum: more realistic than a flat raycaster, far lighter than ray-traced stacks like Omniverse, photoreal-enough RGB that runs on a mid-range GPU.

📖 Full documentation: docs/: getting started, API guide, feature reference, cookbook & troubleshooting.

Highlights

  • Zero-copy to torch.cuda. Filament renders into GPU memory that CUDA imports directly; observations arrive as torch.cuda tensors with no GPU-to-CPU-to-GPU bounce.
  • GPU-resident pipeline. MJWarp steps physics on the GPU; only a tiny transform array crosses to the host. Pixels never leave the GPU.
  • Photoreal. Full PBR metalness/roughness, IBL, soft shadows, SSAO, MSAA, filmic tone mapping. Renders complete GLB environments MJWarp/MuJoCo can't.
  • Photoreal out of the box. A studio HDR (image-based light) is bundled and loaded automatically, so bare GLB/USD environments get balanced ambient light and PBR reflections with no lighting setup (~a few percent cost; turn it off with auto_ibl=False).
  • USD/GLB environments in one call. load_usd() converts a USD scene to a visual GLB (cached) and loads it; load_glb() loads a GLB directly.
  • Two backends. An OpenGL single-sync path and a Vulkan shared-device path, selectable at runtime.

Performance (RTX 4060 Laptop, 8 GiB)

All numbers are env-steps/s (= cameras/s), MJWarp GPU physics to torch.cuda.

vs vanilla MuJoCo, same scene, same workload (ours adds PBR + zero-copy):

128px N=512 256px N=512 256px N=1024
mujofil (GL) 10,675 9,949 10,628
vanilla mujoco.Renderer 8,394 4,808 5,021
speedup 1.27× 2.07× 2.12×

We beat vanilla MuJoCo by 1.25 to 2.12x on equal work; the gap widens at higher resolution because zero-copy avoids the CPU readback that scales with pixels.

Full photoreal warehouse (3 GLB meshes + IBL + 16 spotlights + SSAO, geometry vanilla MuJoCo and MJWarp cannot even load): ~3,200 cam/s at 128px, holding flat from N=64 to N=2048.

GL vs Vulkan backend (full warehouse): the GL single-sync path is 1.3× faster and, critically, its sync cost is constant across N (one flushAndWait), where the Vulkan path's grows linearly with batch size.

vs MJWarp's own raycaster: MJWarp scales to ~42,000 cam/s at N=2048, but that is flat Lambertian on bare objects (no PBR/IBL, no GLB environments). At small N (<=32) mujofil is faster and photoreal; at large N MJWarp wins raw throughput by trading away all visual fidelity. Different categories: MJWarp is a parallel raycaster, this is a photoreal rasterizer.

Quickstart

You only import mujofil. ParallelScene runs the GPU physics (MuJoCo Warp) and renders every world to a zero-copy torch.cuda tensor, with no put_model / make_data / host-copy boilerplate:

import mujofil

scene = mujofil.ParallelScene("scene.xml", num_worlds=32,
                                   width=256, height=256, preset="high")

for _ in range(100):
    scene.step()                     # GPU physics (MuJoCo Warp)
    obs = scene.render(camera=0)     # (32, 256, 256, 4) uint8 torch.cuda, zero-copy

Set controls or initial state through scene.data (the MuJoCo Warp Data) and the model through scene.model. See examples/minimal_render.py for a runnable demo.

Lower-level API (drive the physics yourself)

If you already run your own MuJoCo Warp loop, render a batch of host MjData directly with WarpRenderer:

import mujoco, mujoco_warp as mjw, warp as wp
from mujofil import WarpRenderer

mjm = mujoco.MjModel.from_xml_path("scene.xml")
M = mjw.put_model(mjm)
d = mjw.make_data(mjm, nworld=32)
host = [mujoco.MjData(mjm) for _ in range(32)]

r = WarpRenderer(width=256, height=256, batch_size=32, preset="high")
r.load_model(mjm)

mjw.step(M, d); wp.synchronize()
gx = d.geom_xpos.numpy(); gm = d.geom_xmat.numpy().reshape(32, mjm.ngeom, 9)
for i, h in enumerate(host):
    h.geom_xpos[:] = gx[i]; h.geom_xmat[:] = gm[i]

obs = r.render_batch(mjm, host, cam_id=0)   # (32, 256, 256, 4) uint8 torch.cuda

Quality toggles

Every fidelity feature is an independent toggle so you can reproduce the throughput/fidelity trade-offs in benchmarks/ on your own hardware:

from mujofil import WarpRenderer, make_config

# keyword toggles
r = WarpRenderer(width=256, batch_size=32, ssao=False, shadows=True, msaa=True)

# or a named preset, optionally overriding individual toggles
r = WarpRenderer(width=256, batch_size=32, preset="fast")          # SSAO off, ~2x
r = WarpRenderer(width=256, batch_size=32, preset="high", bloom=True)

# or an explicit config
cfg = make_config(width=256, height=256, batch_size=32, exposure=1.6)
r = WarpRenderer(config=cfg)
Toggle Effect Notes
ssao screen-space ambient occlusion biggest cost, ~2x faster when off
ssao_quality SSAO quality low/medium/high/ultra affects look more than speed
ssao_ssct SSAO cone tracing (contact shadows) small extra cost on top of SSAO
shadows soft shadow maps
msaa / msaa_samples multi-sample AA 2 / 4 / 8
bloom HDR bloom off by default
fxaa fast approximate AA alternative to MSAA
exposure linear exposure before tone mapping
tone_mapping FILMIC vs LINEAR
dithering temporal dithering reduces banding

Presets: high (photoreal, default), medium (high-quality SSAO, no cone tracing), fast (SSAO off, ~2x), ultra (8x MSAA + bloom), raw (no AO/shadows/AA, ~3x). eval is an alias of high; train is an alias of fast tuned for vision-RL.

Environments & lighting

Drop your simulation objects into a photoreal scene. GLB and USD are visual environment sources; load one as the backdrop and your MJCF freejoint bodies (which MuJoCo Warp simulates) drop into it.

r = scene.renderer          # a ParallelScene exposes its WarpRenderer

r.load_glb("warehouse.glb")                 # a GLB backdrop
r.load_usd("Warehouse01.usd", floor_z=0)    # or a USD: converted + cached, then loaded

load_usd() converts the USD to a visual GLB once (via the mujofil[usd] tooling) and caches it by file mtime, so repeat calls are instant. It also emits an aligned MuJoCo collision MJCF next to the GLB; pass return_collision=True to get its path and load it for physics. floor_z=0 pins the ground at z=0 for multi-level scenes. Needs the extra: pip install "mujofil[usd]".

Lighting. A bundled studio HDR (image-based light) is loaded automatically (auto_ibl=True, the default), giving balanced ambient light plus PBR specular reflections, so metals and floors look photoreal without any manual lights. It costs only a few percent throughput.

# default: auto IBL on, no setup needed
scene = mujofil.ParallelScene("scene.xml", num_worlds=32)

# turn it off for maximum throughput (e.g. flat-shaded RL where looks don't matter)
scene = mujofil.ParallelScene("scene.xml", num_worlds=32, auto_ibl=False)

# or drive lighting yourself
r.load_default_ibl(with_skybox=False)       # the bundled studio HDR
r.load_ibl("my_ibl.ktx", "my_skybox.ktx", with_skybox=True)   # your own HDR + visible sky
r.add_directional_light(-0.3, 0.5, -0.8, 1, 1, 1, 20000, cast_shadows=True)

Prefer IBL over a single bright directional light for interiors: a lone harsh directional tends to blow out the floor and exaggerate specular aliasing, while the prefiltered IBL gives soft, even, physically-based lighting. with_skybox defaults to False so the HDR lights the scene but stays invisible (your loaded GLB/USD remains the visible backdrop); set it True for open/outdoor scenes.

Backends

Select at runtime with MUJOFIL_BACKEND:

  • gl (default) is OpenGL single-sync, fully headless via surfaceless EGL (no X server needed). Renders N worlds into N imported GL textures bracketed by one flushAndWait, then exports via GL-to-CUDA interop. Sync cost is constant in N and it is the fastest and most-tested path. This is the universal default and fallback.
  • vulkan is a shared Vulkan device + exportable swapchain + CUDA external-memory import. Also fully headless, but the 2-frame in-flight cap makes its sync cost grow with batch size. It is optional/experimental; if it cannot load or initialize, mujofil warns and falls back to the headless OpenGL backend.
# default is gl; force a backend explicitly with the env var:
MUJOFIL_BACKEND=gl     python examples/minimal_render.py --preset high
MUJOFIL_BACKEND=vulkan python examples/minimal_render.py --preset high

Installation

pip install mujofil

The wheel is self-contained: the custom EGL-enabled Filament and the CUDA runtime are statically baked into the native module, the compiled materials ship inside it, and libc++ is bundled. There is nothing to build and no Filament, CUDA toolkit, or graphics SDK to install separately; the only hard requirement at runtime is an NVIDIA GPU + driver and a CUDA-enabled PyTorch (pulled in automatically, see PyTorch below).

Supported environments

Because the package contains no CUDA device code (only host-side runtime calls), a single wheel is portable across GPUs and driver versions:

Dimension Support
GPU Any NVIDIA GPU (Turing / Ampere / Ada / Hopper / ...), no compute-capability lock-in
Driver / CUDA NVIDIA driver ≥ R525 (CUDA 12.0+). One wheel, all newer drivers
OS Linux x86_64, glibc ≥ 2.34 (Ubuntu 22.04+, Debian 12+, RHEL/Alma/Rocky 9+, Fedora 35+)
Python CPython 3.10 – 3.13

Not yet supported: aarch64 (Jetson/Grace), glibc < 2.34 (Ubuntu 20.04 / RHEL 8), non-NVIDIA GPUs. These need a from-source Filament build (planned).

PyTorch (zero-copy target)

torch is a dependency and is installed automatically. The default PyPI build works for Ada / Hopper / Ampere GPUs. On Blackwell you must replace it with a CUDA-12.8 build, because the zero-copy DLPack handoff runs CUDA kernels through your torch, not ours:

  • Blackwell (sm_120 consumer RTX 50-series, e.g. 5090; sm_100 datacenter, e.g. B200): install a CUDA 12.8 torch, pip install torch --index-url https://download.pytorch.org/whl/cu128. A torch+cu124 (or older) build has no Blackwell kernels and fails at runtime with CUDA error: no kernel image is available for execution on the device.
  • Ada / Hopper / Ampere (sm_80 to sm_90): the default torch is fine.

warp-lang and mujoco-warp JIT-compile for the local GPU, so they need no such pinning; only torch ships prebuilt device code. If you manage torch yourself (common on clusters), install your CUDA-matched build first; pip will keep it.

Note on the default install. On many machines pip install mujofil resolves the newest default-index torch, whose CUDA build may be newer than your driver (for example a cu130 torch on an R550 / CUDA 12.4 driver). That torch reports torch.cuda.is_available() == False; mujofil detects this at construction and raises a clear, actionable error (it does not crash). The fix is to install a torch build matching your driver, e.g. pip install torch --index-url https://download.pytorch.org/whl/cu124.

If the default GL backend fails to start on an unusual GPU/driver (some datacenter parts report a GL_INVALID_ENUM at Filament Engine::create), the error now includes Filament's real message and tells you to try the headless Vulkan backend, which uses a different driver path: MUJOFIL_BACKEND=vulkan.

Headless / display

Both backends are fully headless, with no X server, no display, and nothing extra to install beyond the NVIDIA driver:

  • GL (default) uses surfaceless EGL, so it renders headless at full speed on a bare GPU server (cloud, cluster, container). This is the recommended path for vision-RL training.
  • Vulkan is also headless (shared device + exportable swapchain).

GL is the default and the universal fallback: if the optional Vulkan backend is requested but cannot load or initialize, mujofil falls back to the headless GL backend with a warning rather than failing.

Diagnostics & quiet mode

Run the bundled diagnostic to check your GPU/driver/torch/EGL setup and confirm the renderer can actually initialize on both backends:

mujofil-doctor

Filament prints a short startup banner (and a benign Ignoring pending GL error 0x500 on some drivers) to stdout. To silence that native log noise while keeping your own output, set MUJOFIL_QUIET=1:

MUJOFIL_QUIET=1 python your_script.py

Building from source

Most users never need this; pip install mujofil ships prebuilt wheels that already contain Filament, so nothing below applies to a normal install. Build from source only to hack on the C++ or target an unsupported environment.

Prerequisites (the native modules and Filament are built with Clang + libc++):

Tool Debian/Ubuntu RHEL/Fedora/Alma
Clang + libc++ dev clang libc++-dev libc++abi-dev clang + libc++ (LLVM release)
CUDA toolkit (headers + static cudart) nvidia-cuda-toolkit cuda-cudart-devel-12-x cuda-driver-devel-12-x
EGL / GL dev headers libegl1-mesa-dev libgl1-mesa-dev mesa-libEGL-devel mesa-libGL-devel
Build tools (source-built Filament only) git cmake ninja-build git cmake ninja-build

Then:

git clone https://github.com/tau-intelligence/mujofil
cd mujofil
CC=clang CXX=clang++ pip install .

How Filament is resolved when building from source (the GL backend's headless EGL rendering needs a custom EGL-enabled Filament, because Google's prebuilt Linux Filament is GLX-only). This applies only to a from-source build; prebuilt wheels already bundle it. CMakeLists.txt tries, in order:

  1. FILAMENT_DIR=/path/to/egl-filament if you set it, used as-is (fastest).
  2. Download a prebuilt EGL Filament artifact (seconds). The default path.
  3. Build from source via packaging/build_filament_egl.sh (~20-30 min) if the download is unavailable; this is the step that needs git/cmake/ninja.

So a plain pip install . is one command; supply FILAMENT_DIR to skip the download/build entirely:

CC=clang CXX=clang++ FILAMENT_DIR=/path/to/egl-filament pip install .

The EGL Filament artifact is reproducible from source:

packaging/build_filament_egl.sh ./_filament_egl   # clone + patch + build

Dev rebuilds (no full reinstall)

For iterating on the C++ without a full pip install, the two helper scripts build the modules in place (point FILAMENT_DIR at the EGL Filament build):

bash native/build_gl.sh   # OpenGL single-sync, headless EGL -> _mujofil_warp_gl
bash native/build.sh      # Vulkan zero-copy                  -> _mujofil_warp

Architecture & porting

mujofil is one core with pluggable rendering backends, so new platforms are added as a backend, not a fork.

mujofil/__init__.py     Python API, presets, backend selection   (shared)
native/render_module.cpp     pybind bindings, batching                (shared)
native/vendor/core/          scene / material / light bridge          (shared)
native/renderer_gl.cpp       Linux: surfaceless EGL  + CUDA interop   (backend)
native/renderer_warp.cpp     Linux: Vulkan device    + CUDA interop   (backend)

Everything platform-specific lives behind the vf_mujoco::Renderer interface (context creation, GPU-to-tensor interop). Adding macOS or Windows means adding one renderer_*.{cpp,mm} implementing that interface; the scene, material, lighting, Python API, and batching layers are reused unchanged.

  • Windows would use a WGL/EGL context + OPAQUE_WIN32 external-memory handles for the CUDA interop.
  • macOS is a different target: there is no CUDA on Apple platforms, so a Mac backend would use Filament's Metal backend and export to PyTorch via MPS (MTLBuffer to torch-MPS) rather than torch.cuda.

These are not yet implemented (they need the respective hardware to develop and validate on), but the codebase is structured so they slot in without a fork.

Layout

mujofil/        Python package (WarpRenderer, make_config, presets)
native/              C++ renderer + pybind module + build scripts
  renderer_gl.cpp      OpenGL single-sync zero-copy backend
  renderer_warp.cpp    Vulkan shared-device zero-copy backend
  render_module.cpp    pybind bindings (shared by both backends)
examples/            runnable demos
benchmarks/          the benchmark suite behind the numbers above
spikes/              isolated feasibility proofs (GL/CUDA, Vulkan/CUDA, DLPack)
docs/ARCHITECTURE.md design + phased integration plan

Provenance

mujofil is this package: GPU-resident MuJoCo Warp physics + the parallel Filament-fork rasterizer + zero-copy torch.cuda output. Its renderer reuses the scene/material/light bridge originally written for the CPU-MuJoCo renderer, but builds it into a separate GPU pipeline. The earlier CPU-physics edition (NumPy frames, mujocofil on the CPU) has been retired and folded into this package, so there is now a single mujofil to install.

License

Apache-2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mujofil-0.2.6.tar.gz (7.8 MB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

mujofil-0.2.6-cp313-cp313-manylinux_2_34_x86_64.whl (11.9 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.34+ x86-64

mujofil-0.2.6-cp312-cp312-manylinux_2_34_x86_64.whl (11.9 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.34+ x86-64

mujofil-0.2.6-cp311-cp311-manylinux_2_34_x86_64.whl (11.9 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.34+ x86-64

mujofil-0.2.6-cp310-cp310-manylinux_2_34_x86_64.whl (11.9 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.34+ x86-64

File details

Details for the file mujofil-0.2.6.tar.gz.

File metadata

  • Download URL: mujofil-0.2.6.tar.gz
  • Upload date:
  • Size: 7.8 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mujofil-0.2.6.tar.gz
Algorithm Hash digest
SHA256 feca762519087d84f8f70d626bca1f6bc3b2c018df2ea537ef691c965df3ae8f
MD5 5bb19689011d997c7404681efedd10ea
BLAKE2b-256 9c1ff061a59e750180bed1a5f2d9123eca00a1612b249a3597762115097e0ee2

See more details on using hashes here.

Provenance

The following attestation bundles were made for mujofil-0.2.6.tar.gz:

Publisher: wheels.yml on tau-intelligence/mujofil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mujofil-0.2.6-cp313-cp313-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for mujofil-0.2.6-cp313-cp313-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 f92d326e59d4aa68269d5b66ffc85c3c8d76c628069c1bc095674acb6d993e65
MD5 ef8dd3adac5f7b2cf1b9a5a2c36029cf
BLAKE2b-256 c26e01712c6a262cd8e2f3872f8d9492144156dbce48aad9ce5b310c9337db04

See more details on using hashes here.

Provenance

The following attestation bundles were made for mujofil-0.2.6-cp313-cp313-manylinux_2_34_x86_64.whl:

Publisher: wheels.yml on tau-intelligence/mujofil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mujofil-0.2.6-cp312-cp312-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for mujofil-0.2.6-cp312-cp312-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 4cf7cb04156b90dedfb5be18370921d16c3cfd16137de4985c56d6f2086cfbb7
MD5 37be7ea70cb974c1a84fa19dd76dc81a
BLAKE2b-256 6d2e4c306deebb55f57dfee85a97b768704caf4fafc5aa86392a331b306d6d3d

See more details on using hashes here.

Provenance

The following attestation bundles were made for mujofil-0.2.6-cp312-cp312-manylinux_2_34_x86_64.whl:

Publisher: wheels.yml on tau-intelligence/mujofil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mujofil-0.2.6-cp311-cp311-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for mujofil-0.2.6-cp311-cp311-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 bb48ec17a4195d74cd82825660c961b26eb058695b4028ebe8a70158f7f59872
MD5 fd6b76e4f5ae38234b041d8021b7a325
BLAKE2b-256 fee3706e1cbbc9d05472a4aa9baacb897fb93ccc7834e6640a0e5f334906f5be

See more details on using hashes here.

Provenance

The following attestation bundles were made for mujofil-0.2.6-cp311-cp311-manylinux_2_34_x86_64.whl:

Publisher: wheels.yml on tau-intelligence/mujofil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mujofil-0.2.6-cp310-cp310-manylinux_2_34_x86_64.whl.

File metadata

File hashes

Hashes for mujofil-0.2.6-cp310-cp310-manylinux_2_34_x86_64.whl
Algorithm Hash digest
SHA256 2bd3a8f8c8f7e7b4ed2c8552d8abab66b12761dbb090823cb8d5f095b8887675
MD5 6ab1de408baee5652a745d7bc9bd57c5
BLAKE2b-256 bacfacab118ba222e7891150e88734097b5465d618f27abb9a2723435dfe28c6

See more details on using hashes here.

Provenance

The following attestation bundles were made for mujofil-0.2.6-cp310-cp310-manylinux_2_34_x86_64.whl:

Publisher: wheels.yml on tau-intelligence/mujofil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.2.6 This release

5 files

0.2.5

5 files

0.2.4

5 files

0.2.3

5 files

0.2.2

5 files

0.2.1

5 files

0.2.0

5 files

0.1.7

5 files

0.1.6

5 files

0.1.5

5 files

0.1.4

5 files

0.1.3

5 files

0.1.2

5 files

0.1.1

5 files

0.1.0

5 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page