PersistBench: Can 4D Foundation Models Remember?
NeurIPS 2026 Evaluations and Datasets (Spotlight)
Guangzhao He, Hadar Averbuch-Elor*, Wei-Chiu Ma* ("*" denotes equal advising)
Overview
PersistBench is a dataset and evaluation suite for benchmarking visual memory in 4D foundation models, including camera-controlled video generators, video-to-360° video generators, and 4D reconstruction models. It uses 360° videos as reference observations to evaluation objects after they leave the input camera's view, measuring object permanence, motion continuity, and appearance preservation.
Release checklist
- Custom model inference and evaluation
- Dataset download workflow
- PyPI publication
- Release code for baselines in the paper
Setup
Inference supports Python 3.9+ in your own model environment. There are two options to install the evaluation suite. After installation, both routes share the same workflow.
Option 1: Install from PyPI
conda activate your-model-env # replace with your model's inference environment
python -m pip install 'persistbench>=0.1.4'
persistbench docs --output /path/to/PersistBench # save the documents to this directory
The export creates a new workspace containing these guides, the model example, and the Qwen launcher.
Option 2: Clone and install
git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
conda activate your-model-env # replace with your model's inference environment
cd /path/to/PersistBench
python -m pip install -e .
Dataset
We host our evaluation dataset metadata and reconstruction script on Hugging Face.
Create a separate dataset environment, from your PersistBench workspace:
cd /path/to/PersistBench
conda create -n persistbench-data python=3.12 pip -y
conda activate persistbench-data
conda install -c conda-forge 'nodejs>=22' -y
Install the dataset dependencies using your chosen route:
| Clone | PyPI |
|---|---|
python -m pip install -e '.[data]' |
python -m pip install 'persistbench[data]>=0.1.4' |
Choose one download option, in the dataset environment, from your PersistBench workspace:
Option 1: Local YouTube download (yt-dlp)
persistbench dataset download --output ./data/persistbench_data
Option 2: Google Cloud download
Install the Google Cloud CLI, then run and follow the Google sign-in and project-selection prompts:
persistbench dataset download --output ./data/persistbench_data --backend cloud
Source videos download on a Cloud VM; reconstruction runs locally. See Cloud setup for requirements and VM management.
Both options save data to ./data/persistbench_data and register its
absolute path for all your Conda environments. Inference and evaluation require
registered data. See dataset setup for resuming and small test runs.
Inference
- Switch back to your model environment and go to your model's repo:
conda activate your-model-env cd /path/to/your-model
- Create
my_model.pyin that repo. Replace the loading and generation calls below with your model's API; set its frame count and resolution:from persistbench import BaseModel, ModelOutput class MyModel(BaseModel): num_frames = 49 resolution = (384, 512) # height, width; -1 preserves dataset shape def __init__(self, device="cuda:0"): super().__init__(device=device) self.model = load_your_model(device=device) def predict(self, sample): frames = self.model.generate(sample.input_frames, sample.input_cameras, sample.target_cameras) return ModelOutput(frames) # CPU NumPy RGB [T, H, W, 3]
See the executable example and interface details. - Run inference from your model's repo, in the same model environment:
persistbench infer --model ./my_model.py:MyModel --name my-model
Predictions are saved to/path/to/your-model/outputs/null/my-model/.
Evaluation
-
Create the evaluation Conda environment. Run from your PersistBench workspace:
cd /path/to/PersistBench conda create -n persistbench-eval python=3.11 pip -y conda activate persistbench-eval
Install the evaluation dependencies using your chosen route:
Clone PyPI python -m pip install '.[eval]'python -m pip install 'persistbench[eval]>=0.1.4' -
Install the metric models and download their checkpoints, in this same environment and workspace: follow the complete checkpoint guide (SAM2, DINOv3, DINOv2).
-
Configure Qwen using either option in the Qwen judge guide: Alibaba Cloud / DashScope API with your own key (
qwen3.8-27b), or self-hosting the pinned BF16 checkpoint (development setup: 2 × 48 GB A6000; one 48 GB GPU is insufficient). Self-hosting uses a separate Conda environment and terminal. -
Evaluate and aggregate. In the evaluation terminal, from your PersistBench workspace:
cd /path/to/PersistBench conda activate persistbench-eval export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt" export DINOV3_REPO="$PWD/vendor/dinov3" export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth" persistbench evaluate --predictions /path/to/your-model/outputs/null/my-model persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
Keep the Qwen variables configured in step 3 in this terminal.
Read /path/to/your-model/outputs/persist_bench/my-model/summary.json for
static/dynamic × visible/invisible averages and valid-case counts.
Evaluation saves final scores; aggregation only averages them.
See score definitions for formulas and missing-score handling.
Commands use the active Conda environment; they do not switch environments.
Leaderboard submission
Follow the submission guide to package the evaluation report from all 2,000 cases and upload it to the leaderboard. Packaging uses the files already produced by evaluation and aggregation:
persistbench package \
--results /path/to/your-model/outputs/persist_bench/my-model \
--output submission.zip
This command is included in the PyPI package; no separate helper script is needed.
Visualizations and progress
Per-case evaluation folders include matching images, tracked-mask overlays, object crops, and judge inputs. From your PersistBench workspace, in the evaluation environment:
persistbench dashboard --output /path/to/your-model/outputs --port 8080
Open http://127.0.0.1:8080. For splitting, resuming, and configuration, see
usage; for checks and limitations, see validation.
Acknowledgements
We thank the authors of SAM2, DINOv3, DINOv2, and Qwen for their open-source projects.
License
This project is licensed under the MIT License.
Maintainers: automatic PyPI releases.
Citation
If you use PersistBench in your research, please cite:
@misc{he2026persistbench,
title={Can {4D} Foundation Models Remember?},
author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
year={2026},
eprint={2609.20819},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2609.20819}
}
Metadata
Release files for persistbench 0.1.4
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| persistbench-0.1.4.tar.gz | 796.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| persistbench-0.1.4-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.5 MB
Release files / persistbench-0.1.4.tar.gz
| Download URL | persistbench-0.1.4.tar.gz |
|---|---|
| Size | 796.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
0013a0e7fa44ca0dd8f13c6e708fecf8c67f35a0a2454420465fe6ffd229f061
|
|
BLAKE2b-256 checksum How to use checksums |
cf836b3fb2cab11df5fa93b00bd7838a5088efd819b005f0ed0805e502749c89
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / persistbench-0.1.4-py3-none-any.whl
| Download URL | persistbench-0.1.4-py3-none-any.whl |
|---|---|
| Size | 688.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
541483a09dd43253371cb9a31468dbb370ef988f1b090bc6c84805a6ed53d683
|
|
BLAKE2b-256 checksum How to use checksums |
7375d83e01d4493ad98f815db93f612aa501a5f68facd19d0984185221dc4a33
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log