PersistBench: Can 4D Foundation Models Remember?
NeurIPS 2026 Evaluations and Datasets (Spotlight)
Guangzhao He, Hadar Averbuch-Elor*, Wei-Chiu Ma* ("*" denotes equal advising)
Overview
PersistBench is a dataset and metric suite for evaluating visual memory in 4D foundation models, including camera-controlled video generators and 4D reconstruction models. It uses 360° videos as reference observations to assess objects after they leave the input camera's view, measuring object permanence, motion continuity, and appearance preservation. See the paper abstract.
This early code release supports your own model and the reference diagnostic.
Release checklist
- Custom model inference and evaluation
- Dataset download workflow
- PyPI publication
- Release code for baselines in the paper
Setup
Use Python 3.10+ in your model environment. Replace /path/to/... and your-model-env with your own paths and Conda environment name.
- Clone PersistBench into your workspace:
git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
- Activate your model's existing Conda environment, then install from the PersistBench repo:
conda activate your-model-env cd /path/to/PersistBench python -m pip install -e .
Alternatively, install from PyPI in your model environment:python -m pip install persistbench.
PyPI only (0.1.1+): run persistbench docs for the complete workflow, or
persistbench docs --output /path/to/persistbench-workspace to export the guides,
model example, and Qwen launcher. Follow the
PyPI guide without cloning this repo.
Dataset
Prepare the dataset separately at /data/benchmark. The download workflow is coming;
currently use the dataset format.
Inference
- Go to your model's repo, keeping your model environment active:
cd /path/to/your-model
- Create
my_model.pyin that repo. Replace the loading and generation calls below with your model's API; set its frame count and resolution:from persistbench import BaseModel, ModelOutput class MyModel(BaseModel): num_frames = 49 resolution = (384, 512) # height, width; -1 preserves dataset shape def __init__(self, device="cuda:0"): super().__init__(device=device) self.model = load_your_model(device=device) def predict(self, sample): frames = self.model.generate(sample.input_frames, sample.input_cameras, sample.target_cameras) return ModelOutput(frames) # CPU NumPy RGB [T, H, W, 3]
See the executable example and interface details. - Run inference from your model's repo, in the same model environment:
persistbench infer --model ./my_model.py:MyModel --name my-model --dataset /data/benchmark
Predictions are saved to/path/to/your-model/outputs/null/my-model/.
Evaluation
- Create the evaluation Conda environment. Run from the PersistBench repo:
cd /path/to/PersistBench conda create -n persistbench-eval python=3.11 pip -y conda activate persistbench-eval python -m pip install '.[eval]'
- Install the metric models and download their checkpoints, in this same environment and repo: follow the complete checkpoint guide (SAM2, DINOv3, DINOv2).
- Start or connect to Qwen. Follow the independent Qwen judge guide. Hosting uses a separate Conda environment and terminal; keep the evaluation terminal open.
- Evaluate and aggregate. In the evaluation terminal, from the PersistBench repo:
cd /path/to/PersistBench conda activate persistbench-eval export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt" export DINOV3_REPO="$PWD/vendor/dinov3" export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth" export PERSISTBENCH_VLM_BASE_URL=http://127.0.0.1:8000/v1 persistbench evaluate --dataset /data/benchmark --predictions /path/to/your-model/outputs/null/my-model persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
For a remote Qwen server, replace the endpoint above with its URL.
Read /path/to/your-model/outputs/persist_bench/my-model/summary.json for
static/dynamic × visible/invisible averages and valid-case counts.
Evaluation saves final scores; aggregation only averages them.
See score definitions for formulas and missing-score handling.
Commands use the active Conda environment; they do not switch environments.
Visualizations and progress
Per-case evaluation folders include matching images, tracked-mask overlays, object crops, and judge inputs. From the PersistBench repo, in the evaluation environment:
persistbench dashboard --output /path/to/your-model/outputs --port 8080
Open http://127.0.0.1:8080. For splitting, resuming, and configuration, see
usage; for checks and limitations, see validation.
Acknowledgements
We thank the authors of SAM2, DINOv3, DINOv2, and Qwen for their open-source projects.
License
This project is licensed under the MIT License.
Citation
If you use PersistBench in your research, please cite:
@misc{he2026persistbench,
title={Can {4D} Foundation Models Remember?},
author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
year={2026},
eprint={2609.20819},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2609.20819}
}
Metadata
Release files for persistbench 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| persistbench-0.1.1.tar.gz | 779.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| persistbench-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 924.3 kB
Release files / persistbench-0.1.1.tar.gz
| Download URL | persistbench-0.1.1.tar.gz |
|---|---|
| Size | 779.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
cb219e3fdb17d51bc656e238a1e06b6852ab4e0b6e15ba274f7d9c89f455c687
|
|
BLAKE2b-256 checksum How to use checksums |
22c30e4e382b049663773329ca4e08319270971fdc9cee4918fc3b2aedb7c4f7
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|
Release files / persistbench-0.1.1-py3-none-any.whl
| Download URL | persistbench-0.1.1-py3-none-any.whl |
|---|---|
| Size | 145.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
22cfade6d03025d9d1621600b2619890d5bb2ac05978a9593478608bdd3c9a87
|
|
BLAKE2b-256 checksum How to use checksums |
0582910a5ed826ca7168a3716804bd767b1b73c0f16297ec510f5f2adaf711f6
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.16
|