Skip to main content

PersistBench: Can 4D Foundation Models Remember?

arXiv PersistBench Video PersistBench Website Leaderboard Dataset Code Visitors

NeurIPS 2026 Evaluations and Datasets (Spotlight)

Guangzhao He, Hadar Averbuch-Elor*, Wei-Chiu Ma* ("*" denotes equal advising)

Overview

PersistBench is a dataset and metric suite for evaluating visual memory in 4D foundation models, including camera-controlled video generators and 4D reconstruction models. It uses 360° videos as reference observations to assess objects after they leave the input camera's view, measuring object permanence, motion continuity, and appearance preservation. See the paper abstract.

This early code release supports your own model and the reference diagnostic.

Release checklist

  • Custom model inference and evaluation
  • Dataset download workflow
  • PyPI publication
  • Release code for baselines in the paper

Setup

Use Python 3.10+ in your model environment. Replace /path/to/... and your-model-env with your own paths and Conda environment name.

  1. Clone PersistBench into your workspace:
    git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
    
  2. Activate your model's existing Conda environment, then install from the PersistBench repo:
    conda activate your-model-env
    cd /path/to/PersistBench
    python -m pip install -e .
    
    Alternatively, install from PyPI in your model environment: python -m pip install persistbench.

Dataset

Prepare the dataset separately at /data/benchmark. The download workflow is coming; currently use the dataset format.

Inference

  1. Go to your model's repo, keeping your model environment active:
    cd /path/to/your-model
    
  2. Create my_model.py in that repo. Replace the loading and generation calls below with your model's API; set its frame count and resolution:
    from persistbench import BaseModel, ModelOutput
    
    class MyModel(BaseModel):
        num_frames = 49
        resolution = (384, 512)  # height, width; -1 preserves dataset shape
    
        def __init__(self, device="cuda:0"):
            super().__init__(device=device)
            self.model = load_your_model(device=device)
    
        def predict(self, sample):
            frames = self.model.generate(sample.input_frames,
                                         sample.input_cameras, sample.target_cameras)
            return ModelOutput(frames)  # CPU NumPy RGB [T, H, W, 3]
    
    See the executable example and interface details.
  3. Run inference from your model's repo, in the same model environment:
    persistbench infer --model ./my_model.py:MyModel --name my-model --dataset /data/benchmark
    
    Predictions are saved to /path/to/your-model/outputs/null/my-model/.

Evaluation

  1. Create the evaluation Conda environment. Run from the PersistBench repo:
    cd /path/to/PersistBench
    conda create -n persistbench-eval python=3.11 pip -y
    conda activate persistbench-eval
    python -m pip install '.[eval]'
    
  2. Install the metric models and download their checkpoints, in this same environment and repo: follow the complete checkpoint guide (SAM2, DINOv3, DINOv2).
  3. Start or connect to Qwen. Follow the independent Qwen judge guide. Hosting uses a separate Conda environment and terminal; keep the evaluation terminal open.
  4. Evaluate and aggregate. In the evaluation terminal, from the PersistBench repo:
    cd /path/to/PersistBench
    conda activate persistbench-eval
    export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt"
    export DINOV3_REPO="$PWD/vendor/dinov3"
    export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth"
    export PERSISTBENCH_VLM_BASE_URL=http://127.0.0.1:8000/v1
    persistbench evaluate --dataset /data/benchmark --predictions /path/to/your-model/outputs/null/my-model
    persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
    
    For a remote Qwen server, replace the endpoint above with its URL.

Read /path/to/your-model/outputs/persist_bench/my-model/summary.json for static/dynamic × visible/invisible averages and valid-case counts. Evaluation saves final scores; aggregation only averages them. See score definitions for formulas and missing-score handling. Commands use the active Conda environment; they do not switch environments.

Visualizations and progress

Per-case evaluation folders include matching images, tracked-mask overlays, object crops, and judge inputs. From the PersistBench repo, in the evaluation environment:

persistbench dashboard --output /path/to/your-model/outputs --port 8080

Open http://127.0.0.1:8080. For splitting, resuming, and configuration, see usage; for checks and limitations, see validation.

Acknowledgements

We thank the authors of SAM2, DINOv3, DINOv2, and Qwen for their open-source projects.

License

This project is licensed under the MIT License.

Citation

If you use PersistBench in your research, please cite:

@misc{he2026persistbench,
  title={Can {4D} Foundation Models Remember?},
  author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
  year={2026},
  eprint={2609.20819},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.20819}
}

Metadata

Release files for persistbench 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for persistbench 0.1.0
File Size Uploaded
persistbench-0.1.0.tar.gz 775.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for persistbench 0.1.0
File Interpreter ABI Platform
persistbench-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 889.4 kB

Release files / persistbench-0.1.0.tar.gz

Download URL persistbench-0.1.0.tar.gz
Size 775.3 kB
Tags Source
SHA-256 checksum
How to use checksums
a5a9fad5341fbd7f1982141fcfdbb93f30db93e4ffe06ba73fb0659a77c44752
BLAKE2b-256 checksum
How to use checksums
16b176f352c2400102a8543a2c6e17179939c47973fe5e454656561586aca26d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / persistbench-0.1.0-py3-none-any.whl

Download URL persistbench-0.1.0-py3-none-any.whl
Size 114.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7013efc639f0864ab0191357191e5c1f5d279ff435d5c6d94d202d6a951cfe99
BLAKE2b-256 checksum
How to use checksums
08678ef55961696aab25436933cbcb5d5b653f58c011ef92cfb66fbec0b08597
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.2

2 release files

0.1.1

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page