Skip to main content

PersistBench: Can 4D Foundation Models Remember?

arXiv PersistBench Video PersistBench Website Leaderboard Dataset Code Visitors

NeurIPS 2026 Evaluations and Datasets (Spotlight)

Guangzhao He, Hadar Averbuch-Elor*, Wei-Chiu Ma* ("*" denotes equal advising)

Overview

PersistBench is a dataset and metric suite for evaluating visual memory in 4D foundation models, including camera-controlled video generators and 4D reconstruction models. It uses 360° videos as reference observations to assess objects after they leave the input camera's view, measuring object permanence, motion continuity, and appearance preservation. See the paper abstract.

This early code release supports your own model and the reference diagnostic.

Release checklist

  • Custom model inference and evaluation
  • Dataset download workflow
  • PyPI publication
  • Release code for baselines in the paper

Setup

Use Python 3.10+ in your model environment. Replace /path/to/... and your-model-env with your own paths and Conda environment name.

  1. Clone PersistBench into your workspace:
    git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
    
  2. Activate your model's existing Conda environment, then install from the PersistBench repo:
    conda activate your-model-env
    cd /path/to/PersistBench
    python -m pip install -e .
    
    Alternatively, install from PyPI in your model environment: python -m pip install persistbench.

PyPI only (0.1.1+): run persistbench docs for the complete workflow, or persistbench docs --output /path/to/persistbench-workspace to export the guides, model example, and Qwen launcher. Follow the PyPI guide without cloning this repo.

Dataset

Prepare the dataset in ./data inside the PersistBench repo (/path/to/PersistBench/data). You can store it elsewhere and pass that path with --dataset; relative paths are resolved from the directory where you run the command. The download workflow is coming; currently use the dataset format.

Inference

  1. Go to your model's repo, keeping your model environment active:
    cd /path/to/your-model
    
  2. Create my_model.py in that repo. Replace the loading and generation calls below with your model's API; set its frame count and resolution:
    from persistbench import BaseModel, ModelOutput
    
    class MyModel(BaseModel):
        num_frames = 49
        resolution = (384, 512)  # height, width; -1 preserves dataset shape
    
        def __init__(self, device="cuda:0"):
            super().__init__(device=device)
            self.model = load_your_model(device=device)
    
        def predict(self, sample):
            frames = self.model.generate(sample.input_frames,
                                         sample.input_cameras, sample.target_cameras)
            return ModelOutput(frames)  # CPU NumPy RGB [T, H, W, 3]
    
    See the executable example and interface details.
  3. Run inference from your model's repo, in the same model environment:
    persistbench infer --model ./my_model.py:MyModel --name my-model --dataset /path/to/PersistBench/data
    
    Predictions are saved to /path/to/your-model/outputs/null/my-model/.

Evaluation

  1. Create the evaluation Conda environment. Run from the PersistBench repo:
    cd /path/to/PersistBench
    conda create -n persistbench-eval python=3.11 pip -y
    conda activate persistbench-eval
    python -m pip install '.[eval]'
    
  2. Install the metric models and download their checkpoints, in this same environment and repo: follow the complete checkpoint guide (SAM2, DINOv3, DINOv2).
  3. Start or connect to Qwen. Follow the independent Qwen judge guide. Hosting uses a separate Conda environment and terminal; keep the evaluation terminal open.
  4. Evaluate and aggregate. In the evaluation terminal, from the PersistBench repo:
    cd /path/to/PersistBench
    conda activate persistbench-eval
    export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt"
    export DINOV3_REPO="$PWD/vendor/dinov3"
    export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth"
    export PERSISTBENCH_VLM_BASE_URL=http://127.0.0.1:8000/v1
    persistbench evaluate --dataset ./data --predictions /path/to/your-model/outputs/null/my-model
    persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
    
    For a remote Qwen server, replace the endpoint above with its URL.

Read /path/to/your-model/outputs/persist_bench/my-model/summary.json for static/dynamic × visible/invisible averages and valid-case counts. Evaluation saves final scores; aggregation only averages them. See score definitions for formulas and missing-score handling. Commands use the active Conda environment; they do not switch environments.

Visualizations and progress

Per-case evaluation folders include matching images, tracked-mask overlays, object crops, and judge inputs. From the PersistBench repo, in the evaluation environment:

persistbench dashboard --output /path/to/your-model/outputs --port 8080

Open http://127.0.0.1:8080. For splitting, resuming, and configuration, see usage; for checks and limitations, see validation.

Acknowledgements

We thank the authors of SAM2, DINOv3, DINOv2, and Qwen for their open-source projects.

License

This project is licensed under the MIT License.

Citation

If you use PersistBench in your research, please cite:

@misc{he2026persistbench,
  title={Can {4D} Foundation Models Remember?},
  author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
  year={2026},
  eprint={2609.20819},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.20819}
}

Metadata

Release files for persistbench 0.1.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for persistbench 0.1.2
File Size Uploaded
persistbench-0.1.2.tar.gz 778.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for persistbench 0.1.2
File Interpreter ABI Platform
persistbench-0.1.2-py3-none-any.whl Python 3 none any Details

Total release size: 924.3 kB

Release files / persistbench-0.1.2.tar.gz

Download URL persistbench-0.1.2.tar.gz
Size 778.9 kB
Tags Source
SHA-256 checksum
How to use checksums
9bea9a3fc71e0471be7e839716adcd8bebaa084e8953f34af1d75e88b8c1042b
BLAKE2b-256 checksum
How to use checksums
8ca08fcf8d6e2dfee4667426a4200c3f035417dc73cf745bb92a8869daf9882b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release files / persistbench-0.1.2-py3-none-any.whl

Download URL persistbench-0.1.2-py3-none-any.whl
Size 145.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
d58388e673c4fc5124b150e5bca23ff7d49f42696a6dcf9343f49c207dc0cee1
BLAKE2b-256 checksum
How to use checksums
f74102887aea60b2717bdb6c6bbe37a04d1dd2ebab4f61fe58534ab83d978810
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.16

Release history Release notifications | RSS feed

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

This release

0.1.2 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page