Skip to main content

PersistBench: Can 4D Foundation Models Remember?

arXiv PersistBench Video PersistBench Website Hugging Face Leaderboard Leaderboard Submission PyPI Hugging Face Dataset Dataset Code Visitors

NeurIPS 2026 Evaluations and Datasets (Spotlight)

Guangzhao He, Hadar Averbuch-Elor*, Wei-Chiu Ma* ("*" denotes equal advising)

Overview

PersistBench is a dataset and evaluation suite for benchmarking visual memory in 4D foundation models, including camera-controlled video generators, video-to-360° video generators, and 4D reconstruction models. It uses 360° videos as reference observations to evaluation objects after they leave the input camera's view, measuring object permanence, motion continuity, and appearance preservation.

Release checklist

  • Custom model inference and evaluation
  • Dataset download workflow
  • PyPI publication
  • Release code for baselines in the paper

Setup

Inference supports Python 3.9+ in your own model environment. There are two options to install the evaluation suite. After installation, both routes share the same workflow.

Option 1: Install from PyPI

conda activate your-model-env  # replace with your model's inference environment
python -m pip install 'persistbench>=0.1.4'
persistbench docs --output /path/to/PersistBench  # save the documents to this directory

The export creates a new workspace containing these guides, the model example, and the Qwen launcher.

Option 2: Clone and install

git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
conda activate your-model-env  # replace with your model's inference environment
cd /path/to/PersistBench
python -m pip install -e .

Dataset

We host our evaluation dataset metadata and reconstruction script on Hugging Face.

Create a separate dataset environment, from your PersistBench workspace:

cd /path/to/PersistBench
conda create -n persistbench-data python=3.12 pip -y
conda activate persistbench-data
conda install -c conda-forge 'nodejs>=22' -y

Install the dataset dependencies using your chosen route:

Clone PyPI
python -m pip install -e '.[data]' python -m pip install 'persistbench[data]>=0.1.4'

Choose one download option, in the dataset environment, from your PersistBench workspace:

Option 1: Local YouTube download (yt-dlp)

persistbench dataset download --output ./data/persistbench_data

Option 2: Google Cloud download

Install the Google Cloud CLI, then run and follow the Google sign-in and project-selection prompts:

persistbench dataset download --output ./data/persistbench_data --backend cloud

Source videos download on a Cloud VM; reconstruction runs locally. See Cloud setup for requirements and VM management.

Both options save data to ./data/persistbench_data and register its absolute path for all your Conda environments. Inference and evaluation require registered data. See dataset setup for resuming and small test runs.

Inference

  1. Switch back to your model environment and go to your model's repo:
    conda activate your-model-env
    cd /path/to/your-model
    
  2. Create my_model.py in that repo. Replace the loading and generation calls below with your model's API; set its frame count and resolution:
    from persistbench import BaseModel, ModelOutput
    
    class MyModel(BaseModel):
        num_frames = 49
        resolution = (384, 512)  # height, width; -1 preserves dataset shape
    
        def __init__(self, device="cuda:0"):
            super().__init__(device=device)
            self.model = load_your_model(device=device)
    
        def predict(self, sample):
            frames = self.model.generate(sample.input_frames,
                                         sample.input_cameras, sample.target_cameras)
            return ModelOutput(frames)  # CPU NumPy RGB [T, H, W, 3]
    
    See the executable example and interface details.
  3. Run inference from your model's repo, in the same model environment:
    persistbench infer --model ./my_model.py:MyModel --name my-model
    
    Predictions are saved to /path/to/your-model/outputs/null/my-model/.

Evaluation

  1. Create the evaluation Conda environment. Run from your PersistBench workspace:

    cd /path/to/PersistBench
    conda create -n persistbench-eval python=3.11 pip -y
    conda activate persistbench-eval
    

    Install the evaluation dependencies using your chosen route:

    Clone PyPI
    python -m pip install '.[eval]' python -m pip install 'persistbench[eval]>=0.1.4'
  2. Install the metric models and download their checkpoints, in this same environment and workspace: follow the complete checkpoint guide (SAM2, DINOv3, DINOv2).

  3. Configure Qwen using either option in the Qwen judge guide: Alibaba Cloud / DashScope API with your own key (qwen3.8-27b), or self-hosting the pinned BF16 checkpoint (development setup: 2 × 48 GB A6000; one 48 GB GPU is insufficient). Self-hosting uses a separate Conda environment and terminal.

  4. Evaluate and aggregate. In the evaluation terminal, from your PersistBench workspace:

    cd /path/to/PersistBench
    conda activate persistbench-eval
    export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt"
    export DINOV3_REPO="$PWD/vendor/dinov3"
    export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth"
    persistbench evaluate --predictions /path/to/your-model/outputs/null/my-model
    persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
    

    Keep the Qwen variables configured in step 3 in this terminal.

Read /path/to/your-model/outputs/persist_bench/my-model/summary.json for static/dynamic × visible/invisible averages and valid-case counts. Evaluation saves final scores; aggregation only averages them. See score definitions for formulas and missing-score handling. Commands use the active Conda environment; they do not switch environments.

Leaderboard submission

Follow the submission guide to package the evaluation report from all 2,000 cases and upload it to the leaderboard. Packaging uses the files already produced by evaluation and aggregation:

persistbench package \
  --results /path/to/your-model/outputs/persist_bench/my-model \
  --output submission.zip

This command is included in the PyPI package; no separate helper script is needed.

Visualizations and progress

Per-case evaluation folders include matching images, tracked-mask overlays, object crops, and judge inputs. From your PersistBench workspace, in the evaluation environment:

persistbench dashboard --output /path/to/your-model/outputs --port 8080

Open http://127.0.0.1:8080. For splitting, resuming, and configuration, see usage; for checks and limitations, see validation.

Acknowledgements

We thank the authors of SAM2, DINOv3, DINOv2, and Qwen for their open-source projects.

License

This project is licensed under the MIT License.

Maintainers: automatic PyPI releases.

Citation

If you use PersistBench in your research, please cite:

@misc{he2026persistbench,
  title={Can {4D} Foundation Models Remember?},
  author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
  year={2026},
  eprint={2609.20819},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.20819}
}

Metadata

Release files for persistbench 0.1.7

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for persistbench 0.1.7
File Size Uploaded
persistbench-0.1.7.tar.gz 796.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for persistbench 0.1.7
File Interpreter ABI Platform
persistbench-0.1.7-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / persistbench-0.1.7.tar.gz

Download URL persistbench-0.1.7.tar.gz
Size 796.8 kB
Tags Source
SHA-256 checksum
How to use checksums
b192c8e04744c898ea7db9fac38b0c838a10646a505740d7d8d7bb5b5f205d34
BLAKE2b-256 checksum
How to use checksums
00ab42070ad78ca088b0cec3b456ebacf02a0c3c59af355746de672759ef7b05
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / persistbench-0.1.7-py3-none-any.whl

Download URL persistbench-0.1.7-py3-none-any.whl
Size 688.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6a6c0d05160297e4f57b9ac9e06a6d592a0865778111912a24d3aa19f5125241
BLAKE2b-256 checksum
How to use checksums
24cf4b307558cc67fc47091789dcf9626b2c749e86eeb1f0a304ff04d78cb2e8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.7 This release

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page