Skip to main content

PersistBench: Can 4D Foundation Models Remember?

arXiv PersistBench Video PersistBench Website Hugging Face Leaderboard Leaderboard Submission PyPI Hugging Face Dataset Dataset Code Visitors

NeurIPS 2026 Evaluations and Datasets (Spotlight)

Guangzhao He, Hadar Averbuch-Elor*, Wei-Chiu Ma* ("*" denotes equal advising)

Overview

PersistBench is a dataset and evaluation suite for benchmarking visual memory in 4D foundation models, including camera-controlled video generators, video-to-360° video generators, and 4D reconstruction models. It uses 360° videos as reference observations to evaluation objects after they leave the input camera's view, measuring object permanence, motion continuity, and appearance preservation.

Release checklist

  • Custom model inference and evaluation
  • Dataset download workflow
  • PyPI publication
  • Release code for baselines in the paper

Setup

Inference supports Python 3.9+ in your own model environment. There are two options to install the evaluation suite. After installation, both routes share the same workflow.

Option 1: Install from PyPI

conda activate your-model-env  # replace with your model's inference environment
python -m pip install 'persistbench>=0.1.4'
persistbench docs --output /path/to/PersistBench  # save the documents to this directory

The export creates a new workspace containing these guides, the model example, and the Qwen launcher.

Option 2: Clone and install

git clone https://github.com/guangzhaohe/PersistBench.git /path/to/PersistBench
conda activate your-model-env  # replace with your model's inference environment
cd /path/to/PersistBench
python -m pip install -e .

Dataset

We host our evaluation dataset metadata and reconstruction script on Hugging Face.

Create a separate dataset environment, from your PersistBench workspace:

cd /path/to/PersistBench
conda create -n persistbench-data python=3.12 pip -y
conda activate persistbench-data
conda install -c conda-forge 'nodejs>=22' -y

Install the dataset dependencies using your chosen route:

Clone PyPI
python -m pip install -e '.[data]' python -m pip install 'persistbench[data]>=0.1.4'

Choose one download option, in the dataset environment, from your PersistBench workspace:

Option 1: Local YouTube download (yt-dlp)

persistbench dataset download --output ./data/persistbench_data

Option 2: Google Cloud download

Install the Google Cloud CLI, then run and follow the Google sign-in and project-selection prompts:

persistbench dataset download --output ./data/persistbench_data --backend cloud

Source videos download on a Cloud VM; reconstruction runs locally. See Cloud setup for requirements and VM management.

Both options save data to ./data/persistbench_data and register its absolute path for all your Conda environments. Inference and evaluation require registered data. See dataset setup for resuming and small test runs.

Inference

  1. Switch back to your model environment and go to your model's repo:
    conda activate your-model-env
    cd /path/to/your-model
    
  2. Create my_model.py in that repo. Replace the loading and generation calls below with your model's API; set its frame count and resolution:
    from persistbench import BaseModel, ModelOutput
    
    class MyModel(BaseModel):
        num_frames = 49
        resolution = (384, 512)  # height, width; -1 preserves dataset shape
    
        def __init__(self, device="cuda:0"):
            super().__init__(device=device)
            self.model = load_your_model(device=device)
    
        def predict(self, sample):
            frames = self.model.generate(sample.input_frames,
                                         sample.input_cameras, sample.target_cameras)
            return ModelOutput(frames)  # CPU NumPy RGB [T, H, W, 3]
    
    See the executable example and interface details.
  3. Run inference from your model's repo, in the same model environment:
    persistbench infer --model ./my_model.py:MyModel --name my-model
    
    Predictions are saved to /path/to/your-model/outputs/null/my-model/.

Evaluation

  1. Create the evaluation Conda environment. Run from your PersistBench workspace:

    cd /path/to/PersistBench
    conda create -n persistbench-eval python=3.11 pip -y
    conda activate persistbench-eval
    

    Install the evaluation dependencies using your chosen route:

    Clone PyPI
    python -m pip install '.[eval]' python -m pip install 'persistbench[eval]>=0.1.4'
  2. Install the metric models and download their checkpoints, in this same environment and workspace: follow the complete checkpoint guide (SAM2, DINOv3, DINOv2).

  3. Configure Qwen using either option in the Qwen judge guide: Alibaba Cloud / DashScope API with your own key (qwen3.8-27b), or self-hosting the pinned BF16 checkpoint (development setup: 2 × 48 GB A6000; one 48 GB GPU is insufficient). Self-hosting uses a separate Conda environment and terminal.

  4. Evaluate and aggregate. In the evaluation terminal, from your PersistBench workspace:

    cd /path/to/PersistBench
    conda activate persistbench-eval
    export SAM2_CHECKPOINT="$PWD/checkpoints/sam2.1_hiera_large.pt"
    export DINOV3_REPO="$PWD/vendor/dinov3"
    export DINOV3_CHECKPOINT="$PWD/checkpoints/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth"
    persistbench evaluate --predictions /path/to/your-model/outputs/null/my-model
    persistbench aggregate --results /path/to/your-model/outputs/persist_bench/my-model
    

    Keep the Qwen variables configured in step 3 in this terminal.

Read /path/to/your-model/outputs/persist_bench/my-model/summary.json for static/dynamic × visible/invisible averages and valid-case counts. Evaluation saves final scores; aggregation only averages them. See score definitions for formulas and missing-score handling. Commands use the active Conda environment; they do not switch environments.

Leaderboard submission

Follow the submission guide to package the evaluation report from all 2,000 cases and upload it to the leaderboard. Packaging uses the files already produced by evaluation and aggregation:

persistbench package \
  --results /path/to/your-model/outputs/persist_bench/my-model \
  --output submission.zip

This command is included in the PyPI package; no separate helper script is needed.

Visualizations and progress

Per-case evaluation folders include matching images, tracked-mask overlays, object crops, and judge inputs. From your PersistBench workspace, in the evaluation environment:

persistbench dashboard --output /path/to/your-model/outputs --port 8080

Open http://127.0.0.1:8080. For splitting, resuming, and configuration, see usage; for checks and limitations, see validation.

Acknowledgements

We thank the authors of SAM2, DINOv3, DINOv2, and Qwen for their open-source projects.

License

This project is licensed under the MIT License.

Maintainers: automatic PyPI releases.

Citation

If you use PersistBench in your research, please cite:

@misc{he2026persistbench,
  title={Can {4D} Foundation Models Remember?},
  author={He, Guangzhao and Averbuch-Elor, Hadar and Ma, Wei-Chiu},
  year={2026},
  eprint={2609.20819},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.20819}
}

Metadata

Release files for persistbench 0.1.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for persistbench 0.1.6
File Size Uploaded
persistbench-0.1.6.tar.gz 796.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for persistbench 0.1.6
File Interpreter ABI Platform
persistbench-0.1.6-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / persistbench-0.1.6.tar.gz

Download URL persistbench-0.1.6.tar.gz
Size 796.8 kB
Tags Source
SHA-256 checksum
How to use checksums
cdda24c41fac38221d01386702c450fa8b7ee7ec3ea2aabab844b8033cede3ba
BLAKE2b-256 checksum
How to use checksums
34ad87f04e34e79a1ab42d8b6a4d4c68e555b8306b7cc8eb7e5a565eb9d8f5b3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / persistbench-0.1.6-py3-none-any.whl

Download URL persistbench-0.1.6-py3-none-any.whl
Size 688.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6af1e48d0599eac140b21ebcfabefbe2093997dac7d3e8f68d0f404efd2a2d1d
BLAKE2b-256 checksum
How to use checksums
90ade97b5301196c33a5609249588d141b262437aa28823ab8299bded26b598a
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.7

2 release files

This release

0.1.6 This release

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page