video-performance-analyzer
Predict how a video lands before you publish it — then find out whether the prediction was right, and get better advice because of it.
vpa runs Meta's TRIBE v2 brain-encoding model over your video, reduces the
output to a response curve you can read, compares it against a reference video
and against your own back catalogue, and asks an LLM for concrete editing
recommendations. When you later record what the video actually did — views,
likes, watch-through — it learns, and goes back to revise advice it has already
given you.
vpa analyse my-cut-v3.mp4 --reference competitor-ad.mp4 \
--segments "hook:0-3,montage:3-16,card:16-19,end:19-24"
╭─ my-cut-v3 ──────────────────────────────────────────────────╮
│ 24.1s · 25 timesteps · Audio, Video │
╰──────────────────────────────────────────────────────────────╯
╭─ predicted response over time ───────────────────────────────╮
│ ▃▄▄▅▅▆▇▇▆▇▇▇▆▅▄▃▃▃▃▂▂▂▂█ │
│ 0s 24s │
╰──────────────────────────────────────────────────────────────╯
opening_2s 0.882 Mean response over the first two seconds.
trough_value 0.777 Lowest normalised response.
decay +0.055 Opening minus closing.
What this actually measures — read this first
TRIBE v2 predicts fMRI brain response. It does not predict attention, watch time, clicks, or sales. Treating "more predicted response" as "better advert" is an inference the model does not make and its authors do not claim.
This matters enough that the tool is built around it:
- Every metric shown has a plain-English explanation attached (
vpa explain). - The same caveats are injected into the LLM's context, so recommendations inherit them rather than overclaiming.
- Known artefacts are flagged automatically. A cut to black spikes the curve
every time;
vpadetects that and excludes it from peak-finding rather than reporting it as your ending landing. - Position matters more than content. In testing, the same images scored highest in one edit and lowest in another purely because of where they sat.
Use it to rank variants of the same idea. Don't use it as a verdict on one video.
Install
pip install video-performance-analyzer # core
pip install 'video-performance-analyzer[tribe]' # + scoring dependencies (~3GB of models)
pip install git+https://github.com/facebookresearch/tribev2.git
TRIBE itself is not on PyPI, so that last line is always needed for scoring.
From source instead
git clone https://github.com/christopher-inegbedion/video-performance-analyzer.git
cd video-performance-analyzer
python3.12 -m venv .venv && source .venv/bin/activate
pip install -e '.[dev]'
You also need ffmpeg (brew install ffmpeg / apt install ffmpeg) and an
API key for whichever LLM you point it at:
export OPENROUTER_API_KEY=sk-or-...
vpa doctor # tells you exactly what is missing
Use
vpa analyse cut.mp4 # score one video
vpa analyse cut.mp4 -r reference.mp4 # compare against something
vpa show eval_a1b2c3 # revisit a past evaluation
vpa list evals # everything you have run
vpa ask eval_a1b2c3 # ask questions about it
vpa explain artefacts # what the numbers can't tell you
vpa tui # interactive session
Describing your structure
Even segments are a poor guide. Tell it where your real beats are and the report speaks your language:
vpa analyse cut.mp4 --segments "hook:0-3,montage:3-16,card:16-19,logo:19-24"
The learning loop
This is what makes the tool improve. Record what a published video did:
vpa metrics add my-cut-v3 --views 12400 --likes 380 --platform instagram
vpa metrics template -o metrics.csv && vpa metrics import metrics.csv
vpa metrics show # what it has learned so far
Once outcomes exist, two things change:
- New recommendations weight real performance above predicted response.
- Old evaluations become stale — their advice predates what you now know.
vpa retrofitrevisits them and writes a new generation of recommendations saying what changed. Nothing is overwritten; you can read every generation withvpa show <id> -g 1.
The tool is deliberately honest about sample size. Below five labelled videos it refuses to call correlations findings and says so in plain terms.
Configuration
vpa config init # writes a commented config file
vpa config show # effective settings and where they came from
Any OpenAI-compatible endpoint works — OpenRouter (default), OpenAI, Together, Groq, or a local model:
[llm]
model = "google/gemini-2.5-flash"
base_url = "https://openrouter.ai/api/v1"
[tribe]
target_fps = 24 # 60fps costs ~2.5x for identical footage
enable_language = false # true needs a gated Llama repo + HF token
[analysis]
ignore_tail_s = 1.5 # don't let the cut-to-black spike become a "finding"
For a fully local setup, point base_url at Ollama (http://localhost:11434/v1)
and no API key is needed.
Running TRIBE on a laptop
TRIBE ships configured for Meta's GPU cluster. vpa patches the known problems
automatically and tells you what it changed:
| Problem | What happens without the fix |
|---|---|
device: cuda baked into the checkpoint |
Torch not compiled with CUDA enabled |
num_workers: 20 baked in |
DataLoader workers die silently; the process sits at 3% CPU looking alive |
compute_type hardcoded to float16 |
CPU speech extraction always fails |
uvx whisperx resolves a broken torchaudio |
module 'torchaudio' has no attribute 'list_audio_backends' |
Language pathway needs gated meta-llama/Llama-3.2-1B |
401 on an otherwise working run |
The last two are why word timings are generated locally with faster-whisper and
written to the .tsv cache TRIBE reads — whisperx is never invoked. The language
pathway is off by default, so vision + audition work with no licence gate.
How long it takes
Scoring is dominated entirely by the video encoder. Setup — importing the package, loading the checkpoint, building events — is about 8 seconds. Everything after that scales with how much footage you feed it.
Measured on an Apple Silicon laptop (CPU only), roughly 96x realtime:
| video length | time to score |
|---|---|
| 10s | ~16 min |
| 15s | ~24 min |
| 24s | ~38 min |
| 60s | ~96 min |
This is not an interactive tool. Start a run and come back to it.
Three ways to make it tractable:
- Halve the frame rate.
target_fps = 12roughly halves the encode, because cost is proportional to frames. V-JEPA samples frames rather than reading every one, so the quality cost is plausibly small — but that is untested, so measure it on your own material before trusting it. - Score an excerpt. For a feed asset the first 6-10 seconds is where the scroll decision happens. Scoring only the opening is a legitimate strategy.
- Use a GPU. This is what the model was built for and it is a different order
of magnitude. Set
device = "cuda".
There is no caching or warm-start trick that helps: the cost is the encoder, not the setup.
Licences
This tool is MIT. The models are not:
- TRIBE v2 is CC BY-NC 4.0 — non-commercial. Using it to optimise commercial advertising is arguably outside that licence. That is your call to make deliberately, and the tool says so rather than hiding it.
- The language pathway uses Llama-3.2-1B, which is gated and carries Meta's own terms.
Development
pip install -e '.[dev]'
pytest -q
ruff check vpa
Contributions welcome — see CONTRIBUTING.md. Platform
connectors for metrics ingestion (vpa/metrics.py has a documented seam) and
additional LLM providers (vpa/providers/) are the most useful places to start.
Metadata
Release files for video-performance-analyzer 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| video_performance_analyzer-0.1.0.tar.gz | 58.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| video_performance_analyzer-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 117.7 kB
Release files / video_performance_analyzer-0.1.0.tar.gz
| Download URL | video_performance_analyzer-0.1.0.tar.gz |
|---|---|
| Size | 58.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4ce180df7100a0cba0d91b47436087ebaa109156975a79098413d910f4366a98
|
|
BLAKE2b-256 checksum How to use checksums |
470ebb5000a88f5a2b0395d90bbdbe24788fe7d9c9bafaa1aa703756ffcec942
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency logRelease files / video_performance_analyzer-0.1.0-py3-none-any.whl
| Download URL | video_performance_analyzer-0.1.0-py3-none-any.whl |
|---|---|
| Size | 59.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
0a18072c8cbb3b7b270ca5a3a52aa87557b8aa9661cf15371e73ec618f5d0138
|
|
BLAKE2b-256 checksum How to use checksums |
ea1674af580c66511e3c7ec24fadc3817414c5397a69087cfb62e19d8ff0b109
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.
Transparency log