A Very Big Video Reasoning Suite

We bet on a future that video reasoning is the next fundamental intelligence paradigm, after language reasoning, where spatiotemporal embodied world experiences could be more naturally captured.

Data Engines

View All
traffic_light
GitHub
Knowledge in-domain testset
This scene shows a crossroad with four traffic lights (North, South, East, West). Each light independently follows a 3-color cycle: Red (4s) → Yellow (4s) → Green (4s) → Yellow (4s) → Red. Currently: North light is red with 3s countdown, South light is green with 1s countdown, East light is yellow with 1s countdown, West light is yellow with 3s countdown. Simulate 4 seconds and show the final state of all four traffic lights.
First Frame
Last Frame
shape_outline_fill
GitHub
Abstraction in-domain testset
Complete the A:B :: C:? shape-style analogy. Show how the right shape in the second row changes its fill or outline so that it follows the same style transformation used between the first two shapes.
First Frame
Last Frame
grid_color_sequence
GitHub
Spatiality training set
The scene shows a 10x10 grid with a green start point, a red end point, and colored cells (orange, yellow, and blue). A purple circular agent is positioned at the green start point. The agent can move to adjacent cells (up, down, left, right). Starting from the green start point, the agent must visit the colored cells in order (orange, then yellow, then blue), taking the shortest path between each consecutive pair of colored cells. The agent is allowed to pass through the red end point when visiting the colored cells if needed. After visiting all colored cells in sequence, the agent must reach the red end point, also following the shortest path.
First Frame
Last Frame
combined_objects_spinning
GitHub
Transformation training set
The scene shows 2 objects on the left side and dashed target outlines on the right side. The dashed target outlines remain completely stationary. For each object, first rotate it in place to match the orientation of its corresponding dashed target outline, then move it horizontally to the right so that it aligns exactly with and fits within its corresponding dashed target outline.
First Frame
Last Frame
mark_wave_peaks
GitHub
Perception out-of-domain testset
The scene shows a continuous wave on a white background. Find all peaks (local maxima: each point where the wave value is greater than both immediate neighbors). Circle each peak with a red hollow outline and a solid red dot at its center, from left to right one by one, and show the solution step by step.
First Frame
Last Frame

Inference Results

View Full Bench
Domino Chain Gap Analysis - Samples
00
01
02
03
04
Task Domains 1/5
Domino Chain Gap Analysis
Knowledge in-domain testset
Next Figure (Small-Large Alt)
Abstraction out-of-domain testset
Grid Shortest Path
Spatiality in-domain testset
Shape Sorter
Transformation out-of-domain testset
Identify Unique Figure
Perception out-of-domain testset
Prompt
Loading...
Ground Truth
First
First Frame
Final
Final Frame
Model Outputs
1/
VBVR-Wan2.2
VBVR-Wan2.2
CogVideoX 1.5
Kling 2.6
LTX-2
Runway Gen-4
Sora 2
Veo 3
Wan 2.2 I2V
Hunyuan I2V
Seedance 2.0

Leaderboard

Modality
Split
Type
Category
2026-04-28