A Very Big Video Reasoning Suite

We bet on a future that video reasoning is the next fundamental intelligence paradigm, after language reasoning, where spatiotemporal embodied world experiences could be more naturally captured.

Data Engines

View All
fluid_diffusion_reasoning
GitHub
Knowledge training set
An ink droplet falls from above the center of a glass beaker filled with water. Upon entering the water, the ink forms irregular downward-extending tendrils due to gravity and initial impact. The ink then diffuses through the water, creating swirling patterns and eddies, until it reaches a stable state of uniform color distribution throughout the entire volume of water.
First Frame
Last Frame
raven
GitHub
Abstraction out-of-domain testset
This is Raven's Progressive Matrices like task. Complete the missing pattern in this 3x3 matrix.
First Frame
Last Frame
grid_number_sequence
GitHub
Spatiality in-domain testset
The scene shows a 10x10 grid with a green start point, a red end point, and yellow cells marked with numbers 1, 2, and 3. An orange circular agent is positioned at the green start point. The agent can move to adjacent cells (up, down, left, right). Starting from the green start point, the agent must visit the numbered yellow cells in numerical order (1, then 2, then 3), taking the shortest path between each consecutive pair of numbered cells. The agent is allowed to pass through the red end point when visiting the numbered cells if needed. After visiting all numbered cells in sequence, the agent must reach the red end point, also following the shortest path.
First Frame
Last Frame
2d_geometric_transformation
GitHub
Transformation out-of-domain testset
The scene shows a colored 2D polygon, a rotation center marked by a small circular marker, and a dashed target outline indicating the final orientation. Only orientation changes. Rotate the polygon in the clockwise direction. First note the polygon’s initial orientation and the marked center, then read the dashed outline to know the target orientation. Rotate the polygon clockwise around the center until it completely overlaps the dashed outline with no offset. Keep size unchanged, keep a line to the center, and end with the polygon exactly matching the dashed outline.
First Frame
Last Frame
identify_objects
GitHub
Perception training set
The scene contains multiple objects of different shapes and colors arranged randomly. Keep all objects unchanged in their shape, color, size, and position. Identify all orange objects and mark them by adding a thick blue outline around each one.
First Frame
Last Frame

Inference Results

View Full Bench
Domino Chain Prediction - Samples
00
01
02
03
04
Task Domains 1/5
Domino Chain Prediction
Knowledge in-domain testset
Shape Scaling
Abstraction out-of-domain testset
Outline Innermost Square
Spatiality out-of-domain testset
Symbol Substitute
Transformation out-of-domain testset
Locate Intersection
Perception out-of-domain testset
Prompt
Loading...
Ground Truth
First
First Frame
Final
Final Frame
Model Outputs
1/
VBVR-Wan2.2
VBVR-Wan2.2
CogVideoX 1.5
Kling 2.6
LTX-2
Runway Gen-4
Sora 2
Veo 3
Wan 2.2 I2V
Hunyuan I2V
Seedance 2.0

Leaderboard

Modality
Split
Type
Category
2026-04-28