New benchmark just dropped: MF² tests whether models really understand full movies, not just clips. It focuses on narrative understanding (key events, emotional arcs, causal chains)—things people remember but models often miss. Huge thanks to everyone in this big collab!
🚨Meet MF²: Movie Facts & Fibs: a new benchmark for long-movie understanding!
🤔Do you think your model understands movies?
Unlike existing benchmarks, MF² targets memorable events, emotional arcs 💔, and causal chains 🔗 — things humans recall easily, but even top models like


