Log inSign up
Mohamed El Banani
226 posts
Mohamed El Banani profile banner
@_mbanani

Mohamed El Banani

@_mbanani
MTS @theworldlabs. Prev: @UMichCSE, @GoogleAI, @MetaAI, @GeorgiaTech. I am interested in computer vision, machine learning, and cognitive science. 🇪🇬
San Francisco, CA
mbanani.github.io
Joined June 2020
823
Following
1,006
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @_mbanani
    Mohamed El Banani
    @_mbanani
    Dec 2, 2024
    We’re finally sharing what we’ve been up to @theworldlabs! This is the first step on our journey to build fully interactive and immersive worlds that allow you to bring your creativity to life. Check out the demos, my favorites are the Kandinsky landscape and Van Gogh terrace.
    @theworldlabs
    World Labs
    @theworldlabs
    Dec 2, 2024
    We’ve been busy building an AI system to generate 3D worlds from a single image. Check out some early results on our site, where you can interact with our scenes directly in the browser! worldlabs.ai/blog 1/n
    Image
    00:00
    1
  • @_mbanani
    Mohamed El Banani
    @_mbanani
    Sep 1
    Very excited that I can finally share what we’ve been working on! Atlas is a big step towards fully interactive and controllable world models. Can’t wait to see what you all build with Atlas!!
    @theworldlabs
    World Labs
    @theworldlabs
    Sep 1
    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
    Image
    00:00
    6
  • @_mbanani
    Mohamed El Banani
    @_mbanani
    Oct 17, 2025
    We've been exploring different ways of modeling the world at @theworldlabs. This direction combines our real-time learned renderer with posed frames as a persistent spatial memory. I am excited to see where we go next! Check out the demo here: rtfm.worldlabs.ai
    @theworldlabs
    World Labs
    @theworldlabs
    Oct 16, 2025
    Introducing RTFM (Real-Time Frame Model): a highly efficient World Model that generates video frames in real time as you interact with it, powered by a single H100 GPU. RTFM renders persistent and 3D consistent worlds, both real and imaginary. Try our demo of RTFM today!
    Image
    00:00
    2
  • @_mbanani
    Mohamed El Banani
    @_mbanani
    Apr 5, 2023
    Language models keep getting better, can we use them for better visual learning? In our CVPR paper, we use language models to sample conceptually similar and visually dissimilar image pairs! We use those pairs for stronger contrastive learning.  🧵⬇️
    A figure showing 3 images: left to right, the first image is of a creek, the second is of an owl flying in a creek, and the third is of an owl flying across a clear sky. The captions are "late after post. thunderstorm, north Ohio", "Snowy Owl lifting off", "Snow owl taking off". The figure depicts that the right two images lie close to each other in language embedding space, while the left two are close to each other in visual embedding space. There is a loss term that is bringing the visual embeddings of the two owl pictures close to each other.
    Justin Johnson and 2 others
    2
  • @_mbanani
    Mohamed El Banani
    @_mbanani
    Oct 14, 2021
    Check out "Bootstrap your own correspondences" w/ @jcjohnss at #ICCV2021 (session 5). Come to the Q&A session to chat with me about self-supervised geometric feature learning and registration. ⏰ October 14th: 9-10am EDT 🌐 mbanani.github.io/byoc 📽️ youtu.be/YJ9oTyI0xF0
    1
Advertisement
Advertisement