1. X
  2. Vedanuj Goswami
Log inSign up
Vedanuj Goswami
277 posts
user avatar

Vedanuj Goswami

@vedanujg
Research Engineer @MetaAI
Menlo Park, CA
vedanuj.github.io
Joined October 2016
444
Following
610
Followers
RepliesRepliesMediaMedia
  • user avatar
    Vedanuj Goswami
    @vedanujg
    Jul 23, 2024
    Llama3 405B is out with impressive performance across the board and 128k ctx length! Along with improved models for 8B and 70B. Blog : llama.meta.com/llama-download… Paper : ai.meta.com/research/publi… A short 🧵 on what went into scaling and training the largest compute Llama model yet
    Image
  • user avatar
    Vedanuj Goswami
    @vedanujg
    Apr 24, 2024
    It is really amazing how good these models turned out to be!
    user avatar
    Ivan Fioravanti ᯅ
    @ivanfioravanti
    Apr 23, 2024
    Look at this! Llama-3 70B english only is now at 1st 🥇 place with GPT 4 turbo on @lmsysorg Chatbot Arena Leaderboard🔝 I did some rounds too and both 8B and 70B were always the best models for me. Incredible achievement @AIatMeta
    Image
    Image
  • user avatar
    Vedanuj Goswami
    @vedanujg
    Apr 22, 2024
    In addition to data and model optimization, stability, efficiency and fault tolerance of the entire infra stack is extremely crucial when we scale LLM training to tens of thousands of GPUs. Really glad to be working with @reducescatter and the brilliant AI Infra team @AIatMeta,
    user avatar
    Jenya
    @reducescatter
    Apr 18, 2024
    @karpathy's insights on the complexities of training LLMs really hit home. Keeping the #llama3 training alive was a journey filled with hard challenges across the entire tech stack. Reading it kept me sane, knowing others have a similar experience.
  • user avatar
    Vedanuj Goswami
    @vedanujg
    Apr 18, 2024
    Happy to be part of this incredible journey of Llama3 and to share the best open weight 8B and 70B models! Our largest 400B+ model is still cooking but we are providing a sneak peek into how it is trending! Check more details here
    Image
    ai.meta.com
    Introducing Meta Llama 3: The most capable openly available LLM to date
    Today, we’re introducing Meta Llama 3, the next generation of our state-of-the-art open source large language model. In the coming months, we expect to share new capabilities, additional model sizes,...
  • user avatar
    Vedanuj Goswami
    @vedanujg
    Jul 19, 2023
    Excited to be part of the Llama 2 release today and congrats to the entire team for this massive effort! A short thread on my thoughts on the pretrained models, why the results are interesting and where there is room for improvement. 🧵 (1/n) ai.meta.com/llama

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement