Log inSign up
Avijit Ghosh
6,138 posts
Avijit Ghosh profile banner
@evijit

Avijit Ghosh

@evijit
AI Research & Policy @huggingface 🤗 . Leading: @evaluatingevals @huggingscience
Boston, Massachusetts
evijit.io
Joined January 2012
1,635
Following
3,076
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @evijit
    Avijit Ghosh
    @evijit
    Sep 4
    Such a cool idea!
    @Cohere_Labs
    Cohere Labs
    Cohere
    @Cohere_Labs
    Sep 3
    Today we’re releasing the Agentic Task Ecosystem (ATE) dataset: nearly 700K tools from public MCP servers offering a new window into what developers are building AI agents to do. Our analysis of this ecosystem reveals early patterns in how AI will impact the future of work.
    Image
  • @evijit
    Avijit Ghosh
    @evijit
    Sep 4
    AA index is only good until Astra doesn’t top it and then it’s a bad index? Evals are running on Twitter vibes basically these days. Have been working on a more dynamic composite Eval index method with EvalEval data! More soon :)
  • @evijit
    Avijit Ghosh
    @evijit
    Sep 3
    A surreal news to get on an all-hands call at 7:30am in the morning, and Jensen came to say hi :D
    @JensenHuang
    Jensen Huang
    NVIDIA
    @JensenHuang
    Sep 3
    Exciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you
    8
  • @evijit
    Avijit Ghosh
    @evijit
    Aug 28
    Having seen an internal prototype of this months ago I am simply awed by how cute the consumer product turned out to be!
    @DynamicWebPaige
    👩‍💻 Paige Bailey
    Google AI Studio
    @DynamicWebPaige
    Aug 27
    😍 LOOK AT HOW CUTE HE IS
    Image
    1
  • @evijit
    Avijit Ghosh
    @evijit
    Aug 27
    This makes evaluations less poisoned perhaps but it makes evaluating evaluations harder. There should be some form of representative public subset for auditing purposes, imo.
    @SolomonMg
    Sol Messing
    @SolomonMg
    Aug 27
    Thrilled to announce a step toward fixing benchmark contamination—the first successful test of our confidential evaluation framework “double blind evals” on a frontier class proprietary model with a non-public evaluation set.
    2
Advertisement
Advertisement