1. X
  2. Fan Bai
Log inSign up
Fan Bai
78 posts
user avatar
Fan Bai
@loadingfan
Building agents @TechAtBloomberg; Previously Postdoc @jhuclsp JHU; PhD @GeorgiaTech
bflashcp3f.github.io
Joined September 2015
252
Following
182
Followers
RepliesRepliesMediaMedia
  • user avatar
    Fan Bai
    @loadingfan
    Mar 13
    Excited to share our latest work on the modality gap in multimodal LLMs. Imagine an agent interacting with the real world as humans do. The text it encounters on signs, documents, or screens appears as pixels, not tokens. Yet MLLMs often perform worse when reading the same text
    user avatar
    Kaiser Sun
    @KaiserWhoLearns
    Mar 12
    Multimodal LLMs can read text in images, but why do they often perform worse than when the same text is given as tokens? Our work studies the modality gap of models perceiving text as pixels and shows how to close it. πŸ“„ arxiv.org/abs/2603.09095 πŸ§΅πŸ‘‡ #NLProc #LLM #ComputerVision
    Image
  • user avatar
    Fan Bai
    @loadingfan
    Jan 29
    Really excited to co-lead this work. πŸš€ Agentic AI for scientific discovery is moving fast β€” but it’s not there yet. Across FIRE-BenchπŸ”₯, even SOTA agents (e.g., Claude Code and Codex) stay <50 F1, struggling with research planning and drawing valid conclusions from
    user avatar
    Zhen Wang
    @zhenwang9102
    Jan 28
    πŸ€–πŸ”¬ Can AI actually do science end-to-end? πŸ§ πŸ“ˆ And how would we know when it matches, or surpasses, humans? ⚑πŸ§ͺ AI is rapidly automating scientific discovery, but benchmarking full-cycle discovery, from πŸ’‘ ideation β†’ πŸ§‘β€πŸ’» execution β†’ πŸ“Š conclusions, remains unsolved: 🧐🧐🧐
    Image
    00:00
  • user avatar
    Fan Bai
    @loadingfan
    Nov 5, 2025
    πŸ€” LLMs can ace Olympiad math, yet struggle with something as β€œsimple” as NER β€” even with in-context learning (ICL)? πŸ’‘ Our #EMNLP2025 paper answers why: β€œLLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition.” @mdredze @jhuclsp πŸ‘‰ TL;DR:
    Image
  • user avatar
    Fan Bai
    @loadingfan
    Dec 21, 2024
    Glad to see my first project at JHU has been accepted to #ml4h2024. Here are a few key takeaways: 1. Naive prompting often produces "easy-to-learn" examples, and methods that promote syntactic diversity in LLM output don't address this fundamental issue.
    Image
    Image
    user avatar
    Mark Dredze
    @mdredze
    Dec 15, 2024
    Today at #ML42024: Clinical QA can help doctors find critical information in patient records. But where do we get training data for these systems? Generating this data from an LLM is hard. 🧡 @loadingfan
  • user avatar
    Fan Bai
    @loadingfan
    Nov 16, 2023
    Struggling to sift through endless tables and lengthy webpages for useful information? πŸ‘‰Checkout our paper β€œSchema-Driven Information Extraction from Heterogeneous Tables” to see how LLMs are revolutionizing this process! πŸ”—arXiv: arxiv.org/abs/2305.14336 @ICatGT @mlatgt
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
TermsΒ·PrivacyΒ·CookiesΒ·AccessibilityΒ·Ads InfoΒ·Β© 2026 X Corp.
Advertisement
Advertisement