📢 Can LLMs locate software service failures? 🤔
My student @SiyuexiH's #ICLR2025 paper introduces OpenRCA, the first benchmark dataset for evaluating LLMs' root cause analysis capabilities in software systems. LLMs/Agents need to analyze system telemetry data to infer results
Thrilled to see OpenRCA has been used by @AnthropicAI to evaluate its new @claudeai model's capability on the Root Cause Analysis (RCA) task.
👇Check out the original paper thread below.
📢 Can LLMs locate software service failures? 🤔
My student @SiyuexiH's #ICLR2025 paper introduces OpenRCA, the first benchmark dataset for evaluating LLMs' root cause analysis capabilities in software systems. LLMs/Agents need to analyze system telemetry data to infer results
Heartbroken to receive a Reject for our #ICLR2026 submission (Rating: 8/6/6/6).
The hardest part isn't the rejection itself, but the Meta-Review reasoning. The AC dismissed all reviewers' unanimous support, raised two new concerns (with factual errors themselves), and claimed
When eyes and memory clash, who wins? 👁️🧠
Introducing a comprehensive study on vision-knowledge conflicts in MLLMs, where visual input contradicts the model's internal commonsense knowledge—and the results might surprise you. #ACL2025NLP
📈 We developed an automated framework
Check out my student Xiaoyuan Liu's @xyliu_cs collaboration work with Tencent: RISE (Reinforcing Reasoning with Self-Verification), enabling LLMs to simultaneously level-up BOTH their problem-solving AND self-checking skills.
Trust your AI, but can it trust itself? 🤔
Introducing an online reinforcement learning framework, RISE (Reinforcing Reasoning with Self-Verification), enabling LLMs to simultaneously level-up BOTH their problem-solving AND self-checking skills!
🧐 Problems tackled:
✅