MLflow Ambassador Joana Mesquita recently shipped a two-part series on agent evaluation and observability.
πΉ Part 1: Evaluate a RAG agent end to end (traces, ground truth, multi-framework scorers)
πΉ Part 2: Align a custom judge with SME feedback using MLflow MemAlign
Part 1
The open source developer platform to build AI applications and models with confidence.
- Join us at AI Agent Builder Day in San Francisco! π Yuki Watanabe, Engineering Lead for OSS MLflow at @databricks, will speak on: Automatic Agent Improvement with Agent Traces using Open Source MLflow ποΈ Monday, August 17 π 1:45β2:30 PM ποΈ Register: luma.com/oss4ai-jiv6
- Static accuracy benchmarks say little about how a model behaves across a real, multi-turn session. Join Kota Tsuyuzaki (@nttcom) at AGNTCon + MCPCon Japan 2026 for trace-based evaluation of open-weight coding agents, using OpenTelemetry and MLflow. π Fri, Sept 11 | 15:35β16:00

