Agent Memory

The open benchmark for agent memory

Agent Memory Leaderboard

Measure what your agents remember.
Compare what truly matters.

Agent Memory Leaderboard
OPEN BENCHMARK

Agent Memory Leaderboard

A public benchmark space for comparing textual, multimodal, and coding-agent memory systems under a consistent evaluation flow.

Leaderboard Preview

Benchmark Tracks

Each track keeps its own result table and detailed metric breakdown.

Textual Memory

Long-context, persona, script, and conversation-memory benchmarks.

Multimodal Memory

Memory retrieval and generation over image-rich or multimodal tasks.

Coding Agent Memory

Agent memory support for coding tasks and repository-context recall.

Evaluation Flow

Open-source methods and commercial products both use participant-hosted Add/Search APIs. AML does not deploy repository-only submissions.

1

Deploy Add/Search APIs

Operate publicly reachable Add/Search endpoints and submit the fixed API version for review.

2

Run a smoke test

Use the issued key to verify the synchronous Add/Search flow.

3

Submit a formal evaluation

After smoke passes, submit the full scored evaluation.

Explore the Platform

Use the product pages to inspect rankings, run evaluations, and prepare an integration.

01

Leaderboard

Public ranking with filters, dataset columns, and score bars.

02

Evaluation

Create eval jobs, watch progress, and inspect private results.

03

Participation Guide

Eligibility, submission routes, required materials, timelines, rewards, and publication rules.

04

Documentation

User guide, evaluation workflow, API contract, security, and result publication.

05

Guide

Add/search API contract, request fields, polling, and response schemas.