Oh
Install Oh
Theme
Appearance

Agent memory framework

Agent memory that shows its work.

Oh is an open-source memory framework for developers building agents. Your agent saves what it learns as linked records in a SQLite file, so later you can trace an answer back to the passage or table behind it.

Free and MIT licensed · Bun 1.3.14 or newer · No account needed · v0.13.1

$ oh put --kind evidence --key evidence:table-2 \
    --depends-on edition:trial-report-v1 \
    --depends-on assertion:endpoint-12-weeks \
    --value '{"source":"entity:trial-report","locator":"table 2","relationship":"supports"}'
✓ Saved evidence:table-2 (generation 5).
Next: oh get evidence:table-2

$ oh get evidence:table-2
evidence:table-2 (evidence)
{
  "locator": "table 2",
  "relationship": "supports",
  "source": "entity:trial-report"
}
Depends on: assertion:endpoint-12-weeks, edition:trial-report-v1

$ oh verify
✓ Store checked: 5 records and 5 changes replay to the same state (generation 5).
The citation from an example review of a fictional trial report. The record names what it rests on, and oh verify replays every change to check the store. Add --json for canonical JSON.

Each part of your research gets its own record.

Oh keeps the question, the source, and the claim in separate linked records, so revising one leaves the others intact. An assertion records a stance on a claim; citations link that stance to its evidence.

Question inquiry
Save what you are trying to find out, along with the investigation that follows.
Source entity
Identify the paper, dataset, person, or system you are researching, even if its title or URL changes.
Capture edition
Record the edition or extract you read, separate from the source as it looks today.
Claim statement
Write down the claim itself, and keep who accepts it and the evidence for it in separate records.
Citation evidence
Point to a passage, table, or observation, and record how it bears on a stance toward a claim, such as support or contradiction.
Artifact view
Build a brief or answer that keeps links to the records it draws on.

Trace a brief back to the table it rests on.

Follow one example review from its question to the finished brief. Each key names a record you can open from the CLI, and the log keeps every change in the order it happened.

  1. Questioninquiry:primary-endpoint
  2. Sourceentity:trial-report
  3. Captureedition:trial-report-v1
  4. Claimstatement:endpoint-12-weeks
  5. Stanceassertion:endpoint-12-weeks
  6. Citationevidence:table-2
  7. Artifactview:review-brief
$ oh get evidence:table-2 --json
{
  "dependencies": [
    "assertion:endpoint-12-weeks",
    "edition:trial-report-v1"
  ],
  "key": "evidence:table-2",
  "kind": "evidence",
  "recordSha256": "e19a2a8e0d951c8332c95bd11d07213bf2d99ccd6a46e1c7d2eb30487c86d9e4",
  "v": 1,
  "value": {
    "locator": "table 2",
    "relationship": "supports",
    "source": "entity:trial-report"
  }
}
The citation from the terminal above, indented for reading. With --json the CLI prints the same record on one line.oh.ontology.v1 · current

Interfaces

Work with the same records from a terminal, TypeScript, or an agent.

The CLI and the TypeScript SDK read and write the same SQLite file. The packaged Agent Skill has a coding agent run the commands you would run yourself, so its changes land in the log you verify.

CLI

Read one record from the local database and space you select.

$ oh get evidence:table-2 \
  --db research.db \
  --space default

TypeScript SDK

Open the database in your own code and read the same record.

import { Oh } from "@hraness/oh/sdk";

const oh = Oh.open({ databasePath: "research.db" });
try {
  const citation = oh.get("evidence:table-2");
  console.log(citation?.recordSha256);
} finally {
  await oh.close();
}

Agent Skill

Teach a coding agent to check the specification version and replay the log before it reads.

oh contract
oh verify --db research.db --space default
oh get evidence:table-2 \
  --db research.db --space default

Oh’s semantic search scored 88.87% on LongMemEval-S.

In each comparison, one model answers the same questions from each system’s memory, and every answer is scored the same way. Each result links to its protocol, costs, and limits.

LongMemEval-S, all 500 questions

Share of answers judged correct, averaged over three runs. Higher is better.

LongMemEval-S, all 500 questions, share of answers judged correct, zero to one hundred percent
Lab pipeline on Oh and BM25 retrieval93.07%Every user message plus top-ranked replies within 180,000 bytes, with re-reads chosen by rules
Oh semantic retrieval88.87%Up to 100 turns within 96,000 bytes
BM25 keyword retrieval86.13%Up to 100 turns within 96,000 bytes

Oh semantic retrieval scored 88.87% and BM25 86.13%, and a lab pipeline that adds every message the user wrote to the retrieved replies scored 93.07%. GPT-5 mini answered every question three times with each system, and GPT-4o graded the answers with LongMemEval’s own prompts. On the measure fixed before the run, questions answered correctly in at least two of three runs, Oh’s lead over BM25 is 2.8 points with a 95% interval from 0.0 to 5.6, which does not rule out a tie. The pipeline’s instructions and rules, which are not part of the Oh package, were written after studying all 500 questions, so its score is in-sample.

Other memory systems publish LongMemEval-S scores up to 97%. Each chose its own answering model, judge, prompts, and configuration, so those scores do not compare directly with these.

Oh and Supermemory on 60 LongMemEval-S questions

Share of the 60 questions answered correctly, one run. Higher is better.

Oh, Supermemory, and BM25 on 60 LongMemEval-S questions, share answered correctly, zero to one hundred percent
Supermemory75.00%One document per session, hybrid search with reranking
Oh default SDK search71.67%Single turns, reranked locally
BM25 keyword retrieval68.33%Single turns

On 60 LongMemEval-S questions, Supermemory answered 75.00% correctly, Oh’s default SDK search 71.67%, and BM25 68.33%. GPT-4o answered from at most 20 results per system and graded the answers, and Supermemory stored one document per session, as its published method does, while Oh and BM25 stored single turns. Sixty questions, all used earlier in Oh’s development, cannot separate Oh from Supermemory: the difference is −3.33 points, with a 95% interval from −13.33 to +6.67. Three questions each for Oh and BM25 scored zero because their search failed or never ran.

Results on CloneMem and LoCoMo

On 146 CloneMem questions, Oh’s default SDK search answered 80.59% correctly against 69.86% for Oh semantic retrieval alone, with a median of 10.9 seconds of reranking per search on an Apple M5 Max. GPT-4o mini picked an answer from each question’s options three times, reading the top 10 results within 96 KiB, and a pick counted only if it matched the correct option. The questions come from two personas the project had already studied, so the gain may not carry over to new conversations.

On 861 CloneMem questions from seven personas, an earlier benchmark run of the same reranker answered 77.82% correctly against 70.54% for vector retrieval, a gain of 7.28 points with a 95% interval from 4.61 to 10.27. That run gave the reranker only each candidate’s text, where Oh’s SDK sends the whole record with its key and kind. It used the same answering model, scoring, and 96 KiB reading limit as the 146-question study. The personas were kept out of tuning but had been seen earlier in the project, and two failed attempts were dropped under a retry rule added during the study.

On LoCoMo, filling each question’s context with the nearby turns that best match it found 90.08% of the marked evidence against 88.93% for fixed windows across 1,224 questions, yet answers scored 77.33% against 78.11% on a 300-question sample. Both methods built each context from the same top 20 vector matches within 12,000 bytes, and GPT-4o mini answered each question three times with each method and graded the answers. The answer difference, −0.78 points with a 95% interval from −4.64 to +2.34, does not show either method answering better, and the project had evaluated these conversations before.

AI agents ran these studies, and no person or outside group has audited them.

Wordcell uses Oh to query the graph of your Markdown notes.

Use Oh to build memory into an application. Use Wordcell to work with a knowledge base made of Markdown files.

Oh provides records, search, and query results that carry the path back to their sources. The application built on it decides what counts as a memory, when one may be written, and who can use it.

Wordcell is a Markdown knowledge base that gives agents the decisions behind code. Your Markdown files stay authoritative, and Wordcell derives an Oh graph from them to answer queries with a path back to each note. Rebuilding or deleting that graph never changes a note.

Wordcell’s search is its own pipeline with its own evaluations and does not use Oh’s memory retrieval, so the Oh scores above do not carry over to it.

Install

Install and start with a local database.

Latest release: v0.13.1

bun add --global @hraness/oh@0.13.1
oh --help
oh init
oh put --kind entity --key entity:ada-lovelace \
  --value '{"name":"Ada Lovelace","role":"mathematician"}'
oh get entity:ada-lovelace
oh search "mathematician"
oh verify

The CLI needs Bun 1.3.14 or newer. The first task creates one entity, reads it back, finds it with keyword search, and verifies the log. Oh writes to .oh/oh.sqlite and the default space unless you pass --db or --space. Read the full first run on GitHub.

The release run for this version installed the package on Linux and macOS and ran the CLI before publishing the same bytes to npm and GitHub Releases.

Control

Oh keeps research local by default and never decides what is true.

A search score, a valid digest, or an agent’s output never becomes an accepted claim on its own. Acceptance is a separate record that your application or a reviewer writes.

Local by default
Your records, the log of every change, and the keyword index live in a SQLite file you choose. Semantic search caches are derived from the records and can be rebuilt.
Remote services are opt-in
Hosted embeddings, network sync, and a remote libSQL database are used only when you configure them. Sync sends operations, never search vectors.
A history you can replay
Every write appends an operation to the log. The verify command replays the log from an empty graph and confirms it reproduces every digest and stored record.
Agents work within your permissions
The Agent Skill teaches a coding agent to read, write, search, verify, and sync through the same CLI and SDK you use. It grants the agent no extra permissions and tells it never to pick a database, space, or sync destination on its own.

Questions

What to know before you install.

Do I need an account?

No. The CLI and the local SDK work on a SQLite file you choose, with no sign-in, hosted model, or remote database. A hosted service you connect, such as an embedding or sync provider, may need an account of its own.

How is Oh different from Mem0 or Supermemory?

Mem0 and Supermemory give each user of a product a memory, built from facts a model extracts or documents you sync, and both offer hosted plans. Oh stores the records your code or agent writes, with no model rewriting them, each linked to its sources, in one local SQLite file with a history you can replay. Oh has no per-end-user scoping API, connectors, or hosted service. The comparison page at oh.computer/compare also covers Zep, Letta, and Claude’s memory tool.

What is stored, and where?

One SQLite file holds your records and their digests, the append-only log of every change, a keyword index built from the records, and the specification version the file follows. Oh uses .oh/oh.sqlite and the default space unless you name another path or space. Semantic caches and remote copies exist only where you configure them.

Is semantic search required?

No. Keyword search needs no model, and it is the only search the CLI runs, because the CLI configures no semantic backend. In the SDK, configuring a semantic backend makes hybrid search the default, and adding a local reranker makes reranking the default. Local models run through the optional QMD package. Hosted embeddings come from Cloudflare Workers AI through the separate @hraness/oh/semantic-cloud entry point, which sends record text to Cloudflare and caches the vectors in libSQL. Oh drops any result whose record has changed or been removed since it was indexed.

Does a passing verification mean a claim is true?

No. A passing verification means the records and their history are intact: replaying the log reproduced every digest. Search scores measure relevance, not truth. Whether a claim holds is recorded separately, in assertions and review decisions linked to their evidence.

What happens when two writers diverge?

Oh reports a conflict and overwrites nothing. A write can name the generation it was based on (--expected-generation in the CLI), and Oh rejects it if another write landed first. Sync accepts only a history that extends yours; when two histories have diverged, it stops with a conflict error instead of letting the last write win, so your application can keep both logs and reconcile them.

What does it cost?

Oh itself is free and MIT licensed, published on npm as @hraness/oh. Local search runs on your own hardware, and a hosted provider you connect bills you under its own plan. The package has no required runtime dependencies.

Where can I run it?

The CLI, the local SDK, and the SQLite store need Bun 1.3.14 or newer. The runtime-neutral store interfaces and the direct libSQL adapter also run on Node 24, including in serverless functions.

Who made it?

Oh is made by Hraness, which builds tools for agents and humans. Its source code and releases are on GitHub.