CLI
Read one record from the local database and space you select.
$ oh get evidence:table-2 \
--db research.db \
--space defaultAgent memory framework
Oh is an open-source memory framework for developers building agents. Your agent saves what it learns as linked records in a SQLite file, so later you can trace an answer back to the passage or table behind it.
Free and MIT licensed · Bun 1.3.14 or newer · No account needed · v0.13.1
$ oh put --kind evidence --key evidence:table-2 \
--depends-on edition:trial-report-v1 \
--depends-on assertion:endpoint-12-weeks \
--value '{"source":"entity:trial-report","locator":"table 2","relationship":"supports"}'
✓ Saved evidence:table-2 (generation 5).
Next: oh get evidence:table-2
$ oh get evidence:table-2
evidence:table-2 (evidence)
{
"locator": "table 2",
"relationship": "supports",
"source": "entity:trial-report"
}
Depends on: assertion:endpoint-12-weeks, edition:trial-report-v1
$ oh verify
✓ Store checked: 5 records and 5 changes replay to the same state (generation 5).
oh verify replays every change to check the store. Add --json for canonical JSON.How it works
Oh keeps the question, the source, and the claim in separate linked records, so revising one leaves the others intact. An assertion records a stance on a claim; citations link that stance to its evidence.
inquiryentityeditionstatementevidenceviewThe v1 specification defines 18 record kinds, and Oh rejects any other. See every record kind
Trace
Follow one example review from its question to the finished brief. Each key names a record you can open from the CLI, and the log keeps every change in the order it happened.
inquiry:primary-endpointentity:trial-reportedition:trial-report-v1statement:endpoint-12-weeksassertion:endpoint-12-weeksevidence:table-2view:review-brief$ oh get evidence:table-2 --json
{
"dependencies": [
"assertion:endpoint-12-weeks",
"edition:trial-report-v1"
],
"key": "evidence:table-2",
"kind": "evidence",
"recordSha256": "e19a2a8e0d951c8332c95bd11d07213bf2d99ccd6a46e1c7d2eb30487c86d9e4",
"v": 1,
"value": {
"locator": "table 2",
"relationship": "supports",
"source": "entity:trial-report"
}
}Interfaces
The CLI and the TypeScript SDK read and write the same SQLite file. The packaged Agent Skill has a coding agent run the commands you would run yourself, so its changes land in the log you verify.
Read one record from the local database and space you select.
$ oh get evidence:table-2 \
--db research.db \
--space defaultOpen the database in your own code and read the same record.
import { Oh } from "@hraness/oh/sdk";
const oh = Oh.open({ databasePath: "research.db" });
try {
const citation = oh.get("evidence:table-2");
console.log(citation?.recordSha256);
} finally {
await oh.close();
}Teach a coding agent to check the specification version and replay the log before it reads.
oh contract
oh verify --db research.db --space default
oh get evidence:table-2 \
--db research.db --space defaultBenchmarks
In each comparison, one model answers the same questions from each system’s memory, and every answer is scored the same way. Each result links to its protocol, costs, and limits.
Share of answers judged correct, averaged over three runs. Higher is better.
Oh semantic retrieval scored 88.87% and BM25 86.13%, and a lab pipeline that adds every message the user wrote to the retrieved replies scored 93.07%. GPT-5 mini answered every question three times with each system, and GPT-4o graded the answers with LongMemEval’s own prompts. On the measure fixed before the run, questions answered correctly in at least two of three runs, Oh’s lead over BM25 is 2.8 points with a 95% interval from 0.0 to 5.6, which does not rule out a tie. The pipeline’s instructions and rules, which are not part of the Oh package, were written after studying all 500 questions, so its score is in-sample.
Other memory systems publish LongMemEval-S scores up to 97%. Each chose its own answering model, judge, prompts, and configuration, so those scores do not compare directly with these.
Share of the 60 questions answered correctly, one run. Higher is better.
On 60 LongMemEval-S questions, Supermemory answered 75.00% correctly, Oh’s default SDK search 71.67%, and BM25 68.33%. GPT-4o answered from at most 20 results per system and graded the answers, and Supermemory stored one document per session, as its published method does, while Oh and BM25 stored single turns. Sixty questions, all used earlier in Oh’s development, cannot separate Oh from Supermemory: the difference is −3.33 points, with a 95% interval from −13.33 to +6.67. Three questions each for Oh and BM25 scored zero because their search failed or never ran.
On 146 CloneMem questions, Oh’s default SDK search answered 80.59% correctly against 69.86% for Oh semantic retrieval alone, with a median of 10.9 seconds of reranking per search on an Apple M5 Max. GPT-4o mini picked an answer from each question’s options three times, reading the top 10 results within 96 KiB, and a pick counted only if it matched the correct option. The questions come from two personas the project had already studied, so the gain may not carry over to new conversations.
On 861 CloneMem questions from seven personas, an earlier benchmark run of the same reranker answered 77.82% correctly against 70.54% for vector retrieval, a gain of 7.28 points with a 95% interval from 4.61 to 10.27. That run gave the reranker only each candidate’s text, where Oh’s SDK sends the whole record with its key and kind. It used the same answering model, scoring, and 96 KiB reading limit as the 146-question study. The personas were kept out of tuning but had been seen earlier in the project, and two failed attempts were dropped under a retry rule added during the study.
On LoCoMo, filling each question’s context with the nearby turns that best match it found 90.08% of the marked evidence against 88.93% for fixed windows across 1,224 questions, yet answers scored 77.33% against 78.11% on a 300-question sample. Both methods built each context from the same top 20 vector matches within 12,000 bytes, and GPT-4o mini answered each question three times with each method and graded the answers. The answer difference, −0.78 points with a 95% interval from −4.64 to +2.34, does not show either method answering better, and the project had evaluated these conversations before.
AI agents ran these studies, and no person or outside group has audited them.
Built on Oh
Use Oh to build memory into an application. Use Wordcell to work with a knowledge base made of Markdown files.
Oh provides records, search, and query results that carry the path back to their sources. The application built on it decides what counts as a memory, when one may be written, and who can use it.
Wordcell is a Markdown knowledge base that gives agents the decisions behind code. Your Markdown files stay authoritative, and Wordcell derives an Oh graph from them to answer queries with a path back to each note. Rebuilding or deleting that graph never changes a note.
Wordcell’s search is its own pipeline with its own evaluations and does not use Oh’s memory retrieval, so the Oh scores above do not carry over to it.
Install
Latest release: v0.13.1
bun add --global @hraness/oh@0.13.1
oh --helpoh init
oh put --kind entity --key entity:ada-lovelace \
--value '{"name":"Ada Lovelace","role":"mathematician"}'
oh get entity:ada-lovelace
oh search "mathematician"
oh verifyThe CLI needs Bun 1.3.14 or newer. The first task creates one entity, reads it back, finds it with keyword search, and verifies the log. Oh writes to .oh/oh.sqlite and the default space unless you pass --db or --space. Read the full first run on GitHub.
The release run for this version installed the package on Linux and macOS and ran the CLI before publishing the same bytes to npm and GitHub Releases.
Control
A search score, a valid digest, or an agent’s output never becomes an accepted claim on its own. Acceptance is a separate record that your application or a reviewer writes.
Questions
No. The CLI and the local SDK work on a SQLite file you choose, with no sign-in, hosted model, or remote database. A hosted service you connect, such as an embedding or sync provider, may need an account of its own.
Mem0 and Supermemory give each user of a product a memory, built from facts a model extracts or documents you sync, and both offer hosted plans. Oh stores the records your code or agent writes, with no model rewriting them, each linked to its sources, in one local SQLite file with a history you can replay. Oh has no per-end-user scoping API, connectors, or hosted service. The comparison page at oh.computer/compare also covers Zep, Letta, and Claude’s memory tool.
One SQLite file holds your records and their digests, the append-only log of every change, a keyword index built from the records, and the specification version the file follows. Oh uses .oh/oh.sqlite and the default space unless you name another path or space. Semantic caches and remote copies exist only where you configure them.
No. Keyword search needs no model, and it is the only search the CLI runs, because the CLI configures no semantic backend. In the SDK, configuring a semantic backend makes hybrid search the default, and adding a local reranker makes reranking the default. Local models run through the optional QMD package. Hosted embeddings come from Cloudflare Workers AI through the separate @hraness/oh/semantic-cloud entry point, which sends record text to Cloudflare and caches the vectors in libSQL. Oh drops any result whose record has changed or been removed since it was indexed.
No. A passing verification means the records and their history are intact: replaying the log reproduced every digest. Search scores measure relevance, not truth. Whether a claim holds is recorded separately, in assertions and review decisions linked to their evidence.
Oh reports a conflict and overwrites nothing. A write can name the generation it was based on (--expected-generation in the CLI), and Oh rejects it if another write landed first. Sync accepts only a history that extends yours; when two histories have diverged, it stops with a conflict error instead of letting the last write win, so your application can keep both logs and reconcile them.
Oh itself is free and MIT licensed, published on npm as @hraness/oh. Local search runs on your own hardware, and a hosted provider you connect bills you under its own plan. The package has no required runtime dependencies.
The CLI, the local SDK, and the SQLite store need Bun 1.3.14 or newer. The runtime-neutral store interfaces and the direct libSQL adapter also run on Node 24, including in serverless functions.
Oh is made by Hraness, which builds tools for agents and humans. Its source code and releases are on GitHub.