Benchmarks
Each Wordcell and Oh result here links its raw data and states its limits.
Wordcell’s own measurements cover the size of a context handoff and reranking quality. The Oh memory kernel that Wordcell embeds publishes its own memory benchmarks, reported here as Oh’s results, not Wordcell’s.
Latest release: v0.22.5. Oh’s result files are copied byte for byte from hraness/oh at pinned commits.
What Wordcell measures
Across four queries on a seven-note public vault, packed snippets used 79.98% fewer UTF-8 bytes than the same notes in full.
Measured with Wordcell 0.21.3. This is payload size only: it does not measure tokens, answer quality, speed, or an advantage over another search tool. Method, raw results, and reproduction.
A relevant source in the first result
BEIR SciFact · 300 queries · Relevant result at rank 1
Same corpus and candidate windows. Scientific abstracts; this study does not establish answer quality or results on your vault.
Models, method, and limitations
- Recorded
- Model
- jev-1.13.0 reranker; the exact baseline uses no model.
- Answer reader
- None. The study measures source ranking, without generating answers.
- Evaluation
- The public BEIR SciFact relevance judgments over 5,183 abstracts.
- Context
- The same 25-candidate windows, with snippets bounded to 512 UTF-8 bytes each.
- Exposure
- 40 initial queries followed by an unchanged 260-query confirmation. Public data may overlap model training; no private-vault claim.
nDCG at five rose from 0.4024 to 0.5781. It improved for 87 queries and regressed for 8. For 110 queries, neither candidate window contained a judged relevant source.
Reranking sends the query and each candidate’s title, path, and up to 512 bytes of its snippet to a paid provider. It is optional; the local search path runs without it. QMD, Letta, and Supermemory were not evaluated under this protocol.
The Oh kernel Wordcell embeds, measured on its own
Oh’s conversation-memory benchmarks evaluate its own memory-retrieval API, reader models, and evaluation protocols. Those scores do not transfer to a Wordcell vault merely because it uses the same library.
Answers judged correct on LoCoMo
LoCoMo · 1,540 questions · Answers judged correct
Oh measured on its own, not through a Wordcell vault: one run over questions Oh had seen before, with no confidence interval.
Models, method, and limitations
- Published
- Model
- Oh semantic retrieval and a BM25 window, each with a 24,000-byte context budget.
- Answer reader
- GPT-5 mini and GPT-5 nano, each answering every question once.
- Evaluation
- LoCoMo J: a GPT-4o mini judge marks each answer correct or wrong. Single-hop, multi-hop, temporal, and open-domain questions are scored; adversarial questions are not.
- Context
- Up to 100 retrieved items packed into 24,000 bytes per question, the same for both systems.
- Exposure
- 1,226 questions from eight conversations had been evaluated before, and 314 from two conversations were used during development. None are unseen.
Paired by question with GPT-5 mini, Oh semantic retrieval was right where the BM25 window was wrong on 126 questions, wrong where it was right on 83, and matched it on 1,331. With GPT-5 nano the counts were 169, 124, and 1,247.
| Category | Questions | Oh semantic retrieval, GPT-5 mini | BM25 window, GPT-5 mini | Oh semantic retrieval, GPT-5 nano | BM25 window, GPT-5 nano |
|---|---|---|---|---|---|
| Single-hop | 841 | 91.3% | 90.1% | 87.6% | 87.3% |
| Multi-hop | 282 | 79.1% | 69.1% | 76.2% | 63.8% |
| Temporal | 321 | 81.3% | 80.4% | 77.3% | 76.0% |
| Open-domain | 96 | 50.0% | 47.9% | 50.0% | 46.9% |
Limits
- Oh’s own summary says these results “do not establish fresh confirmation, statistical superiority or benchmark saturation.”
- Oh’s summary adds: “Not a pinned-snapshot reproduction of any leaderboard harness.”
- Of the questions evaluated before and those used in development, Oh says “neither means unseen.”
- The file records no run date or hardware. The date shown is when Oh published the result, September 10, 2026.
Matched and published comparisons
Oh ran one small pilot of Supermemory, Oh, and BM25 under one protocol; it is Oh’s result, not Wordcell’s. Figures that other memory systems publish use their own protocols, so they appear in a table, not a chart.
Answers judged correct in a small LongMemEval pilot
LongMemEval-S · 60 questions · Answers judged correct
A development pilot on questions Oh had seen before, with Supermemory indexed per session and Oh and BM25 per turn; it does not rank the three systems.
Models, method, and limitations
- Recorded
- Model
- Supermemory search with one fixed profile, Oh retrieval, and BM25 over the same conversation histories.
- Answer reader
- GPT-4o through an unpinned Vercel AI Gateway alias.
- Evaluation
- A GPT-4o judge through an unpinned gateway alias. Oh and BM25 each did not finish three of the 60 questions; those count as misses.
- Context
- Median context per question whose retrieval finished: 1,688 tokens for Supermemory (60 questions), 5,014 for Oh (57 questions), and 7,675 for BM25 (57 questions).
- Exposure
- 60 LongMemEval-S questions, 10 per question type, all previously exposed. A development pilot, not an unseen test set.
Oh minus Supermemory: −3.33 percentage points, with a 95% interval from −13.33 to +6.67 over 60 paired questions. The interval includes zero, so the pilot does not separate Oh from Supermemory.
Oh minus BM25: +3.33 percentage points, with a 95% interval from −1.67 to +8.33 over 60 paired questions. Oh reports this second comparison as descriptive only.
Limits
- Oh notes that its “bootstrap intervals describe resampling within this sample, not an unseen population.”
- In Oh’s words, “retrieval granularity differs by arm.”
- On Supermemory’s search profile, Oh says “this is not a claim about its defaults or best configuration.”
| System | Benchmark | Published figure | Metric, as the source names it | Model | Source |
|---|---|---|---|---|---|
| Mem0 | LoCoMo | 92.5 | Mem0 Score | Not named | Mem0’s research page |
| Mem0 | LongMemEval | 94.4 | Mem0 Score | Not named | Mem0’s research page |
| Zep | LoCoMo | 94.7% (1,459 / 1,540 correct) | Accuracy | gpt-5.4 reader and gpt-5.4 judge | Zep’s research page |
| Zep | LongMemEval | 90.2% (451 / 500 correct) | Accuracy | gpt-5.4 reader and gpt-5.4 judge | Zep’s research page |
| Supermemory | LongMemEval-S, 500 questions | 97% | Recall@20 with aggregation, in a table titled “LLM-as-judge evaluation” | gpt-4o | Supermemory’s LongMemEval research page |
| Supermemory | LongMemEval-S, 500 questions | 84.6% | Recall@20 with aggregation, in a table titled “LLM-as-judge evaluation” | gpt-5 | Supermemory’s LongMemEval research page |
| Supermemory | LongMemEval-S, 500 questions | 85.2% | Recall@20 with aggregation, in a table titled “LLM-as-judge evaluation” | gemini-3-pro | Supermemory’s LongMemEval research page |
Each row uses its publisher’s own protocol, reader, and judge, and links its source. The rows are not a matched ranking, and they are not comparable with the charts above.
How to reproduce
Each Wordcell and Oh figure on this page comes from a file you can read and a procedure you can rerun. The published figures in the table link their sources.
- Context handoff: method, raw results, and reproduction
- Reranking study: evidence and limits
- Oh’s LoCoMo result and Oh’s LongMemEval pilot result
- Oh’s benchmark guide
- memory-evolution-locomo-sealed-1540-v1.json at hraness/oh commit
3add170, copied intodocs/evaluations/oh/(67,362 bytes, SHA-2560c8955db9546) - memory-framework-pilot-v1.json at hraness/oh commit
9edd9f1, copied intodocs/evaluations/oh/(29,285 bytes, SHA-25687cc50aa692e)