Wordcell
Install Wordcell
Theme
Appearance

Benchmarks

Each Wordcell and Oh result here links its raw data and states its limits.

Wordcell’s own measurements cover the size of a context handoff and reranking quality. The Oh memory kernel that Wordcell embeds publishes its own memory benchmarks, reported here as Oh’s results, not Wordcell’s.

Latest release: v0.22.5. Oh’s result files are copied byte for byte from hraness/oh at pinned commits.

What Wordcell measures

Across four queries on a seven-note public vault, packed snippets used 79.98% fewer UTF-8 bytes than the same notes in full.

12,126 bytesPacked snippets

60,584 bytesThe same notes in full

Measured with Wordcell 0.21.3. This is payload size only: it does not measure tokens, answer quality, speed, or an advantage over another search tool. Method, raw results, and reproduction.

A relevant source in the first result

BEIR SciFact · 300 queries · Relevant result at rank 1

BEIR SciFact: Relevant result at rank 1, out of 300 queries
Wordcell exact search33.7%101 of 300 queries
Exact search + Jev reranking53.7%161 of 300 queries · optional paid provider

Same corpus and candidate windows. Scientific abstracts; this study does not establish answer quality or results on your vault.

Models, method, and limitations
Recorded
Model
jev-1.13.0 reranker; the exact baseline uses no model.
Answer reader
None. The study measures source ranking, without generating answers.
Evaluation
The public BEIR SciFact relevance judgments over 5,183 abstracts.
Context
The same 25-candidate windows, with snippets bounded to 512 UTF-8 bytes each.
Exposure
40 initial queries followed by an unchanged 260-query confirmation. Public data may overlap model training; no private-vault claim.

nDCG at five rose from 0.4024 to 0.5781. It improved for 87 queries and regressed for 8. For 110 queries, neither candidate window contained a judged relevant source.

Reranking sends the query and each candidate’s title, path, and up to 512 bytes of its snippet to a paid provider. It is optional; the local search path runs without it. QMD, Letta, and Supermemory were not evaluated under this protocol.

The Oh kernel Wordcell embeds, measured on its own

Oh’s conversation-memory benchmarks evaluate its own memory-retrieval API, reader models, and evaluation protocols. Those scores do not transfer to a Wordcell vault merely because it uses the same library.

Answers judged correct on LoCoMo

LoCoMo · 1,540 questions · Answers judged correct

LoCoMo: Answers judged correct, out of 1,540 questions
Oh semantic retrieval, GPT-5 mini84.4%1,300 of 1,540 questions
BM25 window, GPT-5 mini81.6%1,257 of 1,540 questions
Oh semantic retrieval, GPT-5 nano81.0%1,248 of 1,540 questions
BM25 window, GPT-5 nano78.1%1,203 of 1,540 questions

Oh measured on its own, not through a Wordcell vault: one run over questions Oh had seen before, with no confidence interval.

Models, method, and limitations
Published
Model
Oh semantic retrieval and a BM25 window, each with a 24,000-byte context budget.
Answer reader
GPT-5 mini and GPT-5 nano, each answering every question once.
Evaluation
LoCoMo J: a GPT-4o mini judge marks each answer correct or wrong. Single-hop, multi-hop, temporal, and open-domain questions are scored; adversarial questions are not.
Context
Up to 100 retrieved items packed into 24,000 bytes per question, the same for both systems.
Exposure
1,226 questions from eight conversations had been evaluated before, and 314 from two conversations were used during development. None are unseen.

Paired by question with GPT-5 mini, Oh semantic retrieval was right where the BM25 window was wrong on 126 questions, wrong where it was right on 83, and matched it on 1,331. With GPT-5 nano the counts were 169, 124, and 1,247.

Answers judged correct by question category, in percent
CategoryQuestionsOh semantic retrieval, GPT-5 miniBM25 window, GPT-5 miniOh semantic retrieval, GPT-5 nanoBM25 window, GPT-5 nano
Single-hop84191.3%90.1%87.6%87.3%
Multi-hop28279.1%69.1%76.2%63.8%
Temporal32181.3%80.4%77.3%76.0%
Open-domain9650.0%47.9%50.0%46.9%

Limits

  • Oh’s own summary says these results “do not establish fresh confirmation, statistical superiority or benchmark saturation.”
  • Oh’s summary adds: “Not a pinned-snapshot reproduction of any leaderboard harness.”
  • Of the questions evaluated before and those used in development, Oh says “neither means unseen.”
  • The file records no run date or hardware. The date shown is when Oh published the result, September 10, 2026.

Matched and published comparisons

Oh ran one small pilot of Supermemory, Oh, and BM25 under one protocol; it is Oh’s result, not Wordcell’s. Figures that other memory systems publish use their own protocols, so they appear in a table, not a chart.

Answers judged correct in a small LongMemEval pilot

LongMemEval-S · 60 questions · Answers judged correct

LongMemEval-S: Answers judged correct, out of 60 questions
Supermemory75.00%45 of 60 questions
Oh71.67%43 of 60 questions
BM2568.33%41 of 60 questions

A development pilot on questions Oh had seen before, with Supermemory indexed per session and Oh and BM25 per turn; it does not rank the three systems.

Models, method, and limitations
Recorded
Model
Supermemory search with one fixed profile, Oh retrieval, and BM25 over the same conversation histories.
Answer reader
GPT-4o through an unpinned Vercel AI Gateway alias.
Evaluation
A GPT-4o judge through an unpinned gateway alias. Oh and BM25 each did not finish three of the 60 questions; those count as misses.
Context
Median context per question whose retrieval finished: 1,688 tokens for Supermemory (60 questions), 5,014 for Oh (57 questions), and 7,675 for BM25 (57 questions).
Exposure
60 LongMemEval-S questions, 10 per question type, all previously exposed. A development pilot, not an unseen test set.

Oh minus Supermemory: −3.33 percentage points, with a 95% interval from −13.33 to +6.67 over 60 paired questions. The interval includes zero, so the pilot does not separate Oh from Supermemory.

Oh minus BM25: +3.33 percentage points, with a 95% interval from −1.67 to +8.33 over 60 paired questions. Oh reports this second comparison as descriptive only.

Limits

  • Oh notes that its “bootstrap intervals describe resampling within this sample, not an unseen population.”
  • In Oh’s words, “retrieval granularity differs by arm.”
  • On Supermemory’s search profile, Oh says “this is not a claim about its defaults or best configuration.”
Selected figures other memory systems publish on their own pages, checked September 26, 2026
SystemBenchmarkPublished figureMetric, as the source names itModelSource
Mem0LoCoMo92.5Mem0 ScoreNot namedMem0’s research page
Mem0LongMemEval94.4Mem0 ScoreNot namedMem0’s research page
ZepLoCoMo94.7% (1,459 / 1,540 correct)Accuracygpt-5.4 reader and gpt-5.4 judgeZep’s research page
ZepLongMemEval90.2% (451 / 500 correct)Accuracygpt-5.4 reader and gpt-5.4 judgeZep’s research page
SupermemoryLongMemEval-S, 500 questions97%Recall@20 with aggregation, in a table titled “LLM-as-judge evaluation”gpt-4oSupermemory’s LongMemEval research page
SupermemoryLongMemEval-S, 500 questions84.6%Recall@20 with aggregation, in a table titled “LLM-as-judge evaluation”gpt-5Supermemory’s LongMemEval research page
SupermemoryLongMemEval-S, 500 questions85.2%Recall@20 with aggregation, in a table titled “LLM-as-judge evaluation”gemini-3-proSupermemory’s LongMemEval research page

Each row uses its publisher’s own protocol, reader, and judge, and links its source. The rows are not a matched ranking, and they are not comparable with the charts above.

How to reproduce

Each Wordcell and Oh figure on this page comes from a file you can read and a procedure you can rerun. The published figures in the table link their sources.