Skip to content

Technology

The research operators the engine is built on, and the benchmark cards recording how each published number was measured.

  • How We Measure AI Memory Honestly (2026-06-23): A 16-point benchmark regression was really a settings change. It taught us to make every AI-memory number we publish re-derivable — the principles here.
  • LoCoMo — June 2026 Matrix (Archived) (2026-06-23): Archived: the June 2026 LoCoMo cross-system matrix (conv-26/conv-47, four judges, k-sweep 10-200). Superseded by the full-corpus bench-v1 result — kept for reference with its imperfections disclosed.
  • LoCoMo — Long-Term Conversational Memory (2026-06-23): LoCoMo benchmark card for Mnemoverse: the full-corpus bench-v1 result (n=1,986), two axes always labeled — recall@k and LLM-judge — stock vs tuned-B vs our own naive-cosine floor, plus dated vendor comparisons.
  • BEAM Benchmark: Long-Term Memory at 10M Tokens (2026-06-21): BEAM benchmark for the Mnemoverse memory engine: what it tests at 10M-token scale, our protocol, committed run (0.61 judge accuracy), and disclosed limits.
  • HotpotQA — Multi-Hop QA Benchmark Card (2026-06-21): How Mnemoverse runs HotpotQA: multi-hop QA combining evidence from 2 gold paragraphs among 8 distractors. Measured Answer F1, with provenance caveats.
  • LongMemEval Benchmark: Long-Term Interactive Memory (2026-06-21): LongMemEval benchmark card for Mnemoverse: what it tests, our protocol, measured runs (not yet in the live matrix), and a documented judge discrepancy.
  • MuSiQue — Compositional Multi-Hop QA (2026-06-21): MuSiQue benchmark card for Mnemoverse: what it tests, our protocol, and our measured run (2-hop-only subset, not in the live matrix), with provenance caveats.
  • Agent Memory Benchmarks: LoCoMo, BEAM, LongMemEval (2026-05-18): How the Mnemoverse memory engine benchmarks: LoCoMo (primary matrix, automated, judge-free recall@k plus two labeled LLM-judge graders), plus BEAM, HotpotQA, MuSiQue, LongMemEval — with config.
  • Semantic Level of Detail (SLoD) (2026-04-06): Semantic Level of Detail (SLoD): multi-scale knowledge representation via heat-kernel diffusion on hyperbolic manifolds — a zoom from facts to themes.
  • Tensor-Hyperbolic Graphs for AI Memory (2026-04-06): Our bet that AI memory is intrinsically hyperbolic — and what a tensor-hyperbolic graph adds. Proven core, labelled conjectures, not yet in production.