6 comments

[ 0.19 ms ] story [ 17.4 ms ] thread
Numbers, with the caveats attached.

LongMemEval, full 500 questions, all six types, nothing sampled — against the figures Zep publish for Graphiti:

                          ours (gemini-3.6-flash)   Zep (gpt-4o)
  overall                          94.0%               71.2%
  multi-session                    90.2%               57.9%
  temporal-reasoning               96.2%               62.4%
  knowledge-update                 94.9%               83.3%
What you should discount: their figures are as published, judged by a single gpt-4o where we use a three-model panel with the answering model excluded from its own jury. The models differ in cost class and vintage — their gpt-4o predates these flash models by about two years, in a direction I can't sign. Read it as informative about the tier, not as a controlled head-to-head. One question of 500 is excluded because neither extraction prompt could turn that session into triples; the harness marks a run non-reportable until that count is stated.

The harness, frozen configs and every failing case are in the repo.

Chat memory is also the easy register. The harder one is earnings call transcripts — sixteen quarters per company, every metric restated each quarter, only the date distinguishing the values. Scored under the protocol that benchmark's own authors use (LLM judge, element-wise, refusal counted separately from wrong):

  post-graph-rag   0.807 correct
  TG-RAG           0.599   (published)
  GraphRAG         0.405   (published)
  LightRAG         0.406   (published)
A second judge from a different model family scores the same answers at 0.805. Two-tenths of a point apart matters more than either number, since the standing objection to a judged rate is that it moves with the judge.

Caveat specific to that table: their judge model isn't on my router and their exact rubric is truncated in the public HTML, so mine reproduces their described categories rather than their text, and their figures are on their corpus slice. A few points of margin there are approximate, not decisive.

(comment deleted)
(comment deleted)
(comment deleted)
(comment deleted)