██╗  ██╗ ██████╗ ███╗   ██╗ ██████╗██╗  ██╗ ██████╗
██║  ██║██╔═══██╗████╗  ██║██╔════╝██║  ██║██╔═══██╗
███████║██║   ██║██╔██╗ ██║██║     ███████║██║   ██║
██╔══██║██║   ██║██║╚██╗██║██║     ██╔══██║██║   ██║
██║  ██║╚██████╔╝██║ ╚████║╚██████╗██║  ██║╚██████╔╝
╚═╝  ╚═╝ ╚═════╝ ╚═╝  ╚═══╝ ╚═════╝╚═╝  ╚═╝ ╚═════╝

MEMORY BENCHMARKS

Honcho is open-source AI memory for agents. We measured it on three long-context memory benchmarks against Claude Haiku 4.5 with the full conversation in context, Mem0, and Zep. On LongMem S it scores 90.4% (Haiku 4.5 alone: 62.6%), on LoCoMo 89.9%, and it's Pareto-dominant on accuracy, cost, speed and tokens across BEAM from 100K to 10M tokens. Methodology and raw results are on GitHub.

Honcho benchmark results
BenchmarkHonchoHaiku 4.5 alone
LongMem S90.4%62.6%
LoCoMo89.9%75.6%
BEAM 100K0.6300.533
BEAM 500K0.646—
BEAM 1M0.618—
BEAM 10M0.409—
Honcho achieves state-of-the-art performance across the Longmem, LoCoMo, and BEAM memory benchmarks.
We do so while maintaining competitive token efficiency.
Unless otherwise stated, all benchmarks use gemini-2.5-flash-lite as deriver and claude-haiku-4.5 as dreamer and dialectic.[1]
LOCOMO
89.9%
LONGMEM S
90.4%
BEAM 100K
0.630
BEAM 500K
0.646
BEAM 1M
0.618
BEAM 10M
0.409
Honcho fully determines the Pareto frontier for long-context memory benchmarks.
Honcho also demonstrates SOTA cost efficiency and can be used to significantly reduce the cost of using expensive LLMs in production applications.

LOCOMO

89.9%Overall
1,540 Qs • 5 categories
0%20%40%60%80%100%OVERALL75.689.9SINGLE-HOP77.384.0MULTI-HOP74.588.2COMMONSENSE90.193.2TEMPORAL75.077.1ACCURACY (%)
HAIKU 4.5
HONCHO
Honcho uses
52%
fewer tokens on average than baseline across all LoCoMo questions
COST SAVINGS CALCULATOR
Context200K
10K200K
$0.200→$0.096
LoCoMo evaluates multi-turn conversation memory abilities across 1,540 questions with answers judged by an LLM[2]. Honcho achieves 89.9% overall.

Providing only 16,000 tokens of context on average per question, LoCoMo is not well suited to test memory systems today. This is evidenced by the fact that Claude Haiku 4.5 with no memory system achieves competitive scores. Even still, Honcho demonstrates consistent superiority across most categories.

LONGMEMEVAL

90.4%Honcho
500 Qs • 6 categories
96.4%94.3%94.9%90.0%88.7%85.0%
HONCHO (90.4%)
HAIKU 4.5 BASELINE (62.6%)
Honcho uses
89%
fewer tokens than baseline, on average
COST SAVINGS CALCULATOR
Context100K
10K200K
$0.100→$0.011
LongMemEval S tests chat memory over 500 conversations spanning about 115,000 tokens each. Compared to the Claude Haiku 4.5 baseline (62.6% overall), Honcho achieves 90.4% — a 27.8 point improvement. The most dramatic gains are in single-session preference (90.0% vs 23.3%) and multi-session reasoning (85.0% vs 46.6%). Honcho uses only 11.4% of the full context on average to answer each question, demonstrating meaningful token efficiency.

For the same reasons as LoCoMo, LongMem is no longer well suited to test memory systems today. Expensive frontier models with large context windows—while not cost effective—can take all LongMem tokens in-context and produce respectable scores (see commentary on the Gemini comparison below). Regardless, Honcho is state-of-the-art on LongMem with efficient and expensive models alike.

BEAM 100K

0.630Overall
400 Qs • 20 convos
.856.844.784.644.631.463.706
HONCHO (100K)
CLAUDE HAIKU 4.5 (52.9)
CORE
REASONING
MEMORY
Honcho uses
16%
fewer tokens than baseline, on average
COST SAVINGS CALCULATOR
Context100K
10K200K
$0.100→$0.084
BEAM 100K (100K tokens) evaluates long-context memory capabilities across 10 distinct abilities. Token efficiency scales with context length, meaning Honcho becomes massively more cost-efficient as the context length increases. Honcho demonstrates excellent preference following (85.6) and instruction adherence (84.4), showing strong alignment with user intent. Information extraction (78.4) and temporal reasoning (64.4) perform well above baseline. Knowledge updates (46.3) remain challenging at extreme context lengths, indicating an area for improvement. Compared to Claude Haiku 4.5 baseline (52.9), Honcho achieves a 10.1 point improvement.

BEAM scoring is different from LongMem and LoCoMo: rather than grading pass/fail and scoring the overall test by pass rate, BEAM's judge defined in the paper grades each question individually, and the overall test grade is the average of these scores. The LLM judge grades in a step-function pattern: each question's rubric makes it relatively easy to "pass" with a 0.5, and quite difficult to "ace" the question and score 1.0. This property gives BEAM scores a much higher ceiling of excellence.
DETAILED BREAKDOWN
BEAM 100K
OVERALL
63.0
(53)
PREF FOLLOW
85.6
(80)
INSTR FOLLOW
84.4
(76)
INFO EXTRACT
78.4
(54)
CONTRADICT
70.6
(19)
TEMPORAL
64.4
(53)
MULTI-SESSION
63.1
(51)
EVENTS
52.3
(51)
SUMMARY
49.0
(48)
KNOWLEDGE
46.3
(40)
ABSTAIN
36.3
(60)
Honcho
Baseline (52.9)

[1] In practice, the managed Honcho service uses a variety of models for information extraction and retrieval. We also tune Honcho for various use cases. For example, the message batch size when ingesting messages and the amount of tokens spent on dreaming both have an effect on performance. Notes on the configuration for each benchmark are included, and the full configuration for each run is included in the data available at github.com/plastic-labs/honcho-benchmarks.

[2] The LoCoMo paper proposes a token-based F1 scoring methodology, but we use LLM-as-judge, in line with other memory frameworks that publish LoCoMo results and our prior research.

honcho.dev • Plastic Labs
Last updated: December 2025