Honcho is open-source AI memory for agents. We measured it on three long-context memory benchmarks against Claude Haiku 4.5 with the full conversation in context, Mem0, and Zep. On LongMem S it scores 90.4% (Haiku 4.5 alone: 62.6%), on LoCoMo 89.9%, and it's Pareto-dominant on accuracy, cost, speed and tokens across BEAM from 100K to 10M tokens. Methodology and raw results are on GitHub.
| Benchmark | Honcho | Haiku 4.5 alone |
|---|---|---|
| LongMem S | 90.4% | 62.6% |
| LoCoMo | 89.9% | 75.6% |
| BEAM 100K | 0.630 | 0.533 |
| BEAM 500K | 0.646 | — |
| BEAM 1M | 0.618 | — |
| BEAM 10M | 0.409 | — |
- 90.4% on Longmem S (92.6% with Gemini 3 Pro)
- 89.9% on LoCoMo (beating our previous score of 86.9%)
- Top scores across all BEAM tests
LOCOMO
Providing only 16,000 tokens of context on average per question, LoCoMo is not well suited to test memory systems today. This is evidenced by the fact that Claude Haiku 4.5 with no memory system achieves competitive scores. Even still, Honcho demonstrates consistent superiority across most categories.
LONGMEMEVAL
For the same reasons as LoCoMo, LongMem is no longer well suited to test memory systems today. Expensive frontier models with large context windows—while not cost effective—can take all LongMem tokens in-context and produce respectable scores (see commentary on the Gemini comparison below). Regardless, Honcho is state-of-the-art on LongMem with efficient and expensive models alike.
BEAM 100K
BEAM scoring is different from LongMem and LoCoMo: rather than grading pass/fail and scoring the overall test by pass rate, BEAM's judge defined in the paper grades each question individually, and the overall test grade is the average of these scores. The LLM judge grades in a step-function pattern: each question's rubric makes it relatively easy to "pass" with a 0.5, and quite difficult to "ace" the question and score 1.0. This property gives BEAM scores a much higher ceiling of excellence.
[1] In practice, the managed Honcho service uses a variety of models for information extraction and retrieval. We also tune Honcho for various use cases. For example, the message batch size when ingesting messages and the amount of tokens spent on dreaming both have an effect on performance. Notes on the configuration for each benchmark are included, and the full configuration for each run is included in the data available at github.com/plastic-labs/honcho-benchmarks.
[2] The LoCoMo paper proposes a token-based F1 scoring methodology, but we use LLM-as-judge, in line with other memory frameworks that publish LoCoMo results and our prior research.