Claude-Mem explained: what a coding agent remembers
Claude-Mem performance and cost: evaluate useful recall
Measure provider tokens, retrieval relevance and operational work on real tasks
What you will learn
- List the spend drivers
- Measure quality before savings
- Watch the queue and storage
Before you start
- A synthetic two-session project
- One supported host
- A deliberate memory provider choice
Use two synthetic sessions and an explicit deletion check before team adoption
Key takeaways
- Upstream savings claims are unverified here.
- Relevance matters more than raw memory volume.
- Queue and storage growth affect operating cost.
List the spend drivers
Observation generation can use a configured model or provider. The local worker stores structured records and vector data; remote or hosted choices may add separate service costs.
The README’s token-savings and trial language and the production guide’s historical savings percentage are upstream claims. We did not reproduce them or check current commercial terms.
Measure quality before savings
For a small project, compare a second-session task with memory disabled and enabled. Count relevant facts recalled, incorrect insertions, prompt tokens and human corrections.
`ContextBuilder.ts` trims selected observations to an output budget; the resulting statistics describe that selected set. Actual savings still depend on task, history and client, so avoid one universal percentage.
Watch the queue and storage
The production guide suggests tracking pending messages, failed work, active sessions, SQLite WAL and Chroma growth. These are useful diagnostic categories even when its threshold numbers do not fit your workload.
Record cold start, observer latency, retrieval latency and storage growth separately. No throughput or cost benchmark was run for this article.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Choose a repeated two-session task.
- 2
Record useful and misleading recalled facts.
- 3
Measure tokens, latency and storage with the same provider.
Copy-ready example
task,provider,baseline_tokens,memory_tokens,useful_facts,wrong_facts,review_minutes
trial,,,,,,not-runFrequently asked questions
Does the README’s saving figure apply to my project?
No measurement in this series supports transferring that figure; run a matched task test.
Is more stored memory always better?
Irrelevant or incorrect records can increase review and prompt costs.
Sources
- Claude-Mem / README.mdSource checked 2026-10-08
- Claude-Mem / docs/production-guide.mdSource checked 2026-10-08
- Claude-Mem / docs/architecture-overview.mdSource checked 2026-10-08
- Claude-Mem / docs/api.mdSource checked 2026-10-08
- Claude-Mem / src/services/context/ContextBuilder.tsSource checked 2026-10-08