Fetching latest headlines…

Dev

I Built a Cross-Client Memory Hub for AI Agents — Here's What I Learned

Dev.toUnited States · NORTH AMERICA

I use Claude Code for coding, Cursor for refactoring, and Windsurf for exploration. Each has its own memory. Switch tools and my AI forgets everything. So I built MemTether — a local-first memory hub...

0 views0 likes0 comments

I use Claude Code for coding, Cursor for refactoring, and Windsurf for exploration. Each has its own memory. Switch tools and my AI forgets everything.

So I built MemTether — a local-first memory hub that lets 23+ AI clients share one physical SQLite database.

Here's what I learned building it.

The Problem Is Simpler Than You Think

Existing memory solutions (mem0, cognee, zep) all treat memory as a service. You send memories to their cloud API, they store and retrieve them. This works, but it means:

  • Your memories live on someone else's server
  • You pay per API call
  • You need an API key even for a local project
  • You can't inspect the raw data

My insight: if all your AI tools run on the same machine, you don't need a service — you need a file. Just make them all point to the same SQLite database.

No cloud. No API fees. No abstraction layer. One memory.db, 23 clients reading and writing to it.

Design Decisions That Matter

1. Supersession, Not Deletion

When a memory needs updating, I don't delete the old one. I mark it superseded and create a new version. This means you can always trace "what did we believe before we learned X?"

This sounds simple, but it changes everything about how the system works. Search must filter out superseded entries. FTS5 needs triggers to auto-sync. The projection system needs to prefer the latest version.

2. Bi-Temporal: Two Clocks, Not One

Every memory has two timestamps:

  • T (valid time): when the fact was true in the real world
  • T′ (recorded time): when the system learned about it

Example: an API key expired on September 15th, but I didn't notice until September 18th. Querying "what did we know on September 15th?" returns "the key is valid" (which is what the system believed). Querying "what was actually true?" returns "expired."

Without this distinction, you get retroactive truth bias — using today's knowledge to judge yesterday's decisions.

3. Q-Value: Memories That Get Used Should Rank Higher

Most memory systems rank by recency or semantic similarity. But a memory from 6 months ago that gets used every day is more valuable than one from yesterday that's never been retrieved.

So I added a Q-Value (inspired by reinforcement learning): every time a memory is searched and actually useful, its Q-Value increases. Next search, it ranks higher. Simple, effective, and nobody else does it.

4. Four-Factor Re-Ranking

Raw semantic similarity isn't enough. I blend four factors:

  • Semantic similarity (45%)
  • Recency (25%)
  • Usage frequency (5%)
  • Type importance (10%)

Then normalize with z-score + sigmoid and blend 70/30 with the RRF score. This fixed asset lookup queries that pure semantic search missed.

5. SQLite Triggers for FTS Sync (Not Python)

My first version had Python code to sync the FTS5 full-text index. It had except Exception: pass around the sync calls. Result: 254 stale entries and 260 missing entries.

The fix: SQLite triggers. Three triggers (INSERT, DELETE, UPDATE) at the SQL level. No Python code can accidentally skip them. The inconsistency went from 514 entries to zero.

Lesson: if SQLite can do it at the SQL level, don't do it in Python.

The Hard Part Isn't Code — It's Packaging

Writing the memory engine took 3 weeks. Making it installable took 2 months:

  • Wheel building with correct py-modules (flat layout, not packages)
  • check_packaging.py to catch version drift and missing modules
  • FTS5 triggers in the right place (inside init_db, not floating in the schema string)
  • Console entry points (memtether, memtether-connect)
  • Optional dependencies ([vector] for chromadb, [server] for fastapi)
  • A Quick Start that actually works (I initially wrote commands that didn't exist)

I'm a solo developer. Every hour spent on packaging is an hour not spent on features. But without packaging, nobody can use your features.

Honest Benchmark Numbers

I ran LongMemEval (500 questions, full run):

Metric Score
Strict match 62.6%
LLM judge 54.6%
Multi-session strict 48.8%
Multi-session LLM judge 60.0%

The multi-session gap (strict vs judge) is the most interesting finding. When an answer is computed (like "3 weeks" from multiple data points), strict substring matching fails because the computed answer doesn't appear verbatim in any single memory. The LLM judge is more forgiving.

I also built an E-Hybrid method (session summaries + flat evidence) that improved multi-session strict from 48.8% to 71.4% on a 15-question test. Small sample, but promising.

I'm not going to claim these are better than mem0's 94.4% — different harness, different methodology. But 48.8% strict is above the industry average of 27.9%.

What's Next

  • MCP Registry submission (done: PR #4946)
  • Community building (this blog post is part of that)
  • NLPCC 2027 paper submission (CCF C, deadline ~April 2027)
  • More examples and better docs

Try It

pip install memtether
memtether init
memtether connect --all

GitHub: MemTether/MemTether
PyPI: memtether

Comments (0)

Sign in to join the discussion

Be the first to comment!