Somewhere around week three of running our first long-lived support agent, something odd started happening. It began confidently telling customers about a pricing tier we had retired two months earlier. Nothing was wrong with the model. Nothing was wrong with the prompts. The agent was simply remembering too much, and too much of it was wrong.
If you are building autonomous agents that run for days, weeks, or months, you will hit this wall too. Memory is the feature that makes agents feel intelligent, and unmanaged memory is the thing that slowly makes them dumb.
In this guide, we will build an agentic memory garbage collector (GC) in Node.js and TypeScript. It sits next to your vector store, scores every memory with time-weighted relevance decay, clusters and merges semantically overlapping chunks, and prunes stale episodic embeddings safely, before they exhaust your context window or poison your agent's reasoning.
By the end, you will have working code, a mental model you can adapt to any vector database, and a checklist of mistakes to avoid.

Why Agent Memory Rots
Let us start with the problem, because "just add more memory" is the default instinct and it is usually the wrong one.
A typical agent loop writes to its episodic memory constantly:
- Raw tool outputs (API responses, search results, file listings)
- Conversation turns and summaries
- Intermediate reasoning traces
- Observations about the user or environment
Retrieval then pulls the top-k most similar chunks into the prompt. Here is where things go sideways.
Context distraction
Research on long-context behavior has shown that models do not use all of their context equally well. The well-known "Lost in the Middle" paper found that performance drops when relevant information sits in the middle of a long input. Irrelevant or near-duplicate chunks do not just waste tokens. They actively compete for the model's attention.
When your top-k results contain five slightly different versions of the same weather lookup from five different days, the one that actually matters is buried.
Context poisoning
This is the more dangerous cousin. Context poisoning happens when a retrieved memory is wrong right now even though it was right when stored. Examples:
- A cached tool output says the deployment is "healthy". It was healthy yesterday.
- The user said "I'm on the Basic plan" in March, and upgraded in June.
- An earlier reasoning trace contained a mistaken assumption that the agent later corrected, but the mistaken version is still in the store.
The model has no built-in way to know which memory is current. It sees text and treats it as evidence.
The hallucination cascade
Once a poisoned memory is retrieved, the agent acts on it. That action produces a new tool output or conversation turn, which is then written back into memory. Now the wrong belief has a fresh timestamp and a high similarity score to future queries. The error has been laundered into something that looks new.
That feedback loop is the cascade. And it is why a memory GC is not a nice-to-have cleanup script. It is core infrastructure for any agent that runs longer than a single session.
What a Memory Garbage Collector Actually Does
The name is borrowed from runtime systems for a reason. A good agent memory GC does four things:
- Marks every memory with a retention score based on age, importance, access, and type.
- Merges semantically overlapping memories into a single canonical record.
- Sweeps low-scoring memories into a reversible tombstone state.
- Compacts by permanently deleting tombstones after a grace period.
The key design principle is the same one runtime GCs follow: be conservative and reversible. A GC that deletes the wrong thing is worse than no GC at all.

Architecture and Schema
For this tutorial we will use Node.js 22+, TypeScript, PostgreSQL with the pgvector extension, and the pg driver. pgvector is a pragmatic choice in 2026 because many teams already run Postgres, and it lets us combine vector similarity with ordinary SQL metadata filters in a single query.
Do not worry if you use Qdrant, Pinecone, or Chroma. We will hide the database behind an interface so you can swap it.
The memory table
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE memories (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
agent_id TEXT NOT NULL,
kind TEXT NOT NULL, -- 'tool_output' | 'conversation' | 'fact' | 'preference'
supersession_key TEXT, -- e.g. 'weather:berlin' or 'plan:user_42'
content TEXT NOT NULL,
embedding VECTOR(1536) NOT NULL,
importance REAL NOT NULL DEFAULT 0.5, -- 0..1
access_count INTEGER NOT NULL DEFAULT 0,
pinned BOOLEAN NOT NULL DEFAULT FALSE,
status TEXT NOT NULL DEFAULT 'active', -- 'active' | 'tombstoned'
superseded_by UUID,
tombstoned_at TIMESTAMPTZ,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
last_accessed_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX memories_agent_status_idx ON memories (agent_id, status);
CREATE INDEX memories_supersession_idx ON memories (agent_id, supersession_key)
WHERE supersession_key IS NOT NULL;
CREATE INDEX memories_embedding_idx ON memories
USING hnsw (embedding vector_cosine_ops);
A few columns deserve a quick explanation, because they carry most of the design:
kindlets us give different memory types different lifespans. A tool output should fade in hours. A user preference should last months.supersession_keyis the cheapest and most underrated trick here. When a new memory shares a key with an older one, the older one is obviously stale. No embeddings required.pinnedis your escape hatch for memories that must never be collected.statusandtombstoned_atgive us reversible deletion.
The core types
// src/types.ts
export type MemoryKind = "tool_output" | "conversation" | "fact" | "preference";
export interface Memory {
id: string;
agentId: string;
kind: MemoryKind;
supersessionKey: string | null;
content: string;
embedding: number[];
importance: number;
accessCount: number;
pinned: boolean;
status: "active" | "tombstoned";
createdAt: Date;
lastAccessedAt: Date;
}
export interface GcConfig {
dryRun: boolean;
similarityThreshold: number; // cosine similarity for dedupe, e.g. 0.93
pruneBelowScore: number; // retention score cutoff, e.g. 0.12
maxPruneRatio: number; // never remove more than this fraction per run
graceDays: number; // days a tombstone stays restorable
neighbourLimit: number; // candidates fetched per memory during clustering
}
export const DEFAULT_GC_CONFIG: GcConfig = {
dryRun: true,
similarityThreshold: 0.93,
pruneBelowScore: 0.12,
maxPruneRatio: 0.2,
graceDays: 14,
neighbourLimit: 15,
};
Notice that dryRun defaults to true. That is on purpose. Your first run should never write anything.
Step 1: Time-Weighted Relevance Decay
This is the heart of the system. We want every memory to have a retention score between 0 and 1 that answers a simple question: "How much should we still care about this?"
Our formula combines four signals:
score = importance × recency × reinforcement
- Recency follows exponential decay with a half-life specific to the memory kind.
- Reinforcement rewards memories that keep getting retrieved. If the agent keeps using a memory, it is probably useful.
- Importance is a value from 0 to 1 assigned at write time, either by heuristics or by a cheap LLM call.
This idea has deep roots. The "Generative Agents" paper from Stanford combined recency, importance, and relevance when retrieving memories, and cognitive science has studied forgetting curves since Ebbinghaus. We are applying the same intuition to infrastructure.
Half-lives by memory kind
// src/decay.ts
import type { Memory, MemoryKind } from "./types.js";
const HOUR = 60 * 60 * 1000;
const DAY = 24 * HOUR;
// How long until a memory's recency weight drops to 50%.
export const HALF_LIFE_MS: Record<MemoryKind, number> = {
tool_output: 6 * HOUR,
conversation: 3 * DAY,
fact: 45 * DAY,
preference: 180 * DAY,
};
export function recencyWeight(
kind: MemoryKind,
lastTouched: Date,
now: Date = new Date()
): number {
const ageMs = Math.max(0, now.getTime() - lastTouched.getTime());
const lambda = Math.LN2 / HALF_LIFE_MS[kind];
return Math.exp(-lambda * ageMs);
}
// Diminishing returns: 0 hits -> 1.0, 10 hits -> ~1.7, 100 hits -> ~2.3
export function reinforcement(accessCount: number): number {
return 1 + Math.log10(1 + accessCount) * 0.7;
}
export function retentionScore(m: Memory, now: Date = new Date()): number {
if (m.pinned) return 1;
// Use whichever is more recent: creation or last retrieval.
const lastTouched =
m.lastAccessedAt > m.createdAt ? m.lastAccessedAt : m.createdAt;
const raw =
m.importance *
recencyWeight(m.kind, lastTouched, now) *
reinforcement(m.accessCount);
// Clamp so the score stays comparable across memories.
return Math.min(1, raw);
}
A couple of choices worth defending:
- We measure age from the last time the memory was touched, not just when it was created. A two-month-old fact that is retrieved every day is not stale.
- Reinforcement uses a logarithm so a memory cannot become immortal just because it was retrieved a thousand times by a looping bug.
- Pinned memories short-circuit to 1. No exceptions.
Real-world example
Say your support agent stores three memories:
| Memory | Kind | Age | Importance | Retrievals | Approx. score |
|---|---|---|---|---|---|
| "Order #881 status: shipped" | tool_output | 2 days | 0.4 | 1 | very low |
| "User prefers email over phone" | preference | 60 days | 0.8 | 6 | high |
| "Discussed refund policy" | conversation | 10 days | 0.5 | 0 | low |
The tool output has gone through roughly eight half-lives, so its recency weight is nearly zero. The preference has barely aged against its 180-day half-life. That is exactly the behavior you want.
Step 2: Supersession, the Cheapest Win
Before reaching for embeddings, handle the obvious case. If your agent writes a memory about the same entity again, the old one is outdated.
// src/supersede.ts
import type { Pool } from "pg";
export async function supersedeByKey(
pool: Pool,
agentId: string,
dryRun: boolean
): Promise<number> {
const sql = `
WITH ranked AS (
SELECT id,
ROW_NUMBER() OVER (
PARTITION BY supersession_key
ORDER BY created_at DESC
) AS rn,
FIRST_VALUE(id) OVER (
PARTITION BY supersession_key
ORDER BY created_at DESC
) AS newest_id
FROM memories
WHERE agent_id = $1
AND status = 'active'
AND pinned = FALSE
AND supersession_key IS NOT NULL
)
SELECT id, newest_id FROM ranked WHERE rn > 1
`;
const { rows } = await pool.query(sql, [agentId]);
if (dryRun || rows.length === 0) return rows.length;
await pool.query(
`UPDATE memories
SET status = 'tombstoned',
tombstoned_at = now(),
superseded_by = r.newest_id
FROM (SELECT unnest($1::uuid[]) AS id,
unnest($2::uuid[]) AS newest_id) r
WHERE memories.id = r.id`,
[rows.map((r) => r.id), rows.map((r) => r.newest_id)]
);
return rows.length;
}
To use this well, you need to be deliberate when your agent writes memories. Give tool outputs a key like tool:get_order_status:881 or weather:berlin. It costs nothing at write time, and it removes the single largest category of stale memory before any similarity math happens.
Step 3: Semantic Clustering and Deduplication
Not everything has a clean key. Conversational memories drift into near-duplicates: "The user wants a refund", "Customer is asking for money back", "User requested a refund for order 881". Same fact, three chunks, three slots in your top-k.
We will cluster these using nearest-neighbour search and a similarity threshold, then merge each cluster into one canonical memory.
Why not run a full clustering algorithm?
You could run HDBSCAN or k-means over every embedding. For a nightly batch job on a modest store, that is fine. But for a Node.js service that wants to run incrementally, a greedy neighbour-based approach with union-find is simpler, predictable, and easy to explain to your team. It also lets the database's HNSW index do the heavy lifting.

A small union-find helper
// src/unionFind.ts
export class UnionFind {
private parent = new Map<string, string>();
find(x: string): string {
if (!this.parent.has(x)) this.parent.set(x, x);
let root = x;
while (this.parent.get(root) !== root) root = this.parent.get(root)!;
// Path compression
let cur = x;
while (cur !== root) {
const next = this.parent.get(cur)!;
this.parent.set(cur, root);
cur = next;
}
return root;
}
union(a: string, b: string): void {
const ra = this.find(a);
const rb = this.find(b);
if (ra !== rb) this.parent.set(rb, ra);
}
groups(): Map<string, string[]> {
const out = new Map<string, string[]>();
for (const id of this.parent.keys()) {
const root = this.find(id);
const list = out.get(root) ?? [];
list.push(id);
out.set(root, list);
}
return out;
}
}
Finding neighbours with pgvector
pgvector's <=> operator returns cosine distance, so similarity is 1 - distance.
// src/neighbours.ts
import type { Pool } from "pg";
export interface Neighbour {
id: string;
similarity: number;
}
export async function findNeighbours(
pool: Pool,
agentId: string,
memoryId: string,
limit: number,
minSimilarity: number
): Promise<Neighbour[]> {
const sql = `
SELECT b.id,
1 - (b.embedding <=> a.embedding) AS similarity
FROM memories a
JOIN memories b
ON b.agent_id = a.agent_id
AND b.id <> a.id
AND b.status = 'active'
AND b.kind = a.kind
WHERE a.id = $1
AND a.agent_id = $2
ORDER BY b.embedding <=> a.embedding
LIMIT $3
`;
const { rows } = await pool.query(sql, [memoryId, agentId, limit]);
return rows
.map((r) => ({ id: r.id as string, similarity: Number(r.similarity) }))
.filter((n) => n.similarity >= minSimilarity);
}
Notice we only compare memories of the same kind. A preference and a tool output can look semantically similar without being interchangeable, and merging across kinds is a classic way to lose information.
Choosing the canonical memory and merging
For each cluster, we keep the memory with the highest retention score and fold the others into it. Access counts are summed so the canonical record inherits the "this keeps being useful" signal.
// src/dedupe.ts
import type { Pool } from "pg";
import type { GcConfig, Memory } from "./types.js";
import { retentionScore } from "./decay.js";
import { UnionFind } from "./unionFind.js";
import { findNeighbours } from "./neighbours.js";
export interface MergePlan {
canonicalId: string;
duplicateIds: string[];
mergedAccessCount: number;
mergedImportance: number;
}
export async function planDedupe(
pool: Pool,
agentId: string,
memories: Memory[],
cfg: GcConfig
): Promise<MergePlan[]> {
const uf = new UnionFind();
const byId = new Map(memories.map((m) => [m.id, m]));
for (const m of memories) {
if (m.pinned) continue; // pinned memories are never merged away
const neighbours = await findNeighbours(
pool,
agentId,
m.id,
cfg.neighbourLimit,
cfg.similarityThreshold
);
for (const n of neighbours) {
const other = byId.get(n.id);
if (other && !other.pinned) uf.union(m.id, n.id);
}
}
const plans: MergePlan[] = [];
for (const ids of uf.groups().values()) {
if (ids.length < 2) continue;
const members = ids.map((id) => byId.get(id)!);
members.sort((a, b) => retentionScore(b) - retentionScore(a));
const [canonical, ...rest] = members;
plans.push({
canonicalId: canonical.id,
duplicateIds: rest.map((r) => r.id),
mergedAccessCount: members.reduce((sum, m) => sum + m.accessCount, 0),
mergedImportance: Math.max(...members.map((m) => m.importance)),
});
}
return plans;
}
export async function applyDedupe(
pool: Pool,
plans: MergePlan[]
): Promise<void> {
const client = await pool.connect();
try {
await client.query("BEGIN");
for (const p of plans) {
await client.query(
`UPDATE memories
SET access_count = $2,
importance = $3,
last_accessed_at = now()
WHERE id = $1`,
[p.canonicalId, p.mergedAccessCount, p.mergedImportance]
);
await client.query(
`UPDATE memories
SET status = 'tombstoned',
tombstoned_at = now(),
superseded_by = $1
WHERE id = ANY($2::uuid[])`,
[p.canonicalId, p.duplicateIds]
);
}
await client.query("COMMIT");
} catch (err) {
await client.query("ROLLBACK");
throw err;
} finally {
client.release();
}
}
Two things to note. First, everything is wrapped in a transaction, so a crash halfway through never leaves a cluster half-merged. Second, duplicates are tombstoned, not deleted, and they point at their canonical memory through superseded_by. That gives you a full audit trail.
An optional upgrade: LLM-assisted merging
Picking the highest-scoring member is cheap and usually good enough. If your clusters contain complementary details, for example one chunk has the order number and another has the refund reason, you can pass the cluster to a small, inexpensive model and ask for a single consolidated memory. Re-embed the result and write it as the new canonical record. Just keep the originals tombstoned so you can verify the merge later.
Step 4: Safe Pruning
Now we use the retention score to sweep out low-value memories. This is the step where careless engineering causes real damage, so we add several guardrails.
// src/prune.ts
import type { Pool } from "pg";
import type { GcConfig, Memory } from "./types.js";
import { retentionScore } from "./decay.js";
export interface PrunePlan {
candidateIds: string[];
skippedByCap: number;
}
export function planPrune(
memories: Memory[],
cfg: GcConfig,
now: Date = new Date()
): PrunePlan {
const scored = memories
.filter((m) => !m.pinned)
.map((m) => ({ id: m.id, score: retentionScore(m, now) }))
.filter((m) => m.score < cfg.pruneBelowScore)
.sort((a, b) => a.score - b.score); // lowest first
// Guardrail: never remove more than maxPruneRatio of the active store.
const cap = Math.floor(memories.length * cfg.maxPruneRatio);
const selected = scored.slice(0, cap);
return {
candidateIds: selected.map((s) => s.id),
skippedByCap: Math.max(0, scored.length - selected.length),
};
}
export async function applyPrune(pool: Pool, plan: PrunePlan): Promise<void> {
if (plan.candidateIds.length === 0) return;
await pool.query(
`UPDATE memories
SET status = 'tombstoned', tombstoned_at = now()
WHERE id = ANY($1::uuid[]) AND pinned = FALSE`,
[plan.candidateIds]
);
}
export async function compact(pool: Pool, agentId: string, cfg: GcConfig) {
const { rowCount } = await pool.query(
`DELETE FROM memories
WHERE agent_id = $1
AND status = 'tombstoned'
AND tombstoned_at < now() - ($2 || ' days')::interval`,
[agentId, String(cfg.graceDays)]
);
return rowCount ?? 0;
}
That maxPruneRatio cap is a seatbelt. If a bug in your importance scoring suddenly marks everything as worthless, the GC will still only touch 20 percent of the store per run, and you will have time to notice.
Step 5: Orchestrating the Full GC Run
Now we wire the phases together into one function that returns a report. Reports matter. A GC you cannot observe is a GC you cannot trust.
// src/gc.ts
import type { Pool } from "pg";
import { DEFAULT_GC_CONFIG, type GcConfig, type Memory } from "./types.js";
import { supersedeByKey } from "./supersede.js";
import { planDedupe, applyDedupe } from "./dedupe.js";
import { planPrune, applyPrune, compact } from "./prune.js";
export interface GcReport {
agentId: string;
dryRun: boolean;
activeBefore: number;
superseded: number;
mergedClusters: number;
duplicatesRemoved: number;
pruned: number;
prunedSkippedByCap: number;
compacted: number;
durationMs: number;
}
async function loadActive(pool: Pool, agentId: string): Promise<Memory[]> {
const { rows } = await pool.query(
`SELECT id, agent_id, kind, supersession_key, content, importance,
access_count, pinned, status, created_at, last_accessed_at
FROM memories
WHERE agent_id = $1 AND status = 'active'`,
[agentId]
);
return rows.map((r) => ({
id: r.id,
agentId: r.agent_id,
kind: r.kind,
supersessionKey: r.supersession_key,
content: r.content,
embedding: [], // not needed in memory; neighbours are queried in SQL
importance: r.importance,
accessCount: r.access_count,
pinned: r.pinned,
status: r.status,
createdAt: r.created_at,
lastAccessedAt: r.last_accessed_at,
}));
}
export async function runMemoryGc(
pool: Pool,
agentId: string,
overrides: Partial<GcConfig> = {}
): Promise<GcReport> {
const cfg = { ...DEFAULT_GC_CONFIG, ...overrides };
const started = Date.now();
// Phase 0: take an advisory lock so two GC runs never overlap.
const lockKey = Buffer.from(agentId).reduce((a, b) => (a * 31 + b) | 0, 7);
const { rows } = await pool.query("SELECT pg_try_advisory_lock($1) AS ok", [
lockKey,
]);
if (!rows[0].ok) throw new Error(`GC already running for ${agentId}`);
try {
const activeBefore = (await loadActive(pool, agentId)).length;
// Phase 1: supersession by key (cheap, deterministic)
const superseded = await supersedeByKey(pool, agentId, cfg.dryRun);
// Phase 2: semantic dedupe on whatever is still active
const afterSupersede = await loadActive(pool, agentId);
const plans = await planDedupe(pool, agentId, afterSupersede, cfg);
if (!cfg.dryRun) await applyDedupe(pool, plans);
// Phase 3: decay-based prune
const afterDedupe = cfg.dryRun ? afterSupersede : await loadActive(pool, agentId);
const prunePlan = planPrune(afterDedupe, cfg);
if (!cfg.dryRun) await applyPrune(pool, prunePlan);
// Phase 4: permanently remove old tombstones
const compacted = cfg.dryRun ? 0 : await compact(pool, agentId, cfg);
return {
agentId,
dryRun: cfg.dryRun,
activeBefore,
superseded,
mergedClusters: plans.length,
duplicatesRemoved: plans.reduce((n, p) => n + p.duplicateIds.length, 0),
pruned: prunePlan.candidateIds.length,
prunedSkippedByCap: prunePlan.skippedByCap,
compacted,
durationMs: Date.now() - started,
};
} finally {
await pool.query("SELECT pg_advisory_unlock($1)", [lockKey]);
}
}
Scheduling it
For a single service, node-cron is plenty. For multi-instance deployments, the advisory lock above prevents duplicate runs, or you can move the job to your existing queue or scheduler.
// src/scheduler.ts
import cron from "node-cron";
import { Pool } from "pg";
import { runMemoryGc } from "./gc.js";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });
// Nightly at 03:15, writes enabled only after you have validated dry runs.
cron.schedule("15 3 * * *", async () => {
const agents = await pool.query("SELECT DISTINCT agent_id FROM memories");
for (const { agent_id } of agents.rows) {
try {
const report = await runMemoryGc(pool, agent_id, {
dryRun: process.env.GC_ENABLE_WRITES !== "true",
});
console.info("[memory-gc]", JSON.stringify(report));
} catch (err) {
console.error("[memory-gc] failed", agent_id, err);
}
}
});

Step 6: Apply Decay at Retrieval Time Too
Here is a point many teams miss. The GC cleans the store periodically, but between runs, stale memories are still sitting there waiting to be retrieved. The fix is to apply the same decay logic to ranking.
Instead of sorting purely by vector similarity, fetch a larger candidate set, rescore it, and then pick the final top-k. Bumping access_count on the memories you actually use closes the reinforcement loop.
// src/retrieve.ts
import type { Pool } from "pg";
import type { Memory } from "./types.js";
import { retentionScore } from "./decay.js";
export async function retrieve(
pool: Pool,
agentId: string,
queryEmbedding: number[],
k = 6
) {
const { rows } = await pool.query(
`SELECT id, kind, content, importance, access_count, pinned,
created_at, last_accessed_at,
1 - (embedding <=> $2::vector) AS similarity
FROM memories
WHERE agent_id = $1 AND status = 'active'
ORDER BY embedding <=> $2::vector
LIMIT $3`,
[agentId, JSON.stringify(queryEmbedding), k * 4]
);
const reranked = rows
.map((r) => {
const m = {
kind: r.kind,
importance: r.importance,
accessCount: r.access_count,
pinned: r.pinned,
createdAt: r.created_at,
lastAccessedAt: r.last_accessed_at,
} as Memory;
// Blend semantic match with retention. Tune the weights on your own evals.
const blended = 0.7 * Number(r.similarity) + 0.3 * retentionScore(m);
return { id: r.id as string, content: r.content as string, blended };
})
.sort((a, b) => b.blended - a.blended)
.slice(0, k);
await pool.query(
`UPDATE memories
SET access_count = access_count + 1, last_accessed_at = now()
WHERE id = ANY($1::uuid[])`,
[reranked.map((r) => r.id)]
);
return reranked;
}
The 70/30 blend is a starting point, not a law. Build a small evaluation set of real agent queries with known correct memories, and tune the weights against it.
Best Practices
After running this pattern in a few production agents, these habits have paid off the most:
- Always start in dry-run mode. Log what would have been removed for a week and read through a sample by hand.
- Write memories with intent. Assign a
kind, animportance, and asupersession_keyat write time. Cleaning up is much harder than labeling correctly in the first place. - Separate memory tiers. Keep short-lived working memory, episodic memory, and durable facts or preferences apart. They deserve different half-lives and different rules.
- Pin what must survive. Compliance notes, user-stated hard constraints, and safety instructions should never be eligible for collection.
- Cap every run. A maximum prune ratio converts a catastrophic bug into a minor annoyance.
- Keep tombstones restorable. A 7 to 14 day grace period has saved us from more than one bad threshold.
- Log structured reports. Track active count, merge count, prune count, and the distribution of retention scores over time.
- Scope by tenant. In multi-user systems, every query, cluster, and merge must be filtered by agent and user. A cross-tenant merge is a privacy incident.
- Test with a golden set. Keep a list of memories that must still be retrievable after a GC run, and assert on it in CI.
Common Mistakes to Avoid
We have made most of these ourselves, so consider this a friendly warning.
1. Using one global TTL
A single expiry for everything is simple and wrong. It deletes your user's long-term preferences at the same time as yesterday's search results.
2. Setting the similarity threshold too low
At 0.80, "User wants to cancel their subscription" and "User asked how to pause their subscription" may merge into one. Those are very different intents. Start high, around 0.93, and loosen only with evidence.
3. Merging across kinds or tenants
Always filter neighbours by kind and by owner. We showed it in the SQL above, and it is worth repeating.
4. Deleting immediately
Hard deletes on the first pass leave you no way to recover from a bad score. Tombstone first, compact later.
5. Ignoring the feedback loop
If you only clean up the store but keep writing raw tool outputs back verbatim, you are mopping the floor with the tap running. Summarize or key tool outputs before they are stored.
6. Trusting LLM-assigned importance blindly
Model-assigned importance scores are noisy and often cluster around 0.7 to 0.8. Combine them with simple heuristics like memory kind and explicit user statements, and spot-check the distribution.
7. Never measuring the effect
If you cannot show that retrieval precision improved or token usage dropped after the GC, you are guessing. Track retrieval hit quality and average prompt size before and after.
🚀 Pro Tips
- Embed once, reuse everywhere. Store the embedding model name and version in metadata. When you upgrade models, similarity thresholds change, so re-tune before re-enabling aggressive dedupe.
- Use a "contradiction check" for facts. When a new fact is highly similar to an older one but the content differs, for example plan names or dates, flag the older one as superseded instead of merging.
- Add a cheap recall probe. After each GC run, replay a handful of known queries and alert if any expected memory disappears from the top results.
- Compress before you collect. For conversation memories that are borderline, summarize them into a durable fact rather than deleting them outright.
- Make half-lives configurable per agent. A coding agent and a customer support agent age information very differently.
- Expose a "why was this removed" endpoint. The
superseded_bylink and tombstone timestamps make debugging a two-minute job instead of a two-hour one. - Watch the long tail. If a large share of your store sits just above the prune threshold, your importance scores are probably too generous.
📌 Key Takeaways
- Long-running agents fill their vector stores with stale tool outputs and near-duplicate chunks, causing context distraction and hallucination cascades.
- A memory GC is built from three mechanisms: time-weighted relevance decay, semantic clustering and deduplication, and safe, reversible pruning.
- Supersession keys are the cheapest way to eliminate stale facts, and they require no embeddings at all.
- Use kind-specific half-lives so short-lived tool outputs fade fast while user preferences persist.
- Guardrails make the system trustworthy: dry-run defaults, pinned memories, a per-run prune cap, tombstones with a grace period, advisory locks, and transactional merges.
- Apply decay at retrieval time as well as during cleanup, so stale memories lose ranking between GC runs.
- Measure everything. Report counts, track retrieval quality, and replay known queries after each run.
Conclusion
Agent memory is often sold as an append-only feature: store everything, retrieve what is relevant, and let the model sort it out. In practice, that approach works for a demo and quietly fails in production. The longer your agent runs, the more its past starts to interfere with its present.
A memory garbage collector gives your agent something humans rely on every day, which is the ability to forget well. Decay lets old information fade, deduplication keeps repeated facts from crowding the prompt, and careful, reversible pruning keeps the whole system healthy without risking what matters.
The code in this guide is intentionally small, and you can run it against a staging database this afternoon. Start in dry-run mode, read the reports, tune your thresholds on real data, and only then let it write. Your agent, and the customers on the other end of it, will notice the difference.
References
- Liu, N. F., et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. arxiv.org/abs/2307.03172
- Park, J. S., et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. arxiv.org/abs/2304.03442
- Packer, C., et al. (2023). MemGPT: Towards LLMs as Operating Systems. arxiv.org/abs/2310.08560
- pgvector documentation and README. github.com/pgvector/pgvector
- node-postgres (
pg) documentation. node-postgres.com - PostgreSQL documentation: Advisory Locks. postgresql.org/docs/current/explicit-locking.html
- Anthropic Engineering. Effective Context Engineering for AI Agents. anthropic.com/engineering
- Campello, R., Moulavi, D., Sander, J. (2013). Density-Based Clustering Based on Hierarchical Density Estimates (HDBSCAN).
- Ebbinghaus, H. (1885). Memory: A Contribution to Experimental Psychology.