Skip to main content
Back to Blog
AI AgentsNode.jsVector DatabaseAgent MemorypgvectorContext EngineeringRAGTypeScript

How to Build an Agentic Memory Garbage Collector in Node.js to Prevent Vector Store Context Poisoning

Learn how to build a Node.js memory garbage collector for AI agents. Implement time-weighted relevance decay, semantic deduplication, and safe pruning to stop vector store context poisoning and hallucination cascades.

October 1, 202624 min readNiraj Kumar

Somewhere around week three of running our first long-lived support agent, something odd started happening. It began confidently telling customers about a pricing tier we had retired two months earlier. Nothing was wrong with the model. Nothing was wrong with the prompts. The agent was simply remembering too much, and too much of it was wrong.

If you are building autonomous agents that run for days, weeks, or months, you will hit this wall too. Memory is the feature that makes agents feel intelligent, and unmanaged memory is the thing that slowly makes them dumb.

In this guide, we will build an agentic memory garbage collector (GC) in Node.js and TypeScript. It sits next to your vector store, scores every memory with time-weighted relevance decay, clusters and merges semantically overlapping chunks, and prunes stale episodic embeddings safely, before they exhaust your context window or poison your agent's reasoning.

By the end, you will have working code, a mental model you can adapt to any vector database, and a checklist of mistakes to avoid.

Diagram showing an agent's vector store filling with stale and duplicate memories, leading to a polluted context window

Why Agent Memory Rots

Let us start with the problem, because "just add more memory" is the default instinct and it is usually the wrong one.

A typical agent loop writes to its episodic memory constantly:

  • Raw tool outputs (API responses, search results, file listings)
  • Conversation turns and summaries
  • Intermediate reasoning traces
  • Observations about the user or environment

Retrieval then pulls the top-k most similar chunks into the prompt. Here is where things go sideways.

Context distraction

Research on long-context behavior has shown that models do not use all of their context equally well. The well-known "Lost in the Middle" paper found that performance drops when relevant information sits in the middle of a long input. Irrelevant or near-duplicate chunks do not just waste tokens. They actively compete for the model's attention.

When your top-k results contain five slightly different versions of the same weather lookup from five different days, the one that actually matters is buried.

Context poisoning

This is the more dangerous cousin. Context poisoning happens when a retrieved memory is wrong right now even though it was right when stored. Examples:

  • A cached tool output says the deployment is "healthy". It was healthy yesterday.
  • The user said "I'm on the Basic plan" in March, and upgraded in June.
  • An earlier reasoning trace contained a mistaken assumption that the agent later corrected, but the mistaken version is still in the store.

The model has no built-in way to know which memory is current. It sees text and treats it as evidence.

The hallucination cascade

Once a poisoned memory is retrieved, the agent acts on it. That action produces a new tool output or conversation turn, which is then written back into memory. Now the wrong belief has a fresh timestamp and a high similarity score to future queries. The error has been laundered into something that looks new.

That feedback loop is the cascade. And it is why a memory GC is not a nice-to-have cleanup script. It is core infrastructure for any agent that runs longer than a single session.

What a Memory Garbage Collector Actually Does

The name is borrowed from runtime systems for a reason. A good agent memory GC does four things:

  1. Marks every memory with a retention score based on age, importance, access, and type.
  2. Merges semantically overlapping memories into a single canonical record.
  3. Sweeps low-scoring memories into a reversible tombstone state.
  4. Compacts by permanently deleting tombstones after a grace period.

The key design principle is the same one runtime GCs follow: be conservative and reversible. A GC that deletes the wrong thing is worse than no GC at all.

Flow diagram of the four GC phases: mark, merge, sweep, compact

Architecture and Schema

For this tutorial we will use Node.js 22+, TypeScript, PostgreSQL with the pgvector extension, and the pg driver. pgvector is a pragmatic choice in 2026 because many teams already run Postgres, and it lets us combine vector similarity with ordinary SQL metadata filters in a single query.

Do not worry if you use Qdrant, Pinecone, or Chroma. We will hide the database behind an interface so you can swap it.

The memory table

CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE memories (
  id               UUID PRIMARY KEY DEFAULT gen_random_uuid(),
  agent_id         TEXT        NOT NULL,
  kind             TEXT        NOT NULL,      -- 'tool_output' | 'conversation' | 'fact' | 'preference'
  supersession_key TEXT,                       -- e.g. 'weather:berlin' or 'plan:user_42'
  content          TEXT        NOT NULL,
  embedding        VECTOR(1536) NOT NULL,
  importance       REAL        NOT NULL DEFAULT 0.5,   -- 0..1
  access_count     INTEGER     NOT NULL DEFAULT 0,
  pinned           BOOLEAN     NOT NULL DEFAULT FALSE,
  status           TEXT        NOT NULL DEFAULT 'active', -- 'active' | 'tombstoned'
  superseded_by    UUID,
  tombstoned_at    TIMESTAMPTZ,
  created_at       TIMESTAMPTZ NOT NULL DEFAULT now(),
  last_accessed_at TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX memories_agent_status_idx ON memories (agent_id, status);
CREATE INDEX memories_supersession_idx ON memories (agent_id, supersession_key)
  WHERE supersession_key IS NOT NULL;
CREATE INDEX memories_embedding_idx ON memories
  USING hnsw (embedding vector_cosine_ops);

A few columns deserve a quick explanation, because they carry most of the design:

  • kind lets us give different memory types different lifespans. A tool output should fade in hours. A user preference should last months.
  • supersession_key is the cheapest and most underrated trick here. When a new memory shares a key with an older one, the older one is obviously stale. No embeddings required.
  • pinned is your escape hatch for memories that must never be collected.
  • status and tombstoned_at give us reversible deletion.

The core types

// src/types.ts
export type MemoryKind = "tool_output" | "conversation" | "fact" | "preference";

export interface Memory {
  id: string;
  agentId: string;
  kind: MemoryKind;
  supersessionKey: string | null;
  content: string;
  embedding: number[];
  importance: number;
  accessCount: number;
  pinned: boolean;
  status: "active" | "tombstoned";
  createdAt: Date;
  lastAccessedAt: Date;
}

export interface GcConfig {
  dryRun: boolean;
  similarityThreshold: number;   // cosine similarity for dedupe, e.g. 0.93
  pruneBelowScore: number;       // retention score cutoff, e.g. 0.12
  maxPruneRatio: number;         // never remove more than this fraction per run
  graceDays: number;             // days a tombstone stays restorable
  neighbourLimit: number;        // candidates fetched per memory during clustering
}

export const DEFAULT_GC_CONFIG: GcConfig = {
  dryRun: true,
  similarityThreshold: 0.93,
  pruneBelowScore: 0.12,
  maxPruneRatio: 0.2,
  graceDays: 14,
  neighbourLimit: 15,
};

Notice that dryRun defaults to true. That is on purpose. Your first run should never write anything.

Step 1: Time-Weighted Relevance Decay

This is the heart of the system. We want every memory to have a retention score between 0 and 1 that answers a simple question: "How much should we still care about this?"

Our formula combines four signals:

score = importance × recency × reinforcement
  • Recency follows exponential decay with a half-life specific to the memory kind.
  • Reinforcement rewards memories that keep getting retrieved. If the agent keeps using a memory, it is probably useful.
  • Importance is a value from 0 to 1 assigned at write time, either by heuristics or by a cheap LLM call.

This idea has deep roots. The "Generative Agents" paper from Stanford combined recency, importance, and relevance when retrieving memories, and cognitive science has studied forgetting curves since Ebbinghaus. We are applying the same intuition to infrastructure.

Half-lives by memory kind

// src/decay.ts
import type { Memory, MemoryKind } from "./types.js";

const HOUR = 60 * 60 * 1000;
const DAY = 24 * HOUR;

// How long until a memory's recency weight drops to 50%.
export const HALF_LIFE_MS: Record<MemoryKind, number> = {
  tool_output: 6 * HOUR,
  conversation: 3 * DAY,
  fact: 45 * DAY,
  preference: 180 * DAY,
};

export function recencyWeight(
  kind: MemoryKind,
  lastTouched: Date,
  now: Date = new Date()
): number {
  const ageMs = Math.max(0, now.getTime() - lastTouched.getTime());
  const lambda = Math.LN2 / HALF_LIFE_MS[kind];
  return Math.exp(-lambda * ageMs);
}

// Diminishing returns: 0 hits -> 1.0, 10 hits -> ~1.7, 100 hits -> ~2.3
export function reinforcement(accessCount: number): number {
  return 1 + Math.log10(1 + accessCount) * 0.7;
}

export function retentionScore(m: Memory, now: Date = new Date()): number {
  if (m.pinned) return 1;

  // Use whichever is more recent: creation or last retrieval.
  const lastTouched =
    m.lastAccessedAt > m.createdAt ? m.lastAccessedAt : m.createdAt;

  const raw =
    m.importance *
    recencyWeight(m.kind, lastTouched, now) *
    reinforcement(m.accessCount);

  // Clamp so the score stays comparable across memories.
  return Math.min(1, raw);
}

A couple of choices worth defending:

  • We measure age from the last time the memory was touched, not just when it was created. A two-month-old fact that is retrieved every day is not stale.
  • Reinforcement uses a logarithm so a memory cannot become immortal just because it was retrieved a thousand times by a looping bug.
  • Pinned memories short-circuit to 1. No exceptions.

Real-world example

Say your support agent stores three memories:

MemoryKindAgeImportanceRetrievalsApprox. score
"Order #881 status: shipped"tool_output2 days0.41very low
"User prefers email over phone"preference60 days0.86high
"Discussed refund policy"conversation10 days0.50low

The tool output has gone through roughly eight half-lives, so its recency weight is nearly zero. The preference has barely aged against its 180-day half-life. That is exactly the behavior you want.

Step 2: Supersession, the Cheapest Win

Before reaching for embeddings, handle the obvious case. If your agent writes a memory about the same entity again, the old one is outdated.

// src/supersede.ts
import type { Pool } from "pg";

export async function supersedeByKey(
  pool: Pool,
  agentId: string,
  dryRun: boolean
): Promise<number> {
  const sql = `
    WITH ranked AS (
      SELECT id,
             ROW_NUMBER() OVER (
               PARTITION BY supersession_key
               ORDER BY created_at DESC
             ) AS rn,
             FIRST_VALUE(id) OVER (
               PARTITION BY supersession_key
               ORDER BY created_at DESC
             ) AS newest_id
      FROM memories
      WHERE agent_id = $1
        AND status = 'active'
        AND pinned = FALSE
        AND supersession_key IS NOT NULL
    )
    SELECT id, newest_id FROM ranked WHERE rn > 1
  `;
  const { rows } = await pool.query(sql, [agentId]);

  if (dryRun || rows.length === 0) return rows.length;

  await pool.query(
    `UPDATE memories
        SET status = 'tombstoned',
            tombstoned_at = now(),
            superseded_by = r.newest_id
       FROM (SELECT unnest($1::uuid[]) AS id,
                    unnest($2::uuid[]) AS newest_id) r
      WHERE memories.id = r.id`,
    [rows.map((r) => r.id), rows.map((r) => r.newest_id)]
  );
  return rows.length;
}

To use this well, you need to be deliberate when your agent writes memories. Give tool outputs a key like tool:get_order_status:881 or weather:berlin. It costs nothing at write time, and it removes the single largest category of stale memory before any similarity math happens.

Step 3: Semantic Clustering and Deduplication

Not everything has a clean key. Conversational memories drift into near-duplicates: "The user wants a refund", "Customer is asking for money back", "User requested a refund for order 881". Same fact, three chunks, three slots in your top-k.

We will cluster these using nearest-neighbour search and a similarity threshold, then merge each cluster into one canonical memory.

Why not run a full clustering algorithm?

You could run HDBSCAN or k-means over every embedding. For a nightly batch job on a modest store, that is fine. But for a Node.js service that wants to run incrementally, a greedy neighbour-based approach with union-find is simpler, predictable, and easy to explain to your team. It also lets the database's HNSW index do the heavy lifting.

Visualization of embedding space with overlapping memory clusters collapsing into single canonical memories

A small union-find helper

// src/unionFind.ts
export class UnionFind {
  private parent = new Map<string, string>();

  find(x: string): string {
    if (!this.parent.has(x)) this.parent.set(x, x);
    let root = x;
    while (this.parent.get(root) !== root) root = this.parent.get(root)!;
    // Path compression
    let cur = x;
    while (cur !== root) {
      const next = this.parent.get(cur)!;
      this.parent.set(cur, root);
      cur = next;
    }
    return root;
  }

  union(a: string, b: string): void {
    const ra = this.find(a);
    const rb = this.find(b);
    if (ra !== rb) this.parent.set(rb, ra);
  }

  groups(): Map<string, string[]> {
    const out = new Map<string, string[]>();
    for (const id of this.parent.keys()) {
      const root = this.find(id);
      const list = out.get(root) ?? [];
      list.push(id);
      out.set(root, list);
    }
    return out;
  }
}

Finding neighbours with pgvector

pgvector's <=> operator returns cosine distance, so similarity is 1 - distance.

// src/neighbours.ts
import type { Pool } from "pg";

export interface Neighbour {
  id: string;
  similarity: number;
}

export async function findNeighbours(
  pool: Pool,
  agentId: string,
  memoryId: string,
  limit: number,
  minSimilarity: number
): Promise<Neighbour[]> {
  const sql = `
    SELECT b.id,
           1 - (b.embedding <=> a.embedding) AS similarity
      FROM memories a
      JOIN memories b
        ON b.agent_id = a.agent_id
       AND b.id <> a.id
       AND b.status = 'active'
       AND b.kind = a.kind
     WHERE a.id = $1
       AND a.agent_id = $2
     ORDER BY b.embedding <=> a.embedding
     LIMIT $3
  `;
  const { rows } = await pool.query(sql, [memoryId, agentId, limit]);
  return rows
    .map((r) => ({ id: r.id as string, similarity: Number(r.similarity) }))
    .filter((n) => n.similarity >= minSimilarity);
}

Notice we only compare memories of the same kind. A preference and a tool output can look semantically similar without being interchangeable, and merging across kinds is a classic way to lose information.

Choosing the canonical memory and merging

For each cluster, we keep the memory with the highest retention score and fold the others into it. Access counts are summed so the canonical record inherits the "this keeps being useful" signal.

// src/dedupe.ts
import type { Pool } from "pg";
import type { GcConfig, Memory } from "./types.js";
import { retentionScore } from "./decay.js";
import { UnionFind } from "./unionFind.js";
import { findNeighbours } from "./neighbours.js";

export interface MergePlan {
  canonicalId: string;
  duplicateIds: string[];
  mergedAccessCount: number;
  mergedImportance: number;
}

export async function planDedupe(
  pool: Pool,
  agentId: string,
  memories: Memory[],
  cfg: GcConfig
): Promise<MergePlan[]> {
  const uf = new UnionFind();
  const byId = new Map(memories.map((m) => [m.id, m]));

  for (const m of memories) {
    if (m.pinned) continue; // pinned memories are never merged away
    const neighbours = await findNeighbours(
      pool,
      agentId,
      m.id,
      cfg.neighbourLimit,
      cfg.similarityThreshold
    );
    for (const n of neighbours) {
      const other = byId.get(n.id);
      if (other && !other.pinned) uf.union(m.id, n.id);
    }
  }

  const plans: MergePlan[] = [];
  for (const ids of uf.groups().values()) {
    if (ids.length < 2) continue;

    const members = ids.map((id) => byId.get(id)!);
    members.sort((a, b) => retentionScore(b) - retentionScore(a));

    const [canonical, ...rest] = members;
    plans.push({
      canonicalId: canonical.id,
      duplicateIds: rest.map((r) => r.id),
      mergedAccessCount: members.reduce((sum, m) => sum + m.accessCount, 0),
      mergedImportance: Math.max(...members.map((m) => m.importance)),
    });
  }
  return plans;
}

export async function applyDedupe(
  pool: Pool,
  plans: MergePlan[]
): Promise<void> {
  const client = await pool.connect();
  try {
    await client.query("BEGIN");
    for (const p of plans) {
      await client.query(
        `UPDATE memories
            SET access_count = $2,
                importance   = $3,
                last_accessed_at = now()
          WHERE id = $1`,
        [p.canonicalId, p.mergedAccessCount, p.mergedImportance]
      );
      await client.query(
        `UPDATE memories
            SET status = 'tombstoned',
                tombstoned_at = now(),
                superseded_by = $1
          WHERE id = ANY($2::uuid[])`,
        [p.canonicalId, p.duplicateIds]
      );
    }
    await client.query("COMMIT");
  } catch (err) {
    await client.query("ROLLBACK");
    throw err;
  } finally {
    client.release();
  }
}

Two things to note. First, everything is wrapped in a transaction, so a crash halfway through never leaves a cluster half-merged. Second, duplicates are tombstoned, not deleted, and they point at their canonical memory through superseded_by. That gives you a full audit trail.

An optional upgrade: LLM-assisted merging

Picking the highest-scoring member is cheap and usually good enough. If your clusters contain complementary details, for example one chunk has the order number and another has the refund reason, you can pass the cluster to a small, inexpensive model and ask for a single consolidated memory. Re-embed the result and write it as the new canonical record. Just keep the originals tombstoned so you can verify the merge later.

Step 4: Safe Pruning

Now we use the retention score to sweep out low-value memories. This is the step where careless engineering causes real damage, so we add several guardrails.

// src/prune.ts
import type { Pool } from "pg";
import type { GcConfig, Memory } from "./types.js";
import { retentionScore } from "./decay.js";

export interface PrunePlan {
  candidateIds: string[];
  skippedByCap: number;
}

export function planPrune(
  memories: Memory[],
  cfg: GcConfig,
  now: Date = new Date()
): PrunePlan {
  const scored = memories
    .filter((m) => !m.pinned)
    .map((m) => ({ id: m.id, score: retentionScore(m, now) }))
    .filter((m) => m.score < cfg.pruneBelowScore)
    .sort((a, b) => a.score - b.score); // lowest first

  // Guardrail: never remove more than maxPruneRatio of the active store.
  const cap = Math.floor(memories.length * cfg.maxPruneRatio);
  const selected = scored.slice(0, cap);

  return {
    candidateIds: selected.map((s) => s.id),
    skippedByCap: Math.max(0, scored.length - selected.length),
  };
}

export async function applyPrune(pool: Pool, plan: PrunePlan): Promise<void> {
  if (plan.candidateIds.length === 0) return;
  await pool.query(
    `UPDATE memories
        SET status = 'tombstoned', tombstoned_at = now()
      WHERE id = ANY($1::uuid[]) AND pinned = FALSE`,
    [plan.candidateIds]
  );
}

export async function compact(pool: Pool, agentId: string, cfg: GcConfig) {
  const { rowCount } = await pool.query(
    `DELETE FROM memories
      WHERE agent_id = $1
        AND status = 'tombstoned'
        AND tombstoned_at < now() - ($2 || ' days')::interval`,
    [agentId, String(cfg.graceDays)]
  );
  return rowCount ?? 0;
}

That maxPruneRatio cap is a seatbelt. If a bug in your importance scoring suddenly marks everything as worthless, the GC will still only touch 20 percent of the store per run, and you will have time to notice.

Step 5: Orchestrating the Full GC Run

Now we wire the phases together into one function that returns a report. Reports matter. A GC you cannot observe is a GC you cannot trust.

// src/gc.ts
import type { Pool } from "pg";
import { DEFAULT_GC_CONFIG, type GcConfig, type Memory } from "./types.js";
import { supersedeByKey } from "./supersede.js";
import { planDedupe, applyDedupe } from "./dedupe.js";
import { planPrune, applyPrune, compact } from "./prune.js";

export interface GcReport {
  agentId: string;
  dryRun: boolean;
  activeBefore: number;
  superseded: number;
  mergedClusters: number;
  duplicatesRemoved: number;
  pruned: number;
  prunedSkippedByCap: number;
  compacted: number;
  durationMs: number;
}

async function loadActive(pool: Pool, agentId: string): Promise<Memory[]> {
  const { rows } = await pool.query(
    `SELECT id, agent_id, kind, supersession_key, content, importance,
            access_count, pinned, status, created_at, last_accessed_at
       FROM memories
      WHERE agent_id = $1 AND status = 'active'`,
    [agentId]
  );
  return rows.map((r) => ({
    id: r.id,
    agentId: r.agent_id,
    kind: r.kind,
    supersessionKey: r.supersession_key,
    content: r.content,
    embedding: [], // not needed in memory; neighbours are queried in SQL
    importance: r.importance,
    accessCount: r.access_count,
    pinned: r.pinned,
    status: r.status,
    createdAt: r.created_at,
    lastAccessedAt: r.last_accessed_at,
  }));
}

export async function runMemoryGc(
  pool: Pool,
  agentId: string,
  overrides: Partial<GcConfig> = {}
): Promise<GcReport> {
  const cfg = { ...DEFAULT_GC_CONFIG, ...overrides };
  const started = Date.now();

  // Phase 0: take an advisory lock so two GC runs never overlap.
  const lockKey = Buffer.from(agentId).reduce((a, b) => (a * 31 + b) | 0, 7);
  const { rows } = await pool.query("SELECT pg_try_advisory_lock($1) AS ok", [
    lockKey,
  ]);
  if (!rows[0].ok) throw new Error(`GC already running for ${agentId}`);

  try {
    const activeBefore = (await loadActive(pool, agentId)).length;

    // Phase 1: supersession by key (cheap, deterministic)
    const superseded = await supersedeByKey(pool, agentId, cfg.dryRun);

    // Phase 2: semantic dedupe on whatever is still active
    const afterSupersede = await loadActive(pool, agentId);
    const plans = await planDedupe(pool, agentId, afterSupersede, cfg);
    if (!cfg.dryRun) await applyDedupe(pool, plans);

    // Phase 3: decay-based prune
    const afterDedupe = cfg.dryRun ? afterSupersede : await loadActive(pool, agentId);
    const prunePlan = planPrune(afterDedupe, cfg);
    if (!cfg.dryRun) await applyPrune(pool, prunePlan);

    // Phase 4: permanently remove old tombstones
    const compacted = cfg.dryRun ? 0 : await compact(pool, agentId, cfg);

    return {
      agentId,
      dryRun: cfg.dryRun,
      activeBefore,
      superseded,
      mergedClusters: plans.length,
      duplicatesRemoved: plans.reduce((n, p) => n + p.duplicateIds.length, 0),
      pruned: prunePlan.candidateIds.length,
      prunedSkippedByCap: prunePlan.skippedByCap,
      compacted,
      durationMs: Date.now() - started,
    };
  } finally {
    await pool.query("SELECT pg_advisory_unlock($1)", [lockKey]);
  }
}

Scheduling it

For a single service, node-cron is plenty. For multi-instance deployments, the advisory lock above prevents duplicate runs, or you can move the job to your existing queue or scheduler.

// src/scheduler.ts
import cron from "node-cron";
import { Pool } from "pg";
import { runMemoryGc } from "./gc.js";

const pool = new Pool({ connectionString: process.env.DATABASE_URL });

// Nightly at 03:15, writes enabled only after you have validated dry runs.
cron.schedule("15 3 * * *", async () => {
  const agents = await pool.query("SELECT DISTINCT agent_id FROM memories");
  for (const { agent_id } of agents.rows) {
    try {
      const report = await runMemoryGc(pool, agent_id, {
        dryRun: process.env.GC_ENABLE_WRITES !== "true",
      });
      console.info("[memory-gc]", JSON.stringify(report));
    } catch (err) {
      console.error("[memory-gc] failed", agent_id, err);
    }
  }
});

Dashboard mockup showing GC run reports: memories before and after, clusters merged, tombstones pending compaction

Step 6: Apply Decay at Retrieval Time Too

Here is a point many teams miss. The GC cleans the store periodically, but between runs, stale memories are still sitting there waiting to be retrieved. The fix is to apply the same decay logic to ranking.

Instead of sorting purely by vector similarity, fetch a larger candidate set, rescore it, and then pick the final top-k. Bumping access_count on the memories you actually use closes the reinforcement loop.

// src/retrieve.ts
import type { Pool } from "pg";
import type { Memory } from "./types.js";
import { retentionScore } from "./decay.js";

export async function retrieve(
  pool: Pool,
  agentId: string,
  queryEmbedding: number[],
  k = 6
) {
  const { rows } = await pool.query(
    `SELECT id, kind, content, importance, access_count, pinned,
            created_at, last_accessed_at,
            1 - (embedding <=> $2::vector) AS similarity
       FROM memories
      WHERE agent_id = $1 AND status = 'active'
      ORDER BY embedding <=> $2::vector
      LIMIT $3`,
    [agentId, JSON.stringify(queryEmbedding), k * 4]
  );

  const reranked = rows
    .map((r) => {
      const m = {
        kind: r.kind,
        importance: r.importance,
        accessCount: r.access_count,
        pinned: r.pinned,
        createdAt: r.created_at,
        lastAccessedAt: r.last_accessed_at,
      } as Memory;
      // Blend semantic match with retention. Tune the weights on your own evals.
      const blended = 0.7 * Number(r.similarity) + 0.3 * retentionScore(m);
      return { id: r.id as string, content: r.content as string, blended };
    })
    .sort((a, b) => b.blended - a.blended)
    .slice(0, k);

  await pool.query(
    `UPDATE memories
        SET access_count = access_count + 1, last_accessed_at = now()
      WHERE id = ANY($1::uuid[])`,
    [reranked.map((r) => r.id)]
  );

  return reranked;
}

The 70/30 blend is a starting point, not a law. Build a small evaluation set of real agent queries with known correct memories, and tune the weights against it.

Best Practices

After running this pattern in a few production agents, these habits have paid off the most:

  • Always start in dry-run mode. Log what would have been removed for a week and read through a sample by hand.
  • Write memories with intent. Assign a kind, an importance, and a supersession_key at write time. Cleaning up is much harder than labeling correctly in the first place.
  • Separate memory tiers. Keep short-lived working memory, episodic memory, and durable facts or preferences apart. They deserve different half-lives and different rules.
  • Pin what must survive. Compliance notes, user-stated hard constraints, and safety instructions should never be eligible for collection.
  • Cap every run. A maximum prune ratio converts a catastrophic bug into a minor annoyance.
  • Keep tombstones restorable. A 7 to 14 day grace period has saved us from more than one bad threshold.
  • Log structured reports. Track active count, merge count, prune count, and the distribution of retention scores over time.
  • Scope by tenant. In multi-user systems, every query, cluster, and merge must be filtered by agent and user. A cross-tenant merge is a privacy incident.
  • Test with a golden set. Keep a list of memories that must still be retrievable after a GC run, and assert on it in CI.

Common Mistakes to Avoid

We have made most of these ourselves, so consider this a friendly warning.

1. Using one global TTL

A single expiry for everything is simple and wrong. It deletes your user's long-term preferences at the same time as yesterday's search results.

2. Setting the similarity threshold too low

At 0.80, "User wants to cancel their subscription" and "User asked how to pause their subscription" may merge into one. Those are very different intents. Start high, around 0.93, and loosen only with evidence.

3. Merging across kinds or tenants

Always filter neighbours by kind and by owner. We showed it in the SQL above, and it is worth repeating.

4. Deleting immediately

Hard deletes on the first pass leave you no way to recover from a bad score. Tombstone first, compact later.

5. Ignoring the feedback loop

If you only clean up the store but keep writing raw tool outputs back verbatim, you are mopping the floor with the tap running. Summarize or key tool outputs before they are stored.

6. Trusting LLM-assigned importance blindly

Model-assigned importance scores are noisy and often cluster around 0.7 to 0.8. Combine them with simple heuristics like memory kind and explicit user statements, and spot-check the distribution.

7. Never measuring the effect

If you cannot show that retrieval precision improved or token usage dropped after the GC, you are guessing. Track retrieval hit quality and average prompt size before and after.

🚀 Pro Tips

  • Embed once, reuse everywhere. Store the embedding model name and version in metadata. When you upgrade models, similarity thresholds change, so re-tune before re-enabling aggressive dedupe.
  • Use a "contradiction check" for facts. When a new fact is highly similar to an older one but the content differs, for example plan names or dates, flag the older one as superseded instead of merging.
  • Add a cheap recall probe. After each GC run, replay a handful of known queries and alert if any expected memory disappears from the top results.
  • Compress before you collect. For conversation memories that are borderline, summarize them into a durable fact rather than deleting them outright.
  • Make half-lives configurable per agent. A coding agent and a customer support agent age information very differently.
  • Expose a "why was this removed" endpoint. The superseded_by link and tombstone timestamps make debugging a two-minute job instead of a two-hour one.
  • Watch the long tail. If a large share of your store sits just above the prune threshold, your importance scores are probably too generous.

📌 Key Takeaways

  • Long-running agents fill their vector stores with stale tool outputs and near-duplicate chunks, causing context distraction and hallucination cascades.
  • A memory GC is built from three mechanisms: time-weighted relevance decay, semantic clustering and deduplication, and safe, reversible pruning.
  • Supersession keys are the cheapest way to eliminate stale facts, and they require no embeddings at all.
  • Use kind-specific half-lives so short-lived tool outputs fade fast while user preferences persist.
  • Guardrails make the system trustworthy: dry-run defaults, pinned memories, a per-run prune cap, tombstones with a grace period, advisory locks, and transactional merges.
  • Apply decay at retrieval time as well as during cleanup, so stale memories lose ranking between GC runs.
  • Measure everything. Report counts, track retrieval quality, and replay known queries after each run.

Conclusion

Agent memory is often sold as an append-only feature: store everything, retrieve what is relevant, and let the model sort it out. In practice, that approach works for a demo and quietly fails in production. The longer your agent runs, the more its past starts to interfere with its present.

A memory garbage collector gives your agent something humans rely on every day, which is the ability to forget well. Decay lets old information fade, deduplication keeps repeated facts from crowding the prompt, and careful, reversible pruning keeps the whole system healthy without risking what matters.

The code in this guide is intentionally small, and you can run it against a staging database this afternoon. Start in dry-run mode, read the reports, tune your thresholds on real data, and only then let it write. Your agent, and the customers on the other end of it, will notice the difference.

References

Frequently asked questions

What is context poisoning in an AI agent's vector store?

Context poisoning happens when outdated, contradictory, or redundant memories are retrieved into the prompt and mislead the model. The agent treats an old tool output or a superseded fact as current truth, and its next action builds on that mistake. Over time the errors compound into a hallucination cascade.

How is a memory garbage collector different from a simple TTL?

A TTL deletes everything after a fixed age, regardless of value. A memory GC scores each memory using age, importance, access frequency, and kind, then merges duplicates and prunes only what is genuinely low value. Important old memories survive, and noisy recent ones can still be removed.

What cosine similarity threshold should I use for deduplication?

For most modern embedding models, start around 0.92 to 0.95 and tune against your own data. Lower thresholds merge more aggressively but risk collapsing distinct facts. Always run in dry-run mode first and inspect a sample of the proposed clusters before enabling writes.

Can I run this GC with Pinecone, Qdrant, or Chroma instead of pgvector?

Yes. The GC logic lives behind a small VectorStore interface. You only need to implement listing, nearest-neighbour lookup, metadata updates, and delete or tombstone operations for your database. The decay math and clustering code stay the same.

How often should the memory GC run?

Most teams run a light pass hourly or nightly and a deeper clustering pass weekly. Match the cadence to your write volume. An agent that stores thousands of memories a day needs more frequent cleanup than one that stores a few dozen.

Will pruning memories make my agent forget important things?

Not if you design for it. Pin critical memories, give durable kinds like user preferences very long half-lives, cap how much any single run can remove, and keep tombstoned records restorable for a grace period before permanent deletion.

Discussion

All Articles
AI AgentsNode.jsVector DatabaseAgent MemorypgvectorContext EngineeringRAGTypeScript

Written by

Niraj Kumar

Software Developer — building scalable systems for businesses.

Building this for real? LangChain Developer Services — Production LangChain & LangGraph systems for RAG, agents, and AI workflows.