Why does my vector DB retrieval keep returning duplicate noise for my chat agent, and how to fix it?
TL;DR
Duplicate results come back for three reasons, and each has its own setting.
- Look-alikes, by design. A similarity search returns the items closest to your query, and a near-copy of a good match sits almost as close. Setting: maximal marginal relevance.
- One document's chunks. A document split into overlapping chunks can fill the list with its own pieces. Setting: group or collapse by a document field.
- One fact stored more than once. A re-ingest that hands out fresh IDs, or a chat agent writing back what it already read, leaves copies. Setting: stable IDs on the write, and lifecycle rules for an old and a new version of one fact.
A higher similarity threshold fixes none of the three.
Disclosure: we build Mnemoverse, one of the memory layers mentioned here. Every fact on this page, ours included, comes from a page or source file published by its owner and read on 27 September 2026, and our own row was checked hardest.
The same map is on film, Why does my vector DB retrieval keep returning duplicate noise for my chat agent, and how to fix it?: sixteen minutes in seven chapters, for anyone who would rather watch the settings than read them. The text below stands on its own, and every source the film opens on screen is linked from it.
Why a higher similarity threshold does not remove duplicate results
The first fix everyone reaches for is a higher similarity threshold. It cannot work, and the reason is structural.
A similarity threshold is a floor against the query: it keeps or drops each result on that result's own score for the question, and it never compares one result with another. Every threshold read for this page works that way:
| Store or layer | Setting | Page |
|---|---|---|
| Qdrant | score_threshold | Search |
| OpenAI vector store | score_threshold | Retrieval guide |
| LangChain | search_type="similarity_score_threshold" | as_retriever |
| LlamaIndex | SimilarityPostprocessor | Node postprocessors |
| Mem0 | threshold | Search Memories API |
| Supermemory | threshold | Search |
Our own min_relevance on a Mnemoverse read is the same kind of setting, and it is described with the rest of our row at the end of this page.
Suppose two chunks express the same fact in slightly different wording. Both are close to the question, so both receive almost the same query score, and raising the floor tends to keep both or drop both. In the toy script at the end of this page, "The user prefers dark mode." and "User likes the dark theme." score 0.89 and 0.90 against the same query: a floor that separated them would have to land in the hundredth between those two numbers, which is tuning to one pair, not a rule. A threshold cuts weak matches, not copies of each other. That consequence is our reasoning from the definitions above, and we attribute it to none of these vendors.
That does not make thresholds useless. They remain the right tool for weak or irrelevant matches: Mem0 calls its threshold a minimum relevance score, Supermemory describes its threshold as a similarity cutoff against the query where a higher value returns fewer results, and Supermemory's advice for coding agents, "Raise the similarity threshold if irrelevant memories keep appearing", is about irrelevant memories, not copies. Duplicate removal asks a different question of each result: not whether it is close enough to the query, but whether it is too close to something already picked. Nothing in a threshold asks that.
Three symptoms, three settings
Name the symptom before the fix. Qdrant's beginner course pairs each problem with the feature that addresses it, and the first two rows below are its pairing. The third row is ours.
| Symptom | Setting |
|---|---|
| The top results are near-identical | Maximal marginal relevance (MMR): pick results relevant to the query and different from those already picked |
| One document fills the list with its own chunks | Group or collapse by a document field, so one source takes one slot |
| An old and a new version of one fact both come back | Lifecycle, not similarity: a similarity score cannot say which version is current. See Agent Memory Consolidation Compared |
These symptoms can look identical in a prompt. Their causes are different, and so are their settings. Copies made by loading the same content twice are stopped earlier, on the write, with stable IDs (below).
This page is about what happens when memories or chunks are read: why a similarity search hands a chat agent near-identical results, and which read-time settings take them out. What never gets stored is the subject of Agent Memory Write Gate. What happens to the older of two contradicting records is covered on Agent Memory Consolidation Compared. How a merge is decided at write or index time is on Agent Memory Deduplication.
Why similarity search returns look-alikes
A top-k similarity search returns the K items closest to the query, ranked by similarity score. Qdrant's course puts it this way: "Qdrant finds the K points in the collection whose vectors are most similar to the query vector, ranked by similarity score." LangChain says similarity search "returns the closest embedded documents". Pinecone describes semantic search as looking for "records that are most similar in meaning and context to a query". Mem0's search page: "Embeddings locate the closest memories using cosine similarity across your scoped dataset." Mem0's full read path fuses more than that one signal, and its source is read on RAG vs Agent Memory; the sentence quoted here is the first stage, and it is the stage every store shares.
None of these definitions asks whether two results say the same thing. Near-copies are close to the same query for the same reason they are close to each other, so plain top-k ranking has no reason to reserve separate slots for different evidence. Qdrant calls the result an "echo chamber of similar results". Weaviate's concepts page: "Standard vector search returns the closest matches to a query, which often means a cluster of near-duplicate results."
Hybrid search does not remove this condition by itself. Weaviate says fusing a keyword result set with a vector result set "often means the top of the fused list is a cluster of near-duplicates". Merging two lists looks for the same entry, not the same meaning: LangChain's EnsembleRetriever decides which documents are the same by page_content, or by a metadata key you name in id_key, and Pinecone's hybrid guide has you link dense and sparse results by a shared ID field "so you can merge and deduplicate search results later". All three merge by identity, not by similarity between results.
The cost is a slot. Qdrant's clean-up post makes the point for agents: an agent that reads only the first few results has only those slots as evidence, so every duplicate spends one. The problem is also old. In 1998 Carbonell and Goldstein wrote that when relevant documents are highly redundant, or even duplicates, ranking by relevance alone is not enough. Letta's word for what a raw retrieval does to the window is context pollution, and Letta uses it for irrelevant data in general, not for duplicates: "RAG often places irrelevant data into the context window, resulting in context pollution".
Where the copies enter the system
Overlap, chosen on purpose. It preserves information that crosses a chunk boundary. LangChain's base text splitter defaults to an overlap of 200 characters; LlamaIndex documents a default chunk size of 1024 with an overlap of 20, both in tokens in the source. The price is written down too: Qdrant's course lists sliding windows with "duplicate content across results", and Chroma's chunking guide says of overlap: "The downside: you’re storing and embedding duplicate content." Separately, a chunking guide on Supermemory's blog, written as general RAG advice and not about Supermemory's own pipeline, says overlap can "return near-duplicates that push other useful evidence out of the result set".
Loading the same content again. Qdrant: "Duplicates accumulate when each ingest assigns fresh IDs to content the collection already holds." Zep: "Reingesting the same file creates new episodes, even when its contents are unchanged."
The chat agent itself. An agent that writes its own turns back into memory can store what it has just read. Mem0's plugin changelog records a fix: "Repeated and concurrent response hooks no longer duplicate an answer, while identical answers after separate prompts are preserved." Supermemory's changelog, for its Codex plugin: "Codex capture no longer stores injected context or ingests the same fallback turn twice." Zep's AG2 guide warns that when two agents in one conversation each attach a manager pointing at the same session, "every turn is persisted twice"; the fix there is in the wiring. These are recorded fixes and integration constraints, and they show why duplicate retrieval is not always a ranking problem in the database.
Maximal marginal relevance in vector search, store by store
Maximal marginal relevance (MMR) picks results that are relevant to the query and different from those already picked. Carbonell and Goldstein introduced it in 1998 for result sets where relevant documents were highly redundant, and defined it in one sentence: "a document has high marginal relevance if it is both relevant to the query and contains minimal similarity to previously selected documents" (SIGIR 1998, authors' copy).
Each candidate is scored against the question and against the picks already made, which is the comparison a threshold on the query never makes. So MMR is a selection from a pool, not only a new order, and it needs more candidates than it returns. Qdrant documents its own MMR that way: "MMR selects candidates iteratively, starting with the most relevant point". Zep's search guide names the symptom in the sentence that introduces its fix: "Maximal Marginal Relevance addresses a common issue in similarity searches: highly similar top results that don’t add diverse information to your context."
| Store or library | Documented setting | Candidate-pool detail |
|---|---|---|
| Qdrant | Mmr(diversity, candidates_limit) inside a nearest query | candidates_limit falls back to the query limit, which leaves no spare candidates unless you raise it |
| Weaviate | Diversity.mmr(limit, balance) in diversity_selection | the top-level query limit supplies the candidates; the MMR limit is how many come back |
| Weaviate Query Agent | diversity_weight | higher values favour more varied results over the most relevant ones |
| LangChain | search_type="mmr" with k, fetch_k, lambda_mult | defaults k=4, fetch_k=20, lambda_mult=0.5 |
| LlamaIndex | vector_store_query_mode="mmr" with mmr_threshold | despite its name, mmr_threshold weighs relevance against diversity |
| Elasticsearch | the diversify retriever with type: "mmr", lambda, size and rank_window_size | size and rank_window_size default to 10; the page disagrees with itself about the direction of lambda, so none is stated here |
| Zep | reranker="mmr" with mmr_lambda | opt-in per call; Zep calls 0.5 balanced |
| Graphiti | MMR search recipes in the library behind Zep | in the source, each search fetches twice the limit, a candidate below the minimum MMR score is dropped, and the result is cut to the limit |
| Pinecone through LangChain | PineconeVectorStore.as_retriever(search_type="mmr"), on the class page | LangChain's defaults, k=4, fetch_k=20, lambda_mult=0.5 |
| Milvus through LangChain | vector_store.as_retriever(search_type="mmr") | Milvus's own LangChain page shows the retriever switch |
Sources: Qdrant MMR, Weaviate vector search, Weaviate Query Agent, LangChain MMR, LlamaIndex MMR, Elasticsearch diversify retriever, Zep MMR, Graphiti source, LangChain Pinecone, Milvus LangChain integration.
The trade-off knobs do not all run in the same direction, so check the direction before you tune one:
| Setting | Toward relevance | Toward diversity |
|---|---|---|
Qdrant diversity | 0.0 | 1.0 |
Weaviate balance | 1.0 | 0.0 |
Weaviate Query Agent diversity_weight | lower | higher |
LangChain lambda_mult | 1 | 0 |
LlamaIndex mmr_threshold | close to 1 | close to 0 |
Zep mmr_lambda | 1.0 | 0.0 |
One integration needs care. LangChain defines a higher lambda_mult as less diversity, but on Qdrant through langchain-qdrant the value reaches Qdrant as its diversity, where higher means more diversity. An open LangChain bug report records the inversion. At 0.5 the two readings coincide, so the trap opens only when you move it.
If your store sits behind LangChain, the switch is one argument on the retriever, and LangChain's PGVector page, its PineconeVectorStore reference and Milvus's own LangChain page show the same call. With the reference's defaults written out:
retriever = vector_store.as_retriever(
search_type="mmr", # the default is "similarity"
search_kwargs={"k": 4, "fetch_k": 20, "lambda_mult": 0.5},
)k is how many documents come back, fetch_k is how many candidates MMR picks them from, and lambda_mult runs from 0, maximum diversity, to 1, minimum diversity. Ask for more candidates than you keep: MMR can only choose among spares. Qdrant's course names the trap in its own MMR, where candidates_limit defaults to the query's limit,
which leaves MMR nothing spare to choose from, so all it can do is reorder the results it was already given. This is the most common reason MMR looks like it did nothing.
Qdrant's own example keeps 10 from 100 candidates, Weaviate's keeps 5 from 20, and LangChain's defaults keep 4 from 20. Weaviate adds a paging warning: with diversity selection the offset must advance by the query's top-level limit, not by the page size, and nothing raises an error when you get it wrong. "The only symptoms are objects that repeat across pages and objects that are never returned at all."
Group vector search results by document
Grouping by a document field collapses the chunks that share a document identifier, so one source takes one slot or one small group. The results need not be near-identical. They come from the same source.
Qdrant describes the "neighbor problem": one long document becomes many points, and a strong match on it can fill the whole first page with its own chunks. Milvus describes the same pattern: over chunked documents, "the search results may include several paragraphs from the same document, potentially causing other documents to be overlooked".
Flat top-4 retrieval: Grouped by document_id, one per group:
1. doc_A, chunk 12 1. doc_A, chunk 12
2. doc_A, chunk 13 2. doc_B, chunk 01
3. doc_A, chunk 11 3. doc_C, chunk 05
4. doc_B, chunk 01 4. doc_D, chunk 02Do it in the store. Qdrant's course says why the server should group rather than your code: "Deduplicating the results yourself after the search does not fill the page." Ask for ten, receive ten chunks of one document, drop nine of them yourself, and the model sees one result where you budgeted for ten. Grouping in the store collects distinct sources until the limit is reached. A collapse that runs after the fetch cannot: OpenSearch says its collapse response processor "will likely result in fewer than size results being returned", and points to the fix: "To increase the likelihood of returning size hits, use the oversample request processor and truncate_hits response processor". That is the same over-fetch MMR needs.
| Store or library | Setting | What comes back |
|---|---|---|
| Qdrant | Grouping API, group_by="document_id" with limit and group_size | limit groups of up to group_size points each, by a keyword or integer payload field |
| Weaviate | GroupBy(prop=..., objects_per_group=..., number_of_groups=...) on a near search | results grouped by a property or cross-reference |
| Milvus | group_by_field="docId" | one entity per group by default; group_size raises it |
| Chroma Cloud | GroupBy in the Search API | the Search API overview says it is available in Chroma Cloud only, with single-node support planned |
| Elasticsearch | collapse on a field | the top sorted document per collapse key |
| OpenSearch | collapse, or a collapse processor in a search pipeline | the top document within each group |
| PostgreSQL | SELECT DISTINCT ON (document_id) | the first row of each set; unpredictable unless ORDER BY puts the one you want first |
| LangChain | MultiVectorRetriever | searches child chunks, then works by "Collecting unique parent document IDs from chunk metadata" and returns each parent once |
| LlamaIndex | AutoMergingRetriever over a hierarchical index | can "automatically replace retrieved nodes with their parents when a majority of children are retrieved" |
Sources: Qdrant grouping and its course page, Weaviate GroupBy, Milvus grouping, Chroma Cloud GroupBy, Elasticsearch collapse, OpenSearch collapse and collapse processor, PostgreSQL DISTINCT ON, LangChain MultiVectorRetriever, LlamaIndex AutoMergingRetriever.
One fix sits upstream. Chroma's chunking guide ends with a troubleshooting list, and its line for duplicate results is one instruction: decrease chunk overlap.
Stable IDs prevent duplicate writes from becoming duplicate reads
Read-time selection cannot compensate indefinitely for repeated writes. A re-upload with stable IDs does not add a copy, and Qdrant's clean-up post starts there: "Stable point IDs prevent most duplication at the source."
- Qdrant: "points with the same id will be overwritten when re-uploaded". If you do not pass IDs, the client generates random UUIDs.
- Chroma: an add whose ID already exists "will be ignored without throwing an error", and
collection.upsert(update data) updates the record instead. - Pinecone: "If a record ID already exists, upserting overwrites the entire record."
- pgvector: the README upserts on the ID with
ON CONFLICT (id) DO UPDATE SET embedding = EXCLUDED.embedding. - Weaviate: "Use deterministic IDs to avoid inserting duplicate objects", with
generate_uuid5(data_object). - Milvus: an insert with an existing primary key "creates a new entity with the same key"; its page says to use upsert.
- LangChain's indexing API with a record manager hashes each document so a re-run knows which to skip, and LlamaIndex's checklist says to "Use IngestionPipeline with a docstore to deduplicate documents before indexing".
Stable IDs catch the same content loaded twice. They do not decide which of two contradicting facts is current: that is the third row of the map, and the lifecycle settings memory layers offer for it are compared on Agent Memory Consolidation Compared. What a write gate may refuse before anything is stored is the subject of Agent Memory Write Gate.
Memory layers on the read path
A chat agent often reads from a memory layer rather than straight from a store. This section keeps to what each layer does on the read, from its docs or, where marked, its source, and leaves their write paths to the pages linked above.
Mem0. On the hosted Platform, Dream's Merge keeps merged duplicates out of reads: "The merged record is hidden from reads by default (you get the one canonical memory), retained rather than deleted". The lifecycle half of the same page, Supersede and the latest_only read option, is compared with the other layers on Agent Memory Consolidation Compared. In a chat loop, its DeepSeek Harness plugin page says recall "adds unseen results to the model context", and in the source of its coding-agent plugins a memory search returns only memories not already injected into that conversation.
Supermemory. Its changelog of 6 February 2026: "Search results now show one copy of identical content, keeping the highest-scoring match." Its Cursor plugin page says recall "deduplicates results" before injecting them. In the source of its Claude Code plugin, a memory injected once in a session is not injected again, and in the source of its OpenAI middleware for Python, retrieved context is marked "so the next turn can replace it": one memory block, not one per turn.
Zep. Its hosted graph search ships the MMR reranker in the table above.
Cognee. Its docs say that in HYBRID_COMPLETION, which its recall page calls the default, "The lexical and vector chunks are merged and de-duplicated up to chunks_top_k". In the source, a chunk found twice is recognised by its chunk id: that removes the same chunk found through two lanes, not two different chunks with similar text.
Letta. Its documented answer to top-K noise is a change of shape. Its blog: "With agentic RAG, an AI agent isn’t doing a top-K match and dump." Its subagents page moves the search into a subagent, so the main agent sees the final answer instead of everything the search read.
In the source of Mem0's plugin core and of Supermemory's Claude Code plugin, the guard matches the same memory or the same text, so a reworded copy gets through. These controls should not be collapsed into one claim: identity deduplication, MMR, lifecycle filtering, parent grouping and prompt-block replacement solve different problems, and a layer that documents one of them has not thereby documented the others.
How to prevent RAG pipelines from filling context windows with low-quality similar embeddings?
The question holds two problems, low-quality hits and similar ones, and each has its own setting. Both sit on top of a bound on how much reaches the window.
Bound the window. Among the calls read for this page, most bounds are counts. Each default below is the one its own documentation or source states.
| Where | Count setting | Default |
|---|---|---|
| LangChain retriever | k | 4, and MMR fetches fetch_k 20 candidates |
| LlamaIndex retriever | similarity_top_k | 2 |
| LangGraph store search | limit | 10 |
| Chroma query | n_results | 10 |
| OpenAI vector store search | max_num_results | 10, up to 50 |
| Pinecone Assistant context | top_k | 16 |
| Mem0 Platform search | top_k | 10 |
| Mem0 self-hosted library search | top_k | 20 |
| Supermemory search | limit | 10 |
| Zep graph search | limit | 10 |
| Cognee recall and search | top_k | 15 |
Where a vendor's own pages disagree, this page states no default: Mem0's threshold (0.1 on the API reference, 0.3 on the CLI page) and Supermemory's threshold (0.5 on its search guide against 0.6 on its API reference). A number copied from one page would be contradicted by the next. Two more are left out for a different reason: Letta's search reference says top_k "Uses system default if not specified" and states no number, and Zep's SDK reference caps limit at 50 while its search guide's table does not mention the cap.
A few budget in characters or tokens instead:
| Where | Size setting | What the page states |
|---|---|---|
| Zep auto search | max_characters | 2,500 by default, capped at 50,000; the context block is packed to fit it |
| LlamaIndex Memory | token_limit | 30,000 tokens by default in Memory.from_defaults(), over short-term plus long-term content |
| Pinecone Assistant | snippet_size | "default is 2048 tokens" |
| Letta, legacy V1 SDK tool creation | return_char_limit | "The maximum number of characters in the response."; no default stated |
| Mem0 Claude Code plugin | max_context_chars | the injected context is capped at 4,000 characters by default (Claude Code integration) |
Sources: Zep auto search and SDK reference, LlamaIndex Memory, Pinecone Assistant, Letta tools.create, Mem0 search API, Supermemory search API, Cognee recall, Chroma querying, OpenAI retrieval, LangGraph stores.
Cut low-quality hits with a reranker over a wider pool. A reranker cuts only when it runs over more results than it keeps. LangChain's cross encoder page: "Retrieve a relatively large k; the reranker will narrow it down." Qdrant's course: "A common approach is to overfetch: retrieve a larger candidate pool, then rerank it down to the smaller K you actually show the user." The setting is that pair, retrieve more and keep fewer. Each of these scores a result against the query and never compares two results with each other.
Take out similar hits with a step that compares results with each other: MMR, above, or the step in the next section.
In a chat loop, send each memory once. Zep's placement guidance for its Context Block points the same way, for its own reason, to preserve the cacheable prompt prefix: "Replace the previous turn’s block instead of appending a second one." The Mem0 and Supermemory plugin guards above do the same for single memories.
A step you can add yourself: compare results with each other
Where the store or layer you call offers no diversity setting, or you want one step that works over any list, compare the returned results with each other and drop one of each close pair. That takes about ten lines of your own code. Here is a version we wrote for this page, standard library only. The vectors are toy values: the script shows the mechanism and makes no claim about results on real data.
"""Drop near-copies from a list of retrieved results, standard library only.
A threshold on the query cannot do this: it compares each result with the question, and two
copies of one fact score almost the same against it. This compares the results with each other.
"""
import math
def cosine(a, b):
dot = sum(x * y for x, y in zip(a, b))
return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))
def drop_near_copies(results, too_close=0.95):
"""results: (text, embedding) pairs, best match first. Keeps the first of each close pair."""
kept = []
for text, emb in results:
if all(cosine(emb, other) < too_close for _, other in kept):
kept.append((text, emb))
return kept
if __name__ == "__main__":
# Toy embeddings: the first two are the same fact in slightly different words.
results = [
("The user prefers dark mode.", [0.90, 0.10, 0.00]),
("User likes the dark theme.", [0.89, 0.12, 0.01]),
("The user works in Lisbon.", [0.10, 0.90, 0.20]),
("Deploys run on Fridays.", [0.00, 0.20, 0.95]),
]
query = [0.80, 0.40, 0.30]
floor = 0.5
passed = [t for t, e in results if cosine(query, e) >= floor]
print("floor against the query keeps:", len(passed), passed)
kept = drop_near_copies(results)
print("compared with each other keeps:", len(kept), [t for t, _ in kept])Running it on 1 October 2026 with Python 3.12 prints:
floor against the query keeps: 3 ['The user prefers dark mode.', 'User likes the dark theme.', 'The user works in Lisbon.']
compared with each other keeps: 3 ['The user prefers dark mode.', 'The user works in Lisbon.', 'Deploys run on Fridays.']The floor against the query keeps both copies of the dark-mode fact and drops the Friday one. Comparing the results with each other keeps one copy of the dark-mode fact and all three distinct facts. One limit is built in: the step keeps the first of each close pair, and a correction scores close to the fact it replaces for the same reason a copy does, so it cannot tell the two apart. Versions of one fact are the third row of the map, not this step.
The library version is LangChain's EmbeddingsRedundantFilter: "Filter that drops redundant documents by comparing their embeddings." In its source the default similarity_threshold is 0.95. It runs as a document transformer or inside a DocumentCompressorPipeline, and it embeds the documents it is handed, so a memory layer's list can go through it wrapped as documents. It ships in langchain-community, and LangChain's own integration pages carry this notice: "The langchain-community package is no longer maintained."
Where Mnemoverse sits
We build Mnemoverse, a memory layer for AI agents. Its row, from its own API reference and Claude page:
- A REST read has a relevance floor,
min_relevance, 0.3 by default, and in the source the Python client defaults to the same 0.3. Each result has to clear it on its own relevance score, and nothing in it compares two results with each other, so it is a floor like the others above. - A correction names the memory it replaces with
supersedeson the write. - The API reference documents an organisation setting, off by default, that hides replaced versions from reads, and a REST read can pass
include_historyto bring replaced versions back. How that compares with the other layers' lifecycle settings is on Agent Memory Consolidation Compared. - A write can carry a client reference,
external_ref, which makes it idempotent: reuse one and the API answers409instead of storing the write again. - In the MCP server package the default count is five, and the reference calls it a request rather than a hard cap: association expansion can return more, and the relevance floor can return fewer.
What this page does not claim
No performance, latency, accuracy or benchmark number, for anyone: nothing here says one setting gives better results than another in measured terms. No count of how many stores or vendors offer MMR, grouping or a threshold; each is named with its own setting. No default where a vendor's own pages disagree, and the two cases are named in the window section. Where a sentence rests on a vendor's source code rather than its documentation, it says "in the source". The code is ours and runs on toy vectors.
Common questions
Why does my vector DB retrieval keep returning duplicate noise for my chat agent, and how to fix it?
Duplicates come back for three reasons, and each has its own setting. A similarity search returns the items closest to the query, and near-copies of a good match are close too: turn on maximal marginal relevance, which picks results relevant to the query and different from those already picked. One document split into overlapping chunks fills the list: group or collapse results by a document field. One fact stored more than once, by a re-ingest or by the agent itself: give re-ingested content stable IDs, wire the agent's capture so a turn is stored once, and handle old and new versions of a fact with lifecycle settings. A higher similarity threshold fixes none of the three.
How to prevent RAG pipelines from filling context windows with low-quality similar embeddings?
Treat low quality and similarity as two problems. For low-quality hits, run a reranker over a wider pool and keep fewer: retrieve a relatively large k, then narrow it down. For similar hits, add a step that compares results with each other, such as maximal marginal relevance or a filter that drops one of each close pair. Then bound what reaches the window: among the calls this page reads, most bounds are counts, and a few budget in characters or tokens, such as Zep's max_characters and LlamaIndex Memory's token_limit. In a chat loop, send each memory once: replace the previous turn's memory block instead of appending a second one.
Why does a higher similarity threshold not remove duplicate results?
A similarity threshold on a search is a floor against the query. It keeps or drops each result on its own score for the question and never compares two results with each other. Two copies of one fact score almost the same against any query, so they tend to pass together or fail together. A threshold cuts weak matches, not copies of each other.
What is maximal marginal relevance in vector search?
Maximal marginal relevance, defined by Carbonell and Goldstein in 1998, scores each candidate for relevance to the query and for similarity to the results already picked, and picks one at a time. It is a selection from a pool, not only a new order, so it needs spare candidates: ask for more than you keep. Qdrant, Weaviate, LangChain, LlamaIndex, Elasticsearch and Zep each document a setting for it.
How do I stop one document's chunks from filling the search results?
Group or collapse by a document field, so one source takes one slot. Qdrant has a Grouping API, Weaviate a GroupBy, Milvus a group_by_field argument, Chroma Cloud a GroupBy in its Search API, and Elasticsearch and OpenSearch collapse results on a field. Do it on the server: deduplicating the results yourself after the search does not fill the page. Upstream, Chroma's chunking guide answers duplicate results with less chunk overlap.
How do I stop a chat agent from injecting the same memory every turn?
Remember what the conversation already holds, and skip it. Mem0's DeepSeek Harness plugin page says recall adds only unseen results to the model context, and Supermemory's Cursor plugin page says recall deduplicates results before injecting them. In the source of Mem0's plugin core and of Supermemory's Claude Code plugin, the guard matches the same memory or the same text, so a reworded copy gets through. Zep says to replace the previous turn's Context Block instead of appending a second one, for its own reason: to preserve the cacheable prompt prefix.
Sources
Every page below was read on 27 September 2026. Each quotation on this page is taken verbatim from one of them, and where a sentence rests on source code rather than documentation, it says "in the source".
Research and libraries
- Carbonell and Goldstein, The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries, SIGIR 1998, authors' copy.
- LangChain: Vector stores; Recursive text splitter; TextSplitter reference; VectorStore.as_retriever; VectorStore.max_marginal_relevance_search; PGVector; PineconeVectorStore and its max_marginal_relevance_search; issue 39052, open; EnsembleRetriever; MultiVectorRetriever; indexing API; Cross encoder reranker; EmbeddingsRedundantFilter and its source at 7c10a5f; LangGraph stores.
- LlamaIndex: Basic strategies; Maximum Marginal Relevance retrieval; Node postprocessor modules; RAG failure mode checklist; Ingestion pipeline; Node parser modules; Memory.
Vector databases and search systems
- Qdrant: course, Top-K Retrieval, Chunking Strategies, Find Your Problem, Grouping, Diversity: Maximal Marginal Relevance; Search; Search Relevance; Points; OpenAPI schema on master; blog, Balancing Relevance and Diversity with MMR Search and How to Clean Up a Qdrant Collection.
- Weaviate: Vector search concepts; Hybrid search; Vector similarity search; Create objects; Query Agent search mode.
- Milvus: Grouping Search; Basic ANN Search; Insert Entities; LangChain basic usage.
- Chroma: Chunking; Group By and Aggregation, Chroma Cloud; Search API overview, Chroma Cloud; Query and Get; Adding Data; Update Data.
- Pinecone: Semantic search; Hybrid search with separate indexes; Upsert data; Retrieve context snippets.
- pgvector and PostgreSQL: pgvector README; PostgreSQL SELECT.
- Elasticsearch and OpenSearch: Elasticsearch, Collapse search results; Elasticsearch, Diversify retriever; OpenSearch, Collapse search results; OpenSearch, Collapse processor.
- OpenAI: Retrieval guide.
Memory layers
- Mem0: Search; Search Memories API reference; CLI; Claude Code integration; Platform v2 to v3 migration; Dream; SDK and plugin changelog; DeepSeek Harness plugin; source on main: agent-plugin-core memory_core.py, mem0/memory/main.py.
- Supermemory: Search; Search memory entries API reference; Cursor integration; changelog, 6 February 2026 and 28 April 2026; blog, Text chunking strategies for RAG and Infinitely running stateful coding agents; source on main: claude-supermemory recall-directive.js, supermemory_openai utils.py.
- Zep and Graphiti: Searching the graph; graph search SDK reference; Prepare data for ingestion; AG2 memory; Retrieving context; Graphiti searching; Graphiti source on main: search_utils.py, search.py.
- Cognee: Search; Recall; source on main: hybrid pairs.py.
- Letta: blog, RAG Is Not Agent Memory; Subagents; passages.search, V1 SDK (legacy) reference; tools.create, V1 SDK (legacy) reference.
Mnemoverse
Related
- Agent Memory Write Gate: What Never Gets Stored
- Agent Memory Consolidation Compared
- Agent Memory Deduplication: The Missing Error Rate
- RAG vs Agent Memory: What the Source Code Actually Shows
- Outcome Feedback in Agent Memory: Who Reads the Rating?
- Why Agent Memory Needs Sleep
- Memory MCP: How to Give AI Agents Persistent Memory
Edward Izgorodin · Mnemoverse · 2026-10-01
Mnemoverse is a persistent-memory API for AI agents. Free key: console.mnemoverse.com · Plans and limits · Docs: Getting Started
