Your RAG Finds the Documents. But Which Ones Should Reach the LLM?

Published: 2026-09-22 · Rewritten: 2026-09-23
Many glowing paper sheets funnel downward, most fading away while a few pass through a narrow gap into a glass vessel.
Retrieval finds plenty; the real work is deciding which few documents earn a place in the answer. AI-generated illustration

A RAG pipeline has two jobs, and most teams only build one. The first is finding documents that might be relevant — that's retrieval. The second is deciding which of those documents actually earn a place in the model's context window. That second job is the selection layer, and it's where most pipelines quietly fall apart.

Here's the problem in concrete terms. Your vector search returns the top 50 chunks by cosine similarity. You can't paste 50 chunks into a prompt without blowing your budget and burying the answer in noise. So which ones do you keep? If you keep the wrong four, the model either hallucinates or hedges. The retrieval worked. The selection didn't.

Why similarity search alone can't answer this

Embedding similarity measures whether two pieces of text are about similar things. It does not measure whether a chunk contains the answer to a specific question.

Consider a support bot for a SaaS product. A user asks: "Why did my invoice double this month?" The embedding search returns five chunks about billing — the pricing page, the refund policy, a changelog entry, a support macro, and the actual seat-count proration rule. Four of those are topically perfect. Only one answers the question. Cosine similarity ranks all five similarly because they all live in the same semantic neighborhood.

This is the core limitation. Bi-encoders — the models behind vector search — encode the query and the document separately, then compare vectors. That's what makes them fast enough to search millions of chunks. It's also why they can't tell the difference between "about billing" and "answers this billing question."

Reranking: the layer that actually does the selecting

The standard fix is a two-stage pipeline. Stage one retrieves broadly with a cheap bi-encoder. Stage two reranks that shortlist with a cross-encoder.

A cross-encoder takes the query and the document together as a single input and outputs a relevance score. Because it sees both at once, it can model the interaction between them — does this chunk actually address what was asked? That's a much harder question than "are these two things about the same topic," and it's why cross-encoders consistently outperform pure vector similarity on the final ranking.

The cost is speed. You can't run a cross-encoder across a million documents. You run it across 50, or 100, and you take the top handful. That's the whole trade: cheap-and-broad for the recall stage, expensive-and-precise for the precision stage.

Named options in this space include the BGE reranker family, Cohere Rerank, and mxbai-rerank. If you'd rather not host anything, Cohere Rerank is a hosted API; if you want to run it locally, the BGE cross-encoder models are open weights. There are also purpose-built reranking services that wrap these for you. Pricing and model versions in this category move fast, so check the vendor's own page rather than trusting a number from a blog post — including this one.

What to do when you can't rerank

Not every pipeline can afford a second model call. If you're on a tight latency budget or you're running retrieval inside a serverless function with cold-start constraints, a reranker may not be viable.

In that case, the fallback is structural rather than model-based. A few things help:

None of these fully replaces a cross-encoder. They're damage control, not a substitute. Be honest with yourself about which one you're doing.

The long-context trap

There's a tempting shortcut: if the model has a huge context window, why not just stuff everything in and let it sort out the relevant parts?

It's worth knowing how large those windows actually are now. According to our internal AI tool database, both Claude Opus 4.8 and OpenAI's GPT-5.5 ship with 1M-token context windows. That's genuinely enormous — large enough that "just include everything" starts to feel reasonable.

It isn't. Two things go wrong. First, cost and latency scale with input length, so a 200k-token prompt is expensive and slow on every single call. Second, and more insidiously, models get worse at finding a needle in a haystack as the haystack grows. A relevant chunk buried in 190k tokens of filler is less likely to be used than the same chunk sitting in a 4k-token prompt. The context window is a ceiling, not a target.

Treat long context as a safety margin for the cases where you genuinely need a lot of material — not as a reason to skip the selection layer.

A worked example

Say you're building a docs assistant for a developer tool. You have 4,000 pages of documentation, changelogs, and forum threads indexed.

A developer asks: "How do I set a custom timeout on the streaming client?"

Your vector search returns 40 chunks. The top five by similarity are: the streaming client overview, the general configuration guide, a forum post about timeouts on a different SDK, a changelog entry mentioning a timeout default change, and the actual API reference for the timeout parameter.

Only the last one answers the question. The changelog is related but describes a default value, not how to set it. The forum post is about the wrong SDK. Without a reranker, there's a real chance the model sees the forum post ranked highly and gives an answer for the wrong client.

With a cross-encoder reranking those 40 down to the top 3, the API reference rises to the top because it directly addresses "how do I set X." The changelog drops. The model gets a clean context and answers correctly.

That's the entire value of the selection layer: it's the difference between "we retrieved the right document" and "the model used the right document." Those are not the same thing.

Where this gets hard

Reranking isn't free of problems. A few honest caveats:

The teams that get this right usually start with retrieval quality, then add reranking, then measure. Skipping to the reranker before your retrieval is solid is a common and expensive mistake.

Key Takeaways

The short version

If your RAG pipeline is returning documents but the model still gives weak answers, the bug is almost never in the vector search. It's in what you hand the model. Build the selection layer: retrieve broadly, rerank narrowly, and measure whether the final answer actually improved. That last step is the one everyone skips, and it's the only one that tells you if any of this worked.

Sources

Frequently Asked Questions

Do I always need a reranker in a RAG pipeline?

No. If your retrieval already returns a small, high-precision candidate set — say, because you're filtering hard on metadata first — a reranker may add latency for little gain. The reranker earns its cost when your retrieval stage returns a broad shortlist and you need to separate the chunk that answers the question from the chunks that merely share its topic.

Can a large context window replace the selection layer?

It can absorb more input, but it doesn't solve the problem. Cost and latency scale with prompt length, and models get less reliable at using a relevant passage when it's buried among large amounts of unrelated text. Long context is best treated as headroom for cases that genuinely need it, not as a reason to stop ranking what you send.

What's the difference between a bi-encoder and a cross-encoder?

A bi-encoder encodes the query and each document separately, then compares the resulting vectors. That's fast enough to search large collections, which is why it powers retrieval. A cross-encoder processes the query and document together, so it can judge whether the document actually answers the query. It's slower, so it's used to rerank a shortlist rather than search everything.

Oversized soft cubes crowd against a small fixed glass window; only a few small cubes fit inside the frame.
A context window is a fixed aperture, not a warehouse, so filling it with everything weakens the answer. AI-generated illustration
Two conveyor belts of grey blocks: one in arrival order, the other reordered by size and brightness before a lit slot.
Reranking is a second, cheaper sorting pass that changes what actually reaches the model. AI-generated illustration

How this article was produced: it was generated by an automated content pipeline from the sources listed above. No human editor wrote or reviewed it, and we did not personally test the tools described. Facts and prices that appear here come from our own AI tool database, and its verification date is noted where relevant. Spotted an error? Tell us and we will correct or remove it.

Want to try this yourself? AI-Mind generates content from a plain description — no prompt engineering required.

Try AI-Mind