A RAG pipeline has two jobs, and most teams only build one. The first is finding documents that might be relevant — that's retrieval. The second is deciding which of those documents actually earn a place in the model's context window. That second job is the selection layer, and it's where most pipelines quietly fall apart.
Here's the problem in concrete terms. Your vector search returns the top 50 chunks by cosine similarity. You can't paste 50 chunks into a prompt without blowing your budget and burying the answer in noise. So which ones do you keep? If you keep the wrong four, the model either hallucinates or hedges. The retrieval worked. The selection didn't.
Why similarity search alone can't answer this
Embedding similarity measures whether two pieces of text are about similar things. It does not measure whether a chunk contains the answer to a specific question.
Consider a support bot for a SaaS product. A user asks: "Why did my invoice double this month?" The embedding search returns five chunks about billing — the pricing page, the refund policy, a changelog entry, a support macro, and the actual seat-count proration rule. Four of those are topically perfect. Only one answers the question. Cosine similarity ranks all five similarly because they all live in the same semantic neighborhood.
This is the core limitation. Bi-encoders — the models behind vector search — encode the query and the document separately, then compare vectors. That's what makes them fast enough to search millions of chunks. It's also why they can't tell the difference between "about billing" and "answers this billing question."
Reranking: the layer that actually does the selecting
The standard fix is a two-stage pipeline. Stage one retrieves broadly with a cheap bi-encoder. Stage two reranks that shortlist with a cross-encoder.
A cross-encoder takes the query and the document together as a single input and outputs a relevance score. Because it sees both at once, it can model the interaction between them — does this chunk actually address what was asked? That's a much harder question than "are these two things about the same topic," and it's why cross-encoders consistently outperform pure vector similarity on the final ranking.
The cost is speed. You can't run a cross-encoder across a million documents. You run it across 50, or 100, and you take the top handful. That's the whole trade: cheap-and-broad for the recall stage, expensive-and-precise for the precision stage.
Named options in this space include the BGE reranker family, Cohere Rerank, and mxbai-rerank. If you'd rather not host anything, Cohere Rerank is a hosted API; if you want to run it locally, the BGE cross-encoder models are open weights. There are also purpose-built reranking services that wrap these for you. Pricing and model versions in this category move fast, so check the vendor's own page rather than trusting a number from a blog post — including this one.
What to do when you can't rerank
Not every pipeline can afford a second model call. If you're on a tight latency budget or you're running retrieval inside a serverless function with cold-start constraints, a reranker may not be viable.
In that case, the fallback is structural rather than model-based. A few things help:
- Metadata filtering before retrieval. If the query is about a specific product tier, filter to that tier first. This shrinks the candidate set so the top-k you return is already more likely to be relevant.
- Chunk overlap and size tuning. Chunks that are too small lose context; too large and they dilute the signal. There's no universal right answer — it depends on how your documents are structured.
- Reciprocal rank fusion. If you're running more than one retrieval method (say, vector plus keyword), RRF combines their rankings without needing a trained model. It's a cheap way to get some of the benefit of a reranker.
None of these fully replaces a cross-encoder. They're damage control, not a substitute. Be honest with yourself about which one you're doing.
The long-context trap
There's a tempting shortcut: if the model has a huge context window, why not just stuff everything in and let it sort out the relevant parts?
It's worth knowing how large those windows actually are now. According to our internal AI tool database, both Claude Opus 4.8 and OpenAI's GPT-5.5 ship with 1M-token context windows. That's genuinely enormous — large enough that "just include everything" starts to feel reasonable.
It isn't. Two things go wrong. First, cost and latency scale with input length, so a 200k-token prompt is expensive and slow on every single call. Second, and more insidiously, models get worse at finding a needle in a haystack as the haystack grows. A relevant chunk buried in 190k tokens of filler is less likely to be used than the same chunk sitting in a 4k-token prompt. The context window is a ceiling, not a target.
Treat long context as a safety margin for the cases where you genuinely need a lot of material — not as a reason to skip the selection layer.
A worked example
Say you're building a docs assistant for a developer tool. You have 4,000 pages of documentation, changelogs, and forum threads indexed.
A developer asks: "How do I set a custom timeout on the streaming client?"
Your vector search returns 40 chunks. The top five by similarity are: the streaming client overview, the general configuration guide, a forum post about timeouts on a different SDK, a changelog entry mentioning a timeout default change, and the actual API reference for the timeout parameter.
Only the last one answers the question. The changelog is related but describes a default value, not how to set it. The forum post is about the wrong SDK. Without a reranker, there's a real chance the model sees the forum post ranked highly and gives an answer for the wrong client.
With a cross-encoder reranking those 40 down to the top 3, the API reference rises to the top because it directly addresses "how do I set X." The changelog drops. The model gets a clean context and answers correctly.
That's the entire value of the selection layer: it's the difference between "we retrieved the right document" and "the model used the right document." Those are not the same thing.
Where this gets hard
Reranking isn't free of problems. A few honest caveats:
- It adds latency. Every reranked query is a second model pass. For interactive apps, that matters.
- It can be wrong in new ways. A cross-encoder trained on general web data may not understand your domain's notion of relevance. Domain-specific rerankers exist but require training data you probably don't have yet.
- Garbage in, garbage out. If your retrieval stage never surfaces the right chunk in the first 50, no reranker will save you. Reranking fixes precision, not recall.
- Evaluation is the hard part. You can't tune a reranker without a way to measure whether the final answer improved. That's a separate project.
The teams that get this right usually start with retrieval quality, then add reranking, then measure. Skipping to the reranker before your retrieval is solid is a common and expensive mistake.
Key Takeaways
- Retrieval and selection are separate problems; most RAG failures happen in the selection layer, not retrieval.
- Bi-encoders find topically similar chunks; cross-encoders score actual relevance. You need both.
- Named rerankers worth evaluating: BGE reranker, Cohere Rerank, and mxbai-rerank.
- Long context is a safety margin, not a substitute for selection — cost and accuracy both degrade as prompts grow.
- Reranking fixes precision, not recall. If the right chunk never surfaces, no reranker helps.
The short version
If your RAG pipeline is returning documents but the model still gives weak answers, the bug is almost never in the vector search. It's in what you hand the model. Build the selection layer: retrieve broadly, rerank narrowly, and measure whether the final answer actually improved. That last step is the one everyone skips, and it's the only one that tells you if any of this worked.
Sources
- AI Tool Database, Internal AI Tool Snapshot, 2026. Verified pricing and capability records for 360 AI tools, most recently verified 2026-09-18.
- AI Tool Database, Claude — Tool Record, 2026. Anthropic's Claude Opus 4.8, listed with a 1M-token context window.
- AI Tool Database, ChatGPT — Tool Record, 2026. OpenAI's GPT-5.5, listed with a 1M-token context window.
Frequently Asked Questions
Do I always need a reranker in a RAG pipeline?
No. If your retrieval already returns a small, high-precision candidate set — say, because you're filtering hard on metadata first — a reranker may add latency for little gain. The reranker earns its cost when your retrieval stage returns a broad shortlist and you need to separate the chunk that answers the question from the chunks that merely share its topic.
Can a large context window replace the selection layer?
It can absorb more input, but it doesn't solve the problem. Cost and latency scale with prompt length, and models get less reliable at using a relevant passage when it's buried among large amounts of unrelated text. Long context is best treated as headroom for cases that genuinely need it, not as a reason to stop ranking what you send.
What's the difference between a bi-encoder and a cross-encoder?
A bi-encoder encodes the query and each document separately, then compares the resulting vectors. That's fast enough to search large collections, which is why it powers retrieval. A cross-encoder processes the query and document together, so it can judge whether the document actually answers the query. It's slower, so it's used to rerank a shortlist rather than search everything.