Cheap RAG in Go with Gemini File Search: no vector DB, two calls, one hosted store

Published: 2026-09-22 · Rewritten: 2026-09-23
A small glass box holding a glowing orb on a table, with a huge tangle of pipes and gears dissolving behind it.
Most RAG stacks are heavy machinery you maintain forever; the argument here is that one small hosted store can replace the whole tangle. AI-generated illustration

RAG in Go with Gemini File Search means you upload your documents to a store Google hosts, then ask a question and get an answer grounded in those documents — without running embeddings, a vector database, or a retrieval pipeline yourself. The whole flow is two API calls: one to create or populate the store, one to query it.

That sounds like a shortcut, and it is. The real question is whether it holds up for your workload or collapses the moment you need control over retrieval. Most teams reach for a vector database by default — pgvector, Qdrant, Weaviate — and then spend weeks on chunking strategy, embedding model choice, and index tuning before shipping anything. If your documents are ordinary text and your queries are plain questions, you can skip most of that. Here's what you're actually trading away, and the specific symptoms that tell you it's time to build your own pipeline.

What File Search actually does under the hood

You hand Google a set of files. Google chunks them, embeds them, stores the vectors in a managed index, and at query time retrieves the relevant passages, then passes them to a Gemini model to generate a grounded answer. The retrieval step is the part you're outsourcing.

This is why the "no vector DB" claim is accurate but slightly misleading. There is a vector database. You just don't operate it. You don't size it, back it up, migrate its schema, or pay a separate vendor for it. That's the actual saving — not the absence of vectors, but the absence of vector infrastructure on your side.

The trade-off is symmetric: you also can't tune it. You don't choose the chunk size, you don't pick the embedding model, and you don't get to inject a custom reranker between retrieval and generation. For a lot of internal tools and prototypes, none of that matters. For anything where retrieval quality is the product, it matters a lot.

The Go SDK calls, concretely

The Go client for Gemini lives in the google.golang.org/genai module. You authenticate with an API key from AI Studio or a service account, create a client, then create a file search store and import files into it. Import is the slow part — you're uploading bytes and waiting for Google to chunk and embed them.

Once the store is populated, a query is a single generate-content call with the file search store attached as a tool. You pass your question as text, the model retrieves passages from the store, and the response comes back with grounding metadata that tells you which chunks were used. That metadata is the thing people forget to read. It's how you debug a bad answer without guessing.

Two calls, one store, no separate retrieval service. The upload call is one-time per document set; the query call is what runs in your hot path.

File types, size limits, and re-indexing

File Search accepts common document formats — plain text, PDF, and the usual office document types. It does not accept arbitrary binary blobs, and it will reject files it can't extract text from. If your corpus is scanned PDFs with no text layer, you need OCR before upload, and that's on you.

Re-indexing is where the model gets awkward. There's no "update this one document in place" operation that behaves the way a database UPDATE does. You delete the old file from the store and import the new version. For a corpus that changes hourly, that's a real cost — you're re-uploading and re-embedding, and you're paying for those tokens each time.

For a corpus that changes weekly or monthly, it's fine. This is the single biggest practical constraint, and it's the one that pushes teams toward their own pipeline more often than retrieval quality does.

When to abandon it and build your own

Don't switch because a blog post told you managed RAG doesn't scale. Switch when you can point at a specific failure. Three symptoms are worth watching for:

If none of those apply, you're paying for control you don't use. Stay on the managed store and spend your engineering time on the prompt and the answer format instead.

A worked example

Say you're building an internal tool that answers questions about your company's engineering runbooks — a few dozen markdown files, updated when someone changes a process, maybe weekly. Queries are things like "what's the rollback procedure for the payments service."

Conventional route: stand up pgvector, write a chunker, pick an embedding model, build an ingest job, write a retrieval function, wire it to a generation call, then tune chunk overlap because your first answers missed context. That's a real week of work before the first useful answer.

File Search route: create the store, import the markdown files, wire one query call into your handler. You're answering questions the same afternoon. The retrieval quality is whatever Google's chunking gives you — usually adequate for prose runbooks, occasionally wrong on tables and step lists.

Where it breaks: someone asks "what changed in the deploy process last month," and the answer depends on comparing two versions of a file. File Search retrieves passages, not diffs. That query needs your own logic on top, regardless of which retrieval layer you chose.

What this approach does badly

Be honest about the failure modes before you commit. Managed retrieval gives you no visibility into why a passage was chosen. When an answer is wrong, you can read the grounding metadata to see which chunks were used, but you can't see the scores that ranked them or the chunks that were rejected. Debugging is coarser.

You also inherit Google's chunking decisions. If your documents have structure that matters — numbered procedures, tables, code blocks — the chunker may split them in ways that lose meaning. You can't fix that by tweaking a parameter, because there's no parameter.

And the cost model is usage-based, which means it scales with how much you upload and how often. Pricing for these APIs changes frequently — check the vendor's own page rather than trusting a figure from an article, including this one.

Where the model choice actually matters

File Search is a Gemini feature, so you're on Gemini models for generation. That's a real constraint if you've standardized on something else. The current Gemini lineup includes a 1M-token context window and multimodal input, and Google's consumer tier runs $19.99/mo through Google One AI Premium — but that's chat pricing, not API pricing, and the two aren't interchangeable. For an API workload, budget from the usage-based API rates, not from a consumer subscription.

If your team is already standardized on another model family, the honest answer is that File Search may not fit, and a self-hosted pipeline with your preferred model is the better path. The retrieval convenience doesn't outweigh re-platforming your generation layer.

Key Takeaways

The decision, stated plainly

Use File Search when your corpus is prose, changes slowly, and your queries are plain questions. That covers a surprising amount of internal tooling. Build your own retrieval when you need to rank passages yourself, when your documents have structure worth preserving, or when your corpus churns fast enough that re-import becomes the bottleneck.

The mistake isn't choosing managed RAG — it's choosing it by default and discovering the constraint six weeks in, after you've built application logic that assumes retrieval scores you don't have. Pick the failure mode you can live with up front. For most Go services answering questions over a stable document set, the two-call version is the right starting point, and you can migrate later when a specific symptom shows up.

Sources

Frequently Asked Questions

Do I still need a vector database if I use Gemini File Search?

No. File Search maintains a managed vector index on Google's side — it chunks, embeds, and stores your documents, then retrieves relevant passages at query time. You don't run pgvector, Qdrant, or any other vector store, and you don't manage embeddings yourself. The trade-off is that you also can't tune chunking, choose the embedding model, or inject a custom reranker between retrieval and generation.

How do I update a document that's already in the store?

There's no in-place update. You delete the old file from the store and import the new version, which means re-uploading and re-embedding it. For a corpus that changes weekly or monthly, that's fine. For one that changes hourly, the delete-and-reimport cycle becomes the dominant cost and latency, and that's usually the point where teams move to their own pipeline with incremental indexing.

What kinds of documents does File Search handle badly?

Scanned PDFs with no text layer need OCR before upload — File Search extracts text and rejects files it can't read. Documents where structure carries meaning, like tables, numbered procedures, and code blocks, can get split by the chunker in ways that lose context, and there's no parameter to fix it. Structured records are better retrieved by field from a database than by paragraph from a text store.

A short plank bridge spans a gap but splinters and breaks at both ends above a deep drop.
The two-call approach is cheap and fast, but its limits are real: the bridge holds only for the questions it was built to carry. AI-generated illustration

How this article was produced: it was generated by an automated content pipeline from the sources listed above. No human editor wrote or reviewed it, and we did not personally test the tools described. Facts and prices that appear here come from our own AI tool database, and its verification date is noted where relevant. Spotted an error? Tell us and we will correct or remove it.

Want to try this yourself? AI-Mind generates content from a plain description — no prompt engineering required.

Try AI-Mind