Fine-tuning means taking an AI model that has already been trained on a huge amount of general text, then continuing to train it on your own labelled examples so its internal weights shift toward your specific task — and you should only reach for it when prompting and retrieval have genuinely hit a wall, not as a first move.
The short version: prompting changes what you ask, retrieval changes what the model can see, and fine-tuning changes the model itself. That last one is the most expensive and the hardest to undo, which is exactly why it should be your third option, not your first.
The mechanism matters here, so let's walk through it. A base model stores what it learned as billions of numeric weights — think of them as dials that were set during its original training. When you prompt, those dials stay frozen; you're just steering with words.
When you fine-tune, you feed in pairs of input and correct output, and a training process nudges the dials so the model's default behaviour moves toward your examples. A learning rate controls how big each nudge is — set it too high and the model lurches and forgets things; too low and nothing meaningful changes.
That's why fine-tuning is good at teaching style, format, and narrow classification patterns, and bad at teaching fresh facts. If you need the model to know about your company's return policy from last Tuesday, fine-tuning is the wrong tool — retrieval is. If you need every reply to sound exactly like your brand voice and follow a rigid JSON structure, fine-tuning is genuinely strong.
A concrete example makes the decision rule clearer. Say you run a support team and you want incoming tickets auto-tagged into twelve categories. You have 4,000 historical tickets that were already labelled by humans.
Start with prompting: paste the category list into a model and ask it to pick one. If that gets you, say, most of the way there and the misses are tolerable, stop — you're done, and you've spent nothing on training. If it plateaus because your categories are subtle and your tickets are short, retrieval won't help either (there's nothing to look up).
This is the case where fine-tuning earns its keep: you'd export those labelled tickets as JSONL, upload them to a platform like OpenAI's fine-tuning API or Google's Vertex AI, pick a small base model, and set a modest learning rate — commonly something in the low 1e-5 range for a small model — then evaluate on a held-out slice you never trained on. The output is a custom model ID you call exactly like the normal one.
Notice the shape of that story: labelled data existed, the task was narrow, and prompting had already been tried.
Now the honest limits, because this is where people get burned. Fine-tuning costs you two things people underestimate: labelling time and maintenance. Those 4,000 tickets didn't label themselves.
And when the base model you fine-tuned gets updated or retired by the vendor, your fine-tune doesn't automatically come along — you often have to redo the work against the new base. There's also overfitting: train too long on too few examples and the model memorises your training set, acing it while failing on anything new.
A related trap is catastrophic forgetting — push too hard on your narrow task and the model loses general ability it used to have, so it nails your format but writes worse prose everywhere else. Finally, fine-tuning is the wrong choice when your problem is missing knowledge (use retrieval), when you have fewer than a few hundred good examples, or when a simple rules-based script or a small off-the-shelf classifier would do the job for a fraction of the effort.
According to our AI tool database, which tracks 360 AI tools with pricing and capability snapshots verified as of 2026-09-18, the space is crowded enough that a plain classifier or a retrieval setup is often already available inside tools you're paying for. Check what you already have before you build a training pipeline.
A useful rule of thumb: if you can describe your task in a paragraph and the model gets it right most of the time, you don't need fine-tuning — you need a better prompt. Fine-tune when you've proven prompting fails and you own the labelled data to fix it. If the underlying idea of a model's weights still feels fuzzy, our explainer on what an AI model actually means is a good place to start, and for the failure mode that fine-tuning does not fix, see why AI sometimes makes things up.