Fine-tuning means continuing to train an existing AI model on your own labelled examples so its internal weights change, while prompt engineering means changing only the words you send the model at the moment you use it — and the practical rule is to exhaust prompting first, then fine-tune only when you have hundreds or more consistent labelled examples, a stable task, and prompting with a few examples has stopped improving.
## What fine-tuning actually changes A model is a large set of numbers called weights. Those numbers get set during the original training run, when the model reads enormous amounts of text and adjusts itself to predict what comes next. Fine-tuning takes that finished model and runs more training on a much smaller, much more specific dataset — usually examples you wrote, labelled, or corrected yourself.
Each example nudges the weights a little, so the model's default behaviour shifts toward your task. Prompt engineering touches none of that. You leave the weights exactly as they are and change the input: adding instructions, examples, formatting rules, or a role description.
The model's underlying tendencies stay the same; you are just steering them harder. That difference explains almost every trade-off that follows. A prompt change takes effect instantly and costs nothing to undo.
A fine-tune produces a new file you have to store, version, and re-run whenever the base model changes. It also explains why fine-tuning can teach a style or a format reliably but is a poor tool for teaching fresh facts — the model is learning a pattern of behaviour, not a lookup table.
## When prompting is enough
Prompting is enough for the large majority of real tasks, and people jump to fine-tuning far too early. If your task can be described in a paragraph, if the right answer varies a lot between inputs, or if you are still figuring out what you actually want, prompting is the correct tool.
Few-shot prompting — putting two to five worked examples directly in the prompt — often closes most of the gap, because the model copies the pattern you show it. Say you run a support inbox and want every reply to open with an apology, name the customer's order, and close with a link to the returns page.
You can write that as a prompt with three sample replies and get usable output the same afternoon. No dataset, no training run, no new model to maintain. The moment you change your mind about the tone, you edit a sentence.
That agility is the whole point, and it is why prompting should always be the first attempt, not the fallback.
## When to fine-tune
Fine-tuning earns its cost when three conditions line up at once. First, you have hundreds or more labelled examples that are genuinely consistent — not a pile of outputs you half-agree with. Second, the task is stable: the same input shape, the same output shape, week after week.
Third, prompting plus few-shot examples has plateaued, meaning you have tried several formulations and the error rate will not move. A classification job fits well. Suppose you need to tag ten thousand incoming messages as refund request, shipping question, or technical fault, and you have already labelled two thousand of them by hand.
Prompting might get you most of the way, but if the label boundaries are subtle and the prompt is already long, a fine-tune can lock in the pattern and let you drop the examples from every call, which shortens each request. According to our AI tool database, which tracks 360 AI tools with pricing and capability snapshots recorded at verification time, the platforms that support fine-tuning vary widely in what they charge and what they let you tune, so the cost side of this decision is genuinely tool-dependent rather than fixed.
The training run itself is usually the smaller expense. The bigger ongoing cost is maintenance: every time the base model is updated, your fine-tune may need redoing, and every time your task drifts, your labelled set goes stale. Budget for that before you start.
## The limits and the honest costs
Fine-tuning does not fix a vague problem, and it does not reliably add knowledge the model never had. It also will not save money if your volume is low, because you pay the setup cost regardless of how many calls you make. It fails badly when your labels are inconsistent — the model learns your confusion, and it learns it permanently.
There is also a real risk of narrowing the model: a fine-tune tuned hard on one format can get worse at everything else, so if the same assistant handles several jobs, you may need separate models or accept the regression. The practical middle path is worth knowing. Keep prompting as the default, keep a running log of inputs where the output was wrong, and label those as you go.
That log becomes your fine-tuning dataset if you ever need one, and it costs you almost nothing to maintain. Most teams that follow this pattern find prompting plus a few good examples handles the job, and the ones that do fine-tune arrive there with clean data instead of a scramble.