Fine-tuning is the process of taking a pretrained model and continuing its training on a smaller, curated dataset of your own examples, so the model’s weights — the numbers that determine how it responds — actually change. That’s the key distinction from prompting: prompting steers a model at the moment you call it, using instructions and context that vanish once the response comes back, while fine-tuning bakes a pattern permanently into the model itself. It’s powerful for narrow, repeatable behavior, and it’s frequently the wrong first move for problems that prompting or retrieval can already solve more cheaply.
How Does Fine-Tuning Actually Change a Model?
A pretrained model’s behavior is encoded in billions of parameters set during its original training run. Full fine-tuning continues that same process on your dataset: you feed the model input-output pairs, it makes predictions, and its weights get nudged via gradient updates toward your examples. Do this enough, with clean enough data, and the model reliably reproduces the patterns you showed it — a certain output format, a tone, a classification scheme — without you having to re-explain them every time.
Most modern fine-tuning doesn’t touch every parameter, though. Techniques like LoRA (Low-Rank Adaptation) freeze the base model’s weights entirely and instead train small, injected “adapter” matrices alongside each layer. The original LoRA paper reported cutting the number of trainable parameters by roughly 10,000x compared to full fine-tuning of GPT-3, with no added inference latency once the adapter is merged in. That’s why LoRA and its variants (QLoRA, and other methods in Hugging Face’s PEFT library) have become the default way most teams fine-tune open models like Llama or Mistral — it’s dramatically cheaper in compute and storage, and you can swap adapters in and out of the same base model for different tasks.
Prompting, by contrast, changes none of this. It works entirely through in-context learning — the model reads your instructions and examples fresh each call and infers what to do, but nothing persists once the response is generated. That’s both prompting’s weakness (you pay the token cost of re-stating context every time) and its strength (you can change behavior instantly, with no retraining and no risk of damaging the base model).
Fine-Tuning vs. Prompting vs. RAG: Which Should You Use?
These three approaches solve different problems, and conflating them is the single most common reason teams reach for fine-tuning when they didn’t need it.
- Prompting / few-shot examples — best when the model already has the underlying capability and just needs clearer instructions, formatting rules, or a handful of examples of what “good” looks like. Cheapest, fastest to iterate, zero training risk.
- RAG (retrieval-augmented generation) — best when the problem is that the model doesn’t know something: current facts, your internal documents, product data that changes weekly. RAG retrieves relevant content at query time and injects it into the prompt, so the model answers from real, up-to-date source material instead of from what it memorized during training. See our plain-English guide to RAG for how that retrieval step actually works.
- Fine-tuning — best when the model knows the material but doesn’t behave the way you need it to, consistently, at scale: a rigid output schema it keeps drifting from, a house style instructions alone can’t reliably enforce, a classification task with dozens of labels, or domain jargon so specialized that few-shot examples eat your entire context window.
A useful shortcut: if the problem is “the model doesn’t know X,” reach for RAG. If the problem is “the model knows X but won’t consistently do Y with it,” fine-tuning starts to make sense — but only after you’ve genuinely exhausted prompting. Many teams skip straight to fine-tuning and discover months later that a better system prompt and a few dozen well-chosen examples would have closed most of the gap.
What Does Fine-Tuning Actually Cost You?
The upfront compute cost of fine-tuning, especially with LoRA-style methods, has become almost a minor line item. The real costs are elsewhere.
Data curation. Providers set low technical minimums — OpenAI’s supervised fine-tuning guide, for example, lists 10 examples as a hard floor but recommends starting with around 50 well-crafted demonstrations — but “well-crafted” is doing a lot of work in that sentence. Sloppy or inconsistent examples teach the model your inconsistencies just as effectively as your intent.
Overfitting and catastrophic forgetting. Train too hard on a narrow, small dataset and the model can overfit to your examples’ quirks while losing general capabilities it had before — a well-documented failure mode researchers call catastrophic forgetting. A model fine-tuned hard on customer-support tickets can become noticeably worse at unrelated reasoning it used to handle fine.
Maintenance burden. A fine-tuned model is tied to the base model version it was trained on. When the provider deprecates that base model, your fine-tuned checkpoint typically goes with it, meaning you retrain. This isn’t hypothetical: OpenAI’s own deprecations documentation shows it has been winding down direct access to its fine-tuning platform through 2026, restricting new job creation for organizations without prior fine-tuning history, then for accounts without recent usage — while fine-tuned models keep running only until their underlying base model is retired. Google’s Vertex AI still offers supervised tuning for Gemini models, and Anthropic’s Claude 3 Haiku is fine-tunable through Amazon Bedrock rather than a first-party Anthropic API — which is exactly why this is worth checking against current provider docs before you commit, not assuming last year’s offering still stands.
Opportunity cost. Every fine-tuning cycle — collect data, train, evaluate, deploy — is slower than editing a prompt or adjusting a retrieval index. If your requirements are still moving, that iteration speed matters more than the marginal quality gain fine-tuning might buy you.
When Do You Actually Need Fine-Tuning?
Fine-tuning earns its cost in a fairly specific set of situations:
- You need a smaller, cheaper, faster model to match a larger model’s quality on one narrow, high-volume task — fine-tuning a compact model can close that gap without paying frontier-model prices at scale.
- Your task requires a rigid, non-negotiable output format (structured extraction, a specific JSON schema, classification across many fine-grained labels) that prompting keeps drifting from under real-world input variety.
- Your few-shot examples are so long or numerous that they’re bloating your context window and costing more in tokens, per call, than a one-time training run would.
- You have genuinely large volumes of high-quality labeled data — hundreds to thousands of examples — reflecting the real distribution of inputs you’ll see in production, not a handful of idealized cases.
If none of those apply, you likely don’t need it yet. Push prompting and structured examples further than feels comfortable first — it’s reversible, cheap, and faster to test than most teams assume.
Frequently Asked Questions
Does fine-tuning teach a model new facts?
Not reliably. Fine-tuning is better understood as teaching behavior — format, tone, task structure — than as a way to inject knowledge. Facts learned through fine-tuning can be inconsistently recalled or subtly distorted, and they go stale the moment your underlying information changes, since updating them means retraining. If your actual problem is “the model needs access to current or proprietary information,” RAG is almost always the better tool, because it retrieves facts at query time rather than baking them into weights.
How much data do you actually need to fine-tune a model?
It depends heavily on the task and method. Provider platforms often set low technical minimums — OpenAI’s guide lists 10 examples as a floor and recommends roughly 50 as a realistic starting point — but that’s for narrow, well-defined tasks. Open-model fine-tuning with LoRA commonly uses anywhere from a few hundred to several thousand examples depending on task complexity. In every case, diversity and quality of examples matter more than raw count: a few hundred examples that genuinely cover your real input distribution beat several thousand near-duplicates.
Can you combine fine-tuning with prompting or RAG?
Yes, and in production systems this is more common than using any one technique alone. A typical stack fine-tunes a model for consistent output format and domain tone, uses RAG to pull in current or proprietary facts at query time, and still relies on prompting to give per-request instructions on top of both. The three techniques address different weaknesses, so layering them is usually more effective than betting everything on one.
The Bottom Line
Fine-tuning changes what a model is; prompting and RAG change what it sees at the moment you ask it something. That distinction should drive the decision, not which technique sounds more sophisticated. Before committing to a training run, it’s worth reading the current state of the platform you’re considering directly — OpenAI’s own supervised fine-tuning documentation is a good example of how fast provider offerings shift, and Hugging Face’s PEFT/LoRA documentation is the clearest technical reference if you’re fine-tuning an open model yourself rather than going through a provider API.
Photo credit: “BalticServers data center” by BalticServers.com, licensed under CC BY-SA 3.0.



