What is fine-tuning, and when should you use it instead of RAG?

Fine-tuning is one of the most misunderstood levers in AI. It changes how a model behaves, its format, tone and consistency, not what it knows. Here is what fine-tuning actually does, how the cheap version called LoRA works, and the rule for when to fine-tune instead of reaching for retrieval or a longer prompt.

By Yash Malviya

Published

Detailed image of a server rack with glowing lights in a modern data center
Photo: panumas nikhomkhai / Pexels

What fine-tuning actually is

Fine-tuning is the process of taking a model that has already been trained on the open internet and training it a little more on a smaller dataset of your own, so it gets better at one job. Hugging Face's documentation puts it plainly: fine-tuning "continues training a large pretrained model on a smaller dataset specific to a task or domain," and it is "identical to pretraining except you don't start with random weights. It also requires far less compute, data, and time."

The important word is weights. Fine-tuning changes the numbers inside the model. That makes it different from the two things people often confuse it with. Prompting changes what you say to a fixed model and touches nothing inside it. Retrieval-augmented generation, or RAG, leaves the model alone and hands it relevant documents at the moment you ask. Fine-tuning is the only one of the three that actually edits the model.

Full fine-tuning, and the cheap kind

There are two ways to do it. Full fine-tuning updates every weight in the model. It works, but for a large model it is expensive, needs real hardware, and produces a whole new copy of the model to store and serve.

Most people no longer do that. The common approach is parameter-efficient fine-tuning, and the best-known method is LoRA. Instead of editing the model, LoRA freezes it and trains a small set of extra matrices that ride alongside it. The savings are not marginal. In the original 2021 paper, the authors reported that LoRA "can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times" when adapting GPT-3 175B, while matching full fine-tuning on quality. Because the adapter is tiny, it is cheap to train, quick to swap, and easy to throw away if it does not help.

“Fine-tuning is identical to pretraining except you don't start with random weights. It also requires far less compute, data, and time.”

Hugging Face, Transformers documentation
Dynamic 3D digital art featuring vibrant cosmic swirls and particles in a mesmerizing abstract design
Parameter-efficient methods like LoRA freeze the pretrained model and train a small adapter, cutting the cost of fine-tuning by orders of magnitude. Photo: Steve A Johnson / Pexels

The rule that saves you money: behavior, not knowledge

Here is the mistake that costs teams the most: fine-tuning a model to teach it facts. It is the wrong tool for that. Baking information into a model's weights freezes it in time. The moment your prices, policies or product details change, the fine-tuned model is out of date, and there is no clean way to make it cite the document an answer came from.

Fine-tuning is for behavior, not knowledge. Reach for it when you need a model to do something in a consistent, specific way that prompting cannot reliably pin down: emit a strict JSON schema, follow a house style or tone, fill a regulatory form, or respect domain conventions the base model keeps getting slightly wrong. A few hundred to a few thousand clean examples of input and desired output can collapse the variance that a long prompt cannot. That is fine-tuning earning its keep.

When to fine-tune, when to retrieve

The clean split most practitioners now use: if the answer depends on data that changes, or you need to show sources and pass an audit, use retrieval. If the problem is the shape of the output, the tone, or a fixed behavior, fine-tune. And increasingly the answer in production is both. The 2026 default is not fine-tune or RAG, it is fine-tune and RAG, with each doing what it is best at: tune the interface, retrieve the content. A common pattern is to fine-tune the small, stable parts, such as how the model rewrites a question into a search query or formats a cited answer, while keeping the actual facts in a retrieval system you can update without retraining.

The major providers all support it in some form. OpenAI offers fine-tuning for its models, including supervised and preference-based methods. Some Claude models can be fine-tuned through Amazon Bedrock. And open-source models such as Llama and Mistral can be tuned cheaply on your own hardware with Hugging Face and LoRA. What none of them offer is a shortcut around the real work, which is assembling a clean dataset and a way to measure whether the tuned model is actually better.

The order to try things in

For almost every project, the cheapest path that works is a ladder. Start with a better prompt, since newer models need less hand-holding than they used to. If the problem is missing or changing knowledge, add retrieval. Only when a behavior is stable, well-defined, and still wrong after good prompting should you fine-tune, because it is the most expensive option to build and the fastest to go stale. Climb the ladder in that order and you will solve most problems before you ever touch the weights.

Our take

Fine-tuning is real and useful, and it is oversold in exactly one direction. It is a behavior tool wearing the costume of a memory tool. The teams that get value from it treat it as a way to make a model reliably do a specific thing, and they keep their facts in a retrieval system where they can be updated and cited. The teams that waste money on it are trying to pour their company's knowledge into the weights and wondering why the model is confidently out of date a month later. Learn the one distinction, behavior versus knowledge, and you have most of what fine-tuning has to teach.

Frequently asked questions

What does fine-tuning an AI model do?

It continues training a pretrained model on a smaller, task-specific dataset, adjusting the model's weights. Hugging Face describes it as identical to pretraining except that you start from an existing model rather than random weights, so it needs far less compute, data and time. The result is a model tuned toward a particular behavior, format or domain.

What is the difference between fine-tuning and RAG?

Fine-tuning changes the model's weights to change how it behaves; retrieval-augmented generation leaves the model unchanged and feeds it relevant documents at query time. Use RAG for knowledge that changes or needs citations, and fine-tuning for a fixed output format, tone or behavior that prompting cannot pin down.

What is LoRA, and why is it cheaper?

LoRA is a parameter-efficient fine-tuning method that freezes the pretrained model and trains a small set of added matrices instead of all the weights. In its original paper it reduced trainable parameters by about 10,000 times and GPU memory by about 3 times when adapting GPT-3 175B, while matching full fine-tuning on quality.

When should you fine-tune instead of just prompting?

Fine-tune when the behavior you need is stable and prompting is unreliable, such as a strict output schema, a consistent house style, or domain conventions the base model keeps missing. For anything that depends on changing facts, retrieval is usually the better and cheaper choice.

Can you fine-tune models like GPT and Claude?

Yes. OpenAI offers fine-tuning for its models, including supervised and preference-based methods; some Claude models can be fine-tuned through Amazon Bedrock; and open-source models such as Llama and Mistral can be tuned cheaply using Hugging Face and LoRA.

Sources

What each one is, and whose it is.

  1. DocumentationIndependent of the vendor
  2. 2

    LoRA: Low-Rank Adaptation of Large Language Models, Hu et al., arXiv (June 17, 2021)

    PaperIndependent of the vendor
  3. PaperIndependent of the vendor
  4. DocumentationIndependent of the vendor