In-context learning is the ability of a large language model to pick up a new task just from the examples or instructions you put in its prompt, without any of its underlying weights being updated or retrained. Instead of fine-tuning a model on thousands of labeled examples, you simply show it what you want, right there in the conversation, and it adapts its output on the fly. This idea was popularized by OpenAI’s 2020 GPT-3 paper, “Language Models are Few-Shot Learners,” which showed that a large enough model could learn a task from a handful of demonstrations in its prompt almost as well as a smaller model fine-tuned specifically for that task.
What Exactly Happens When a Model “Learns” From a Prompt?
When you type a prompt, the model doesn’t change a single one of its parameters. Nothing is saved, nothing is trained, and the “learning” disappears the moment the conversation ends or the context window is cleared. What actually happens is that the model uses the patterns in your prompt, including any examples, labels, or instructions, as extra context to condition its next-token predictions. Because the model was pretrained on enormous amounts of text containing countless examples of tasks, instructions, and formats, it has already learned a general capacity to recognize “oh, this looks like a translation task” or “this looks like a sentiment-classification task” and to continue the pattern accordingly.
Brown et al. described this directly in the GPT-3 paper, showing that scaling up a language model’s size dramatically improved its “task-agnostic, few-shot performance,” with the model performing tasks purely through text interaction and “without any gradient updates or fine-tuning.” That distinction, no gradient updates, is the whole point of in-context learning: the knowledge used to solve the task lives in the prompt, not in the weights.
How Is In-Context Learning Different From Fine-Tuning?
Fine-tuning and in-context learning both aim to make a general-purpose model better at a specific task, but they work in fundamentally different ways:
- Fine-tuning updates the model’s actual weights using a training dataset, usually with gradient descent. The change is permanent (until the model is fine-tuned again) and persists across every future conversation, but it requires data, compute, and time to set up.
- In-context learning happens entirely inside a single prompt. You supply instructions or examples as text, the model conditions its response on them, and the effect vanishes once that context is gone. There’s no training run, no dataset curation pipeline, and no need for machine learning infrastructure.
In practice, this makes in-context learning the fastest way to steer a model’s behavior. If you need a chatbot to always answer in a certain format, you can often get there by adding two or three examples to the system prompt rather than collecting a dataset and running a fine-tuning job. Fine-tuning still matters when you need consistent behavior baked in permanently, at scale, or when the task requires more examples than fit inside the context window.
What’s the Difference Between Zero-Shot, One-Shot, and Few-Shot Prompting?
Zero-shot, one-shot, and few-shot prompting are the three main flavors of in-context learning, distinguished by how many examples you give the model before asking it to perform the task. This framing comes directly from the GPT-3 paper, which evaluated the model under all three conditions to see how performance scaled with the number of demonstrations.
- Zero-shot: You give the model only an instruction, no examples at all. For instance: “Classify this review as positive or negative: ‘The battery life is amazing.'” The model relies entirely on what it learned during pretraining to infer what “classify” and “positive/negative” mean in context.
- One-shot: You give exactly one example before the real task, showing the model the input-output pattern you expect. For instance, showing one labeled review before asking it to label a new one.
- Few-shot: You give several examples (commonly anywhere from 2 to a few dozen, depending on the task and context window), which generally helps the model infer the intended format, tone, or edge cases more reliably than zero- or one-shot prompts.
The GPT-3 paper found that few-shot performance improved substantially with model scale, and that for many tasks, few-shot prompting closed much of the gap with fine-tuned models, though it didn’t always match state-of-the-art fine-tuned results. The practical lesson for prompt engineering is straightforward: when a zero-shot prompt gives you inconsistent or oddly formatted answers, adding one or two well-chosen examples is often the cheapest way to fix it. For a deeper comparison of when to reach for each approach, see Zero-Shot vs. Few-Shot Prompting: Which Should You Use and When?
How Can You Use In-Context Learning to Get Better AI Outputs?
Because in-context learning is just pattern-matching against what you show the model, the quality and structure of your prompt directly determines the quality of the output. A few practical techniques:
- Show, don’t just tell. Instead of writing a long paragraph describing the exact tone and format you want, provide one or two example input-output pairs in that exact tone and format. Models are often better at imitating a pattern than following an abstract description of it.
- Keep examples consistent. If your few-shot examples use inconsistent formatting, capitalization, or labeling schemes, the model will often pick up the inconsistency instead of the intended pattern.
- Order and diversity matter. Research on few-shot prompting has repeatedly found that the order of examples and how well they cover edge cases can shift results meaningfully, so it’s worth testing a couple of different example sets rather than assuming the first one you write is optimal.
- Combine with reasoning steps. For multi-step problems, showing an example that includes the reasoning process, not just the final answer, tends to produce more reliable outputs. This is the basis of chain-of-thought prompting, which is really a specialized application of in-context learning. See Chain-of-Thought Prompting Explained: How to Make AI Reasoning Visible for more on that technique.
- Mind the context window. Every example you add uses up tokens that could otherwise hold the actual task content, so few-shot prompting is a trade-off between guidance and available space, especially on tasks with long inputs.
Frequently Asked Questions
Is in-context learning the same thing as fine-tuning?
No. Fine-tuning changes a model’s weights through additional training and the effect persists across future sessions, while in-context learning only shapes behavior for the current prompt by giving the model instructions or examples as text. Once that context is removed, the model has no memory of it.
Does in-context learning work on every AI model?
It works best on large language models trained on massive, diverse text corpora, which is why it was first documented clearly in OpenAI’s GPT-3 research. Smaller or more narrowly trained models can show weaker in-context learning ability, since the GPT-3 paper itself found that few-shot performance scaled up notably as model size increased.
How many examples should I include in a few-shot prompt?
There’s no fixed number that works for every task. Many prompt engineers start with two to five well-chosen, diverse examples and adjust from there, since more examples can help the model generalize but also consume more of the context window and can occasionally reinforce an unwanted pattern if the examples aren’t representative.
In-context learning is one of the reasons prompt engineering became its own discipline: the same model can behave very differently depending purely on what you put in front of it. To go deeper on structuring prompts effectively, see Prompt Engineering From Scratch: The Definitive Guide to Mastering Any AI (The P3 Method). For the original research behind this post, read the GPT-3 paper, “Language Models are Few-Shot Learners” (Brown et al., 2020), and IBM’s explainer, “What is in-context learning?”
Featured photo by Ingo Dierking, licensed under CC BY-SA 4.0, via Wikimedia Commons.



