There are three practical ways to get a large language model to produce accurate output in a language other than English: write the prompt itself in the target language, write the prompt in English but explicitly instruct the model to respond in the target language, or run a two-step pipeline that generates the content in English first and translates it afterward with a separate machine-translation or AI-translation call. They differ mainly in where linguistic and cultural judgment happens — inside one model pass, split across the instruction and the output of that same pass, or handed off entirely to a dedicated translation step — and that difference drives everything else: accuracy, cultural fit, cost, latency, and how much control you have over terminology.
Localization vs. Translation, and What “Low-Resource” Means
Translation converts words and sentences from one language into another while preserving their literal meaning; localization goes further, adapting tone, idiom, units, examples, and cultural references so the result reads as if it were written natively for that market rather than converted into it. A prompt that asks for a “translation” and a prompt that asks for “localized” output can produce meaningfully different text for the same source sentence, because localization permits rephrasing that a strict translation would not. The three approaches compared here can each be pushed toward either goal, but they default differently: native-language prompting and well-instructed English prompts tend toward localization, while a raw two-step translation pipeline tends toward literal translation unless you explicitly instruct otherwise.
“Low-resource language” refers to a language for which comparatively little digital text exists for a model to train on — think Yoruba or Swahili next to English, Spanish, or Mandarin. Because large language models learn primarily from the text they are trained on, a language with a smaller training footprint tends to produce weaker, less fluent, or less reliable output, even when the model performs well in high-resource languages. This gap is documented, not anecdotal, and it matters for choosing between the three methods below.
Method 1: Write the Prompt Itself in the Target Language
Writing the prompt in the target language means the model reads the instructions, reasons, and drafts its answer in that language from start to finish, with no translation step anywhere in the process. Anthropic’s own documentation on multilingual support recommends this kind of explicitness: “explicitly stating the desired input and output languages improves reliability,” and it specifically advises submitting “text in its native script rather than transliteration for optimal results.” The same documentation suggests prompting the model to use “idiomatic speech as if it were a native speaker” for better fluency — a nudge that native-language prompting makes natural to include, since the whole prompt is already framed in-language.
The upside is quality: because every instruction, example, and constraint is phrased in the target language, the model has no English scaffolding to drift back toward, and idiomatic register (formality level, regional word choice, courtesy conventions) tends to come through more naturally. The downside is that this method’s quality ceiling is set by how well-represented that language is in the model’s training data. Anthropic’s published benchmark table for Claude models — zero-shot chain-of-thought scores on a translated MMLU test set, expressed relative to English performance fixed at 100% — shows Spanish, French, and Chinese landing in the mid-to-high 90s, while Swahili drops to roughly 91% (Sonnet) and 78% (Haiku), and Yoruba falls further still, to about 80% (Sonnet) and 53% (Haiku). Writing your entire prompt in a lower-resource language inherits that gap directly: there is no English-language fallback inside the same call to compensate.
There’s also a practical authoring cost: whoever writes and maintains the prompt needs real fluency in the target language, including register and idiom, or they cannot reliably judge whether the model followed instructions correctly. For a single flagship market this is manageable; for an application that needs to support fifteen languages, maintaining fifteen separately authored, separately reviewed prompts becomes an operational burden of its own, independent of model quality.
Method 2: Write the Prompt in English, Instruct the Output Language
Writing the prompt in English and instructing the model to answer in another language keeps your instructions, logic, and test cases in one language while only the final response shifts to the target language. Both Anthropic and OpenAI document this pattern directly. Anthropic’s guidance states that “the most reliable place to do this is the system prompt, which keeps the instruction stable across every turn of a conversation,” with an example as simple as: system: "Always respond in French, regardless of the language the user writes in." OpenAI’s help documentation describes the default behavior this method has to override: write a prompt in Spanish and “you’re likely to receive a response in Spanish,” and more generally, the model tends to answer in whatever language the input uses — so when you want English instructions but non-English output, an explicit instruction is what makes that split happen reliably rather than by chance.
The practical advantage is maintainability. One English-language prompt template, versioned and tested in English, can be reused across many target languages just by swapping the “respond in X” instruction (or interpolating a user-selected language, which both vendors recommend over letting the model infer it). That’s a large win for support bots, UI copy generators, and any system that needs to scale across markets from a single source of truth. The trade-off is that you’re relying on the model’s multilingual generation quality under an instruction rather than a fully in-language framing, and that quality still tracks the same resource gap described above — English-prompted output in Yoruba or Swahili will reflect the same underlying weaker multilingual capability as native-language prompting would, because the instruction changes which language the output appears in, not how well the model performs in it.
Method 3: Generate in English, Then Translate as a Separate Step
A two-step pipeline generates the full response in English first, then passes that English text through a dedicated translation step — a machine-translation engine, a separate LLM call whose only job is translation, or a human-in-the-loop translation workflow — before anything reaches the end user. This decouples content generation from language entirely: the model that writes the answer never has to be good at the target language, only the translation layer does.
This separation is the method’s biggest strength for terminology control. Dedicated translation workflows can plug into a maintained glossary or termbase (a standard tool in professional localization: a fixed list mapping source terms to approved target-language equivalents) and translation memory (a database of previously approved translations, reused for consistency), so that a product name, a legal term, or a brand phrase is rendered the same way every time, across every piece of content that passes through the pipeline. Neither native-language prompting nor instructed-output prompting gives you that same mechanical guarantee, because each generation call is reasoning about terminology fresh, with no shared memory across calls unless you build and inject that glossary into every prompt yourself.
The costs are real, though. You’re paying for two model calls (or a model call plus a translation-API call) instead of one, which adds both latency and cost per request. There’s also a second place for errors to enter: the English draft can be fine but the translation step can introduce its own mistakes, or vice versa, and a flaw in the English source propagates into every language you translate it into. Cultural adaptation tends to be the weakest point here unless the translation step is explicitly instructed to localize rather than translate, because a translation pass working from finished English text has less flexibility to restructure a sentence, drop an English-centric idiom, or adjust the example used, than a model generating fresh content directly in the target language would.
Comparing the Three Approaches
The table below compares all three methods on the four dimensions that matter most when choosing between them: how faithfully the output preserves meaning, how well it reads as native (not translated) text, what it costs in time and money, and how much control you have over consistent terminology.
| Dimension | Method 1: Native-language prompt | Method 2: English prompt, instructed output language | Method 3: Generate in English, then translate |
|---|---|---|---|
| Accuracy / fidelity | High for well-resourced languages; degrades in lower-resource languages, since quality depends entirely on that language’s representation in training data. | Similar ceiling to Method 1 for the same model and language — the instruction changes the output language, not the model’s underlying multilingual competence. | Fidelity to the English source depends on translation quality; a second error source is introduced, but errors don’t compound with “weak target-language generation” since the model only ever wrote English. |
| Cultural adaptation | Strongest by default: idiom, register, and local convention are generated natively, with no literal-translation artifacts to correct. | Generally strong, especially with an added instruction to use idiomatic, native-sounding phrasing; still prompt-dependent. | Weakest by default: a translation pass working from finished English text tends toward literal rendering unless explicitly instructed to localize rather than translate. |
| Cost / latency | One model call, same as any single-pass prompt; cost can still rise because some languages are tokenized less efficiently than English, which inflates token counts for the same content. | One model call — the cheapest and fastest of the three in terms of API calls, with the same tokenization caveat as Method 1. | Highest cost and latency: two calls (generation plus translation) per request, run sequentially in most implementations. |
| Control over terminology | Limited and manual: consistent terminology across many prompts or sessions requires the prompt author to maintain and re-paste a glossary themselves. | Same limitation as Method 1; English-language authoring makes it easier to centrally maintain one glossary, but it must still be injected per call. | Strongest: the translation step can be built around a persistent glossary and translation memory, giving mechanical consistency across every request without re-specifying it each time. |
A Worked Example: Same Instruction, Two Phrasings
To make the difference between native-language prompting and English-prompt-with-instruction concrete, here is one instruction expressed both ways, with the kind of divergence you should expect to see between them. This is an illustrative example meant to show why phrasing and terminology differences are plausible, not a benchmarked test result — actual output varies by model, model version, and sampling settings, and running this exact prompt may well produce very similar results in both forms.
Approach A — prompt written natively in Spanish (Method 1):
Eres un agente de atención al cliente de una tienda de ropa en línea.
Escribe un correo breve, cálido pero profesional, para una clienta cuyo
pedido #4521 se retrasó cinco días. Discúlpate, explica que el pedido ya
va en camino, y ofrécele el código de descuento AGRADECIDOS10 para su
próxima compra. Usa un tono cercano pero respetuoso (trato de "usted"),
como lo haría un negocio local dirigiéndose a una clienta habitual.
Approach B — prompt written in English, instructing Spanish output (Method 2):
You are a customer service agent for an online clothing store. Write a
short, warm but professional email to a customer whose order #4521 was
delayed by five days. Apologize, explain that the order is now on its
way, and offer the discount code AGRADECIDOS10 for their next purchase.
Respond entirely in Spanish, using a respectful but friendly tone
(formal "usted").
Both prompts ask for the same content and the same formal register. The divergence to expect is not in the big structural pieces — both should apologize, confirm shipping, and offer the code — but in smaller phrasing choices that reveal which language the instructions were actually reasoned in. Approach A, framed entirely in Spanish, is more likely to reach naturally for a customer-service formula that a Spanish-speaking copywriter would reach for first, such as “Lamentamos mucho la demora” or “quedamos a su disposición” as a closing courtesy phrase. Approach B, built from an English brief, is more likely to produce a correct but slightly more literal rendering of the English instruction’s own phrasing — for example, translating “the order is now on its way” quite directly as “su pedido ya está en camino” (perfectly correct, but a direct mirror of the English clause order) rather than a construction a native Spanish-language support email might lead with instead, such as “¡Buenas noticias! Su pedido ya salió.”
The reason this gap is plausible rather than guaranteed is straightforward: in Approach A, every example, instruction, and constraint the model is weighing is already phrased in the target language, so the register and idiom it draws on when generating are pulled from the same linguistic context as the output. In Approach B, the model is translating a task description that was conceived in English into a Spanish-language response, and while modern models handle this well for a high-resource language like Spanish, the underlying instruction still originates in English phrasing, which makes literal, clause-by-clause carryover slightly more likely at the margins — particularly for connective phrases and courtesy formulas that don’t map one-to-one between languages. This is exactly the effect Anthropic’s own best-practice guidance is aimed at when it recommends explicitly prompting a model to use “idiomatic speech as if it were a native speaker”: that instruction is more necessary, and more effective, specifically when the brief and the output aren’t already in the same language.
Which Method to Use When
Choosing between the three comes down to what you’re optimizing for in a given project, not which method is universally “best.”
- Use native-language prompting when you have genuine fluency available to write and review prompts in the target language, you’re serving a single flagship market where idiomatic quality matters most, and the language in question is reasonably well-resourced so the model’s underlying capability supports it.
- Use an English prompt with an instructed output language when you need one template to scale across many target languages with centralized testing and versioning — the common case for support bots, onboarding flows, and templated UI copy — and you can accept the model’s native multilingual generation quality for each language you support.
- Use a two-step generate-then-translate pipeline when terminology consistency and auditability outweigh raw fluency: regulated industries, legal or medical content, brand-controlled messaging, or any workflow that already has a localization team, glossary, and translation-memory system in place that you want to plug into rather than bypass.
In practice, many teams end up combining elements of all three: a two-step pipeline for anything terminology-sensitive, instructed-output prompting for high-volume, lower-stakes generation across many languages, and native-language prompting reserved for a short list of markets where a human reviewer can validate the result.
FAQ
Can these three methods be combined rather than chosen between?
Yes. A common hybrid is generating in English, then using an in-language review or “polish” pass — rather than a literal translation pass — so a native-language-capable step still adapts tone and idiom instead of just converting words, while terminology from a glossary is still enforced centrally.
Does adding in-language reference material or examples change this comparison?
It helps any of the three approaches, but it’s a separate lever from which method you choose. Supplying a short glossary, a style sample, or a few approved example sentences in the target language as context improves terminology consistency and tone whether the surrounding prompt is written natively, written in English with an output instruction, or feeding a translation step — it doesn’t substitute for the structural choice of where translation happens.
Whichever method you pick, it builds on the same prompt engineering fundamentals that apply to any language, and on context engineering once a glossary or reference material enters the prompt. If you’re standardizing this across a product, it’s also worth revisiting what makes a system prompt hold up reliably across turns, since that’s exactly where Method 2’s language instruction needs to live.
Image: “Bilingual sign on the Welsh side of Boundary Lane, Saltney” by Rept0n1x, licensed under CC BY-SA 3.0, via Wikimedia Commons.



