Reading an AI Model Card Before You Build on a New Model

Person reviewing a model card and technical documentation on a laptop

A model card is the document a model’s creator publishes alongside the model itself, describing what it is, what data trained it, how it scored on specific tests, and where it should not be used. You read one before building on a new model because it is the fastest way to find the gap between what the model can actually do and what your product needs it to do, before that gap shows up in production as a support ticket or a safety incident. This walkthrough reads one real, current model card from top to bottom, section by section, the way you should read it the first time you’re deciding whether to build on it.

The card we’ll use is Meta’s card for Llama 3.1 8B Instruct, published on Hugging Face at huggingface.co/meta-llama/Llama-3.1-8B-Instruct. It’s a good teaching example because it follows a fairly standard structure that most serious model cards converge on, and Hugging Face’s own documentation — the Model Cards guide and the more detailed Annotated Model Card — describes exactly why each section exists. We’ll walk the Llama card section by section, in the order it actually appears, and after each one ask: what does this tell a developer, and what question should you ask if this section is thin, vague, or missing entirely?

Step 1: Confirm you’re reading the card, not the launch post

Before reading anything, check that you’re on the model’s actual card page rather than a press release or a third-party summary. A launch blog post is written to make the model look good; the model card is the closer-to-source document that the provider is willing to put a specific claim on. On Hugging Face, the card is the page you land on when you open a model’s repository — it’s the same file (a README.md with YAML metadata at the top) that drives the model’s listing, license badge, and tags. If you can’t find a dedicated card for the model you’re evaluating — just a blog post and a demo — that absence is itself informative: a provider that skipped the card skipped the discipline of writing down intended use and limitations, and you should treat that model with more caution, not less.

Step 2: Read “Model Information” first, and nothing else yet

The first substantive section on the Llama card is Model Information, which states plainly: “The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes.” This section exists to answer one question before any other: which model, in which configuration, am I actually looking at? That matters because providers publish families, not single models — Llama 3.1 ships as pretrained and instruction-tuned, at three sizes. Hugging Face’s own guidance calls the equivalent section Model Details and says it “provides basic information about what the model is, its current status, and where it came from.” If this section is missing a version number, a parameter count, or a release date, stop and ask the provider’s docs directly: am I looking at the model that will still exist in six months, or a preview that may be renamed or deprecated?

Step 3: Intended Use Cases — the section that defines your legal and product footing

“Intended use” is the model creator’s statement of the applications they built and tested the model for; using the model outside that statement means you’re the one validating safety and performance, not them. The Llama card’s Intended Use Cases section says: “Llama 3.1 is intended for commercial and research use in multiple languages. Instruction tuned text only models are intended for assistant-like chat.” It also names a specific, finite list of supported languages and flags that pretrained models are meant to be adapted for generation tasks, while instruction-tuned models are meant for chat. If you’re building a customer support bot in a language not on that list, the card has already told you: this is an unsupported use, and any safety testing Meta did does not cover you. If an intended-use section only says something generic like “for research purposes,” the question to ask is: research by whom, tested against what, and did anyone validate this for the commercial use I have in mind? A vague intended-use statement usually means the provider hasn’t thought hard about misuse, which should raise your bar for your own testing, not lower it.

Step 4: Out-of-Scope Use — read this one word for word

Out-of-scope use is the provider’s explicit list of ways you are not supposed to use the model; it is often the shortest section on the card and the one teams skip fastest, which is a mistake. The Llama card states it plainly: “Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in any other way that is prohibited by the Acceptable Use Policy and Llama 3.1 Community License. Use in languages beyond those explicitly referenced as supported in this model card.” Notice what this does: it ties out-of-scope use directly to a separate legal document (the Acceptable Use Policy) rather than just listing examples. That means reading the model card alone is not enough — you need to pull the linked policy too. If a card has an intended-use section but no out-of-scope section at all, ask: has this provider actually thought about misuse, or did they only write down the upside? Treat a missing out-of-scope section as an unanswered risk question, not as permission to use the model however you want.

Step 5: Training Data — what the model has and hasn’t seen

The training data section tells you what information the model was built from and, critically, when that information stops. The Llama 3.1 card is specific here: “Llama 3.1 was pretrained on ~15 trillion tokens of data from publicly available sources. The pretraining data has a cutoff of December 2023.” That single sentence answers a question that otherwise causes real production bugs: if you ask the model about anything that happened after December 2023, it either doesn’t know or will guess, and your application needs to handle that — with retrieval, with a disclaimer, or by routing those queries elsewhere. Hugging Face’s annotated template calls for this section to include “1-2 sentences on training data with links to Dataset Cards,” which is a low bar — most providers, Meta included here, don’t name the actual datasets, only that they’re “publicly available.” If a card gives you a cutoff date but no sense of domain coverage (is this mostly web text? code? one language?), the question to ask is: how do I find out what the model has and hasn’t seen on my specific domain, since the card won’t tell me and I’ll need to test it myself.

Step 6: Benchmark scores — numbers that only mean something in context

An evaluation benchmark is a fixed, standardized test with a scoring method, run the same way across many models so their results can be compared; a benchmark score tells you how a model did on that one test, not how it will do on your task. The Llama 3.1 8B Instruct card publishes a results table, and the real numbers are worth looking at directly:

Benchmark Score
MMLU 69.4
MMLU (CoT) 73.0
MMLU-Pro (CoT) 48.3
IFEval 80.4
ARC-C 83.4
GPQA 30.4
HumanEval 72.6
MBPP++ 72.8
GSM-8K (CoT) 84.5
MATH (CoT) 51.9

Read this table by asking what each benchmark actually tests, not just whether the number looks high. MMLU and MMLU-Pro test broad academic knowledge across subjects; GSM-8K and MATH test grade-school and competition-level math reasoning; HumanEval and MBPP++ test whether generated code actually runs and passes test cases; IFEval tests whether the model follows explicit formatting instructions; GPQA is a graduate-level science question set designed to resist lookup. A GPQA score of 30.4 next to an ARC-C score of 83.4 tells you this model is meaningfully stronger on grade-school-level science reasoning (ARC-C) than on graduate-level, resistant-to-memorization questions (GPQA) — which matters if your use case leans toward the latter. If a model card shows benchmark scores with no comparison baseline and no description of how the test was run (zero-shot? few-shot? chain-of-thought prompted, as the “(CoT)” labels here indicate?), the question to ask is: under what prompting conditions was this number produced, because the same model can score very differently depending on how it was asked.

Step 7: Responsibility & Safety — what was actually tested for harm

This section tells you which specific harms the provider went looking for, as opposed to harms they merely gestured at. The Llama 3.1 card states that red teaming — structured, adversarial testing where people deliberately try to make a model produce harmful output — covered three named risk categories: CBRNE uplift (whether the model could help someone plan chemical, biological, radiological, nuclear, or explosive weapons), child safety (expert evaluation of outputs across multiple attack vectors), and cyber attack enablement (whether the model could enhance hacking tasks or autonomous ransomware). The card is also explicit about where responsibility sits after release: “Llama is a foundational technology designed to be used in a variety of use cases. Developers are then in the driver seat to tailor safety for their use case.” That sentence is doing real work — it’s telling you, directly, that passing these three red-team categories does not mean the model is safe for your specific application; it means the provider checked for some specific, serious harms and is handing the rest of the job to you. If a card’s safety section lists broad principles (“we care about safety”) but no named test categories, the question to ask is: what did they actually test, versus what do they merely claim to value?

Step 8: Ethical Considerations, Limitations, and License — the fine print that isn’t optional

The limitations section is where a provider admits, in writing, what the model still gets wrong. Meta’s language here is standard but still useful to read literally: the model “may in some instances produce inaccurate, biased or other objectionable responses,” and developers are told to “perform safety testing and tuning tailored to their specific applications” before deployment. That’s not boilerplate to skip — it’s the provider formally declining to warrant the model’s output for your use case, which should change how much of your own evaluation budget you allocate before shipping. The license section matters just as much for a build decision: Llama 3.1 ships under the Llama 3.1 Community License, which requires that products built with it display “Built with Llama” and that any model name derived from it start with “Llama.” Those aren’t just legal footnotes — they’re product and branding requirements that need to reach your legal and marketing teams before launch, not after. If a license section is missing or just says “open source” with no named license, ask: which specific license, because “open source” covers dozens of licenses with very different obligations around attribution, commercial use, and redistribution.

Putting the walkthrough together

Read in this order — Model Information, Intended Use Cases, Out-of-Scope Use, Training Data, Benchmark Scores, Responsibility & Safety, Limitations, License — a model card stops being a wall of text and becomes a short risk assessment you can act on. For Llama 3.1 8B Instruct specifically, that assessment is: a chat-tuned model, licensed under Meta’s community terms with attribution requirements, trained on data through December 2023, strong on grade-school math and code generation, markedly weaker on graduate-level science reasoning, red-teamed for three named catastrophic-risk categories but explicitly not warranted for your particular application. None of that required testing the model — it required reading the card the provider already published, and knowing which section answers which question.

FAQ

What if the model I want to use has no model card at all?
Treat that as a finding, not a dead end: check the model’s repository README, any linked paper, and the provider’s API documentation for the same information (intended use, training data cutoff, known limitations), and budget extra time for your own evaluation since you can’t rely on the provider’s.

Is a “system card” the same thing as a model card?
No — a system card (the term Anthropic and OpenAI use for documents describing models like Claude or GPT) typically covers the deployed system as a whole, including safety evaluations and policy context, while a model card as Hugging Face defines it focuses more narrowly on the model artifact itself; in practice the two documents answer overlapping questions and are worth reading together when both exist.

Do I need to re-read the card if the provider updates the model?
Yes — providers revise model cards when they retrain, re-evaluate, or change licensing terms, and a card you read six months ago may no longer reflect the current benchmark scores or use restrictions, so re-check it before a major release that depends on the model.

Once you’ve read the card, the next questions are usually about fit: how this model compares to an open-weight alternative, how to go about evaluating it against your own benchmarks rather than the provider’s, how to write a system prompt that respects what the card told you, and where a guardrail layer needs to pick up where the model’s own safety testing stopped.

Image: “Software Developer at work 01” by Daudi mukiibi, licensed under CC BY-SA 4.0, via Wikimedia Commons.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top