When 70B Outclasses 90B
Parameter counts are a lazy proxy for capability, and a fresh evaluation of Meta's Llama family proves the point without even trying. A community benchmark published on the Hugging Face blog ran several open-weight models against medical and healthcare tasks, and the result was quietly embarrassing for the flagship: Llama-3.1-70B-Instruct, the previous generation's text model, edged past the newer, larger Llama-3.2-90B Vision model on specialized exams like MMLU College Biology and Professional Medicine. No fine-tuning was involved on either side. It was a straight walk-on, bare-model comparison, and the smaller model won anyway.
I find this more useful than most polished model-card marketing because it captures something real about how these systems behave in practice. The bigger checkpoint is not automatically the better tool for a domain task. Sometimes it is barely the same tool at all.
The Medical Benchmark Results
Let me put the numbers on the table, because the margins are small enough that you need to see them to appreciate the point. The evaluation, attributed to Ankit Pal and published in late September 2024, scored the models across datasets including MMLU College Biology, Professional Medicine, and PubMedQA.
Llama-3.1-70B-Instruct landed first overall with an average score of 84%. Its breakdown was strong and consistent: 95.14% on MMLU College Biology and 91.91% on MMLU Professional Medicine. The 90B vision model came in a hair behind, tied for second at an average of 83.95%, with 93.06% on College Biology and 91.18% on Professional Medicine. We are talking about fractions of a percentage point on the average, and a two-point gap on the biology question set where the older 70B simply answered better.
Third place went to the original Meta-Llama-3-70B-Instruct at an average of 82.24%, with a standout 93% on Medical Genetics and 90.28% on College Biology. So even the first-generation 70B held its own. The honest read is that the three largest models here are clustered within roughly two points of each other on text-only medical reasoning, and the largest among them is not the leader.
The smaller tier told its own story. Phi-3-4k led the compact models with an average of 68.93%, scoring 84.72% on College Biology and 75.85% on Clinical Knowledge. Meta-Llama-3.2-3B-Instruct followed at 64.15%, and the un-instructed Llama-3.2-3B trailed at 60.36%. Worth noting: the base 3B actually beat the instruct 3B on PubMedQA, 72.8% versus 70.6%, a reminder that instruction tuning can help with format while quietly doing nothing, or sometimes a little harm, for raw factual recall.
Instruct and Base, Score for Score
The finding that genuinely stopped me mid-scroll is that Meta-Llama-3.2-90B Vision Instruct and Meta-Llama-3.2-90B Vision Base produced identical results across every dataset. Identical. Not close. On PubMedQA both scored 72.8%, with no variation between them.
That is strange, and the article is right to flag it. Instruction-tuned models normally diverge from their base siblings on knowledge benchmarks, sometimes in your favor and sometimes the other way. The whole premise of an instruct checkpoint is that a downstream tuning pass reshapes how the model retrieves and presents what it knows. When two checkpoints behave as if they are the same model, one of two things is going on. Either the evaluation was too coarse to register the differences that instruction tuning does introduce, or the vision-oriented tuning on these checkpoints simply did not move text-knowledge behavior much at all.
My own lean is toward the second explanation, with a caveat. These are multimodal models first. Their training emphasis sits on wiring a vision encoder to a language backbone so they can see images. If the instruction component of that tuning spent its budget on image-grounded conversations, a plain text-only medical exam never pokes the parts of the model that changed. So the instruct and base variants look interchangeable on tasks they were not being shaped for. That is a hypothesis, not a proven fact from this evaluation, but it is the one the numbers most comfortably support.
What Changed in 3.2 Anyway
If the 90B does not win on medical text, why does it exist? Because Llama 3.2 Vision was never primarily about beating the old 70B at multiple-choice biology. The release shipped vision models in two sizes, an 11B aimed at consumer-size GPUs and the 90B aimed at large-scale applications, each offered in base and instruction-tuned flavors. The flagship's reason for being is multimodal: it takes an image and text as input and reasons across both, the architecture Hugging Face files under image-text-to-text. A new Llama Guard 3 with vision support shipped alongside it to screen inputs and outputs.
The smaller 1B and 3B siblings round out the family. They follow the same underlying architecture as Llama 3.1, were trained on up to nine trillion tokens, keep the 128k context window, and cover eight languages, with Meta pointing developers toward fine-tuning for broader language coverage. None of that architecture story shows up on a text medical exam, which is precisely the lesson. You would pick the vision 90B because you need it to read a chart, a radiograph, or a screenshot of a lab panel, not because it should out-score the 70B on a written test.
Choosing the Right Tool, Honestly
Strip away the model-card noise and the practical guidance is blunt. If your workload is text-only domain reasoning, the previous-generation 70B is a legitimate and sometimes better choice than the newer 90B vision model, and it likely costs you less to serve. Run your own eval before you believe any benchmark on a single domain, but do not let a larger parameter count talk you into a regression you could have caught in an afternoon.
Two habits come straight out of this. First, treat domain benchmarks as the only fair test of a model's competence in that domain, because aggregate leaderboards hide exactly these inversions. The 70B and 90B sit within noise of each other on the average and still trade blows on individual question sets, which means your specific task, not the family name, decides the winner. Competitions like the NeurIPS LLM fine-tuning competition have made the same point from the other direction: carefully tuned smaller models repeatedly beat much larger ones on the task that actually matters.
Second, be suspicious when an instruct and a base checkpoint agree perfectly on a task. It is either a sign your eval is too blunt to see the difference, or a sign the tuning pass did not touch the capability you are measuring. Both are worth investigating before you ship the instruct model and assume it behaves differently than the thing you started from.
Llama 3.2 is a genuinely good release, especially for anyone who needs open multimodal models that fit on real hardware. It just happens to be a worse medical text specialist than its own predecessor, and there is no shame in admitting that the headline parameter count and the useful capability are, this time, different things. That tension is exactly what makes open-weight AI models worth evaluating hands-on rather than ordering by spec sheet.
Source Note
The benchmark figures above come from the community evaluation Performance Comparison: Llama-3 Models in Medical and Healthcare AI Domains. Architecture, model sizes, and release details for Llama 3.2 are drawn from the Hugging Face Llama 3.2 release post and the Llama-3.2-90B-Vision-Instruct model card.