ProBackend
open source ai models
1 hour ago6 min read

Storm in a Small Model: What Llama 3.1 Storm 8B Changes

A sourced look at the Storm fine-tune's self-curation, Spectrum fine-tuning and SLERP model-merging recipe, its reported benchmark gains over Llama-3.1-8B-Instruct and Hermes-3, downloadable variants, and the caveats around interpreting its results.

What Llama 3.1 Storm 8B Is

Llama 3.1 Storm 8B (published as akjindal53244/Llama-3.1-Storm-8B on Hugging Face) is a fine-tuned derivative of Meta's Llama-3.1-8B-Instruct, built by Ashvini Kumar Jindal, Pawan Kumar Rajpoot, Ankur Parikh and Akshita Sukhlecha. The model is described in its collection and model card as a generalist conversational model aimed at three improvements over its base: stronger reasoning, better conversation and instruction following, and more reliable function calling. The authors frame it squarely as an efficiency result for the small-language-model (SLM) class — the kind of model that runs on a single consumer GPU, the same niche explored by compact releases such as Liquid AI's LFM2.5-2.6B — rather than as a frontier-scale system.

That framing is not incidental. The work descends from the team's NeurIPS LLM Efficiency Challenge 2023 entry, which won first prize by curating roughly 200K high-quality examples from a pool of about 5 million and training a 7B instruct model (the "Birbal" model) on a commodity GPU within 24 hours. Storm carries that same data-quality-first philosophy into the Llama 3.1 generation and scales the curation step up substantially. The model is licensed under the Llama 3.1 license, tagged with tooling such as mergekit and axolotl, marked as supporting eight languages, and linked to a set of arXiv papers from the team.

The Recipe: Self-Curation, Spectrum Fine-Tuning, and Merging

The authors describe a three-step pipeline. Each step is the part of the story that matters most for understanding what Storm actually changes relative to its base model.

Self-Curation

First, the team applied two self-curation methods to select about one million high-quality training examples from a pool of roughly 2.8 million open-source examples. The curation criteria focused on educational value and difficulty level. The notable design choice is who does the judging: instead of leaning on a much larger model (a 70B or 405B annotator) to score the data, they used the small model itself for annotation. This is the same "self-curation" idea that drove their earlier prize-winning work, and it is meant to keep the whole data-selection process cheap and reproducible without a giant teacher model.

Targeted Supervised Fine-Tuning (Spectrum)

Second, rather than fine-tuning every parameter, the team used Spectrum-based targeted fine-tuning over Llama-3.1-8B-Instruct. Spectrum accelerates training by selectively targeting the layer modules with the best signal-to-noise ratio (SNR) and freezing the rest. In the Storm recipe, about 50% of the layers are frozen, so the fine-tune concentrates updates where the model appears most receptive to them. This targeted approach is intended to reduce compute and overfitting risk while still moving benchmark behaviour.

Model Merging with SLERP

Third, the fine-tuned model was merged with the Llama-Spark model using SLERP (spherical linear interpolation). Merging produces a blended model whose weights are smoothly interpolated between the two parent checkpoints, so the resulting Storm weights carry characteristics from both the new fine-tune and Llama-Spark. This step is important to remember when interpreting results: Storm is not purely the product of the curated fine-tune — part of its behaviour is inherited from a separate merged model.

Reported Benchmark Gains

The authors report that Storm improves over Llama-3.1-8B-Instruct across ten diverse benchmarks covering instruction-following, knowledge-driven QA, reasoning, truthful answer generation, and function calling. The headline strengths, as published in the model card, are reproduced below as the authors' own reported absolute gains over Llama-3.1-8B-Instruct — they are not independent validations.

Strength areaReported absolute gain vs Llama-3.1-8B-Instruct
Instruction followingIFEval strict +3.93%
Knowledge-driven QAGPQA +7.21%, MMLU-Pro +0.55%, AGIEval +3.77%
ReasoningARC-C +3.92%, MuSR +2.77%, BBH +1.67%, AGIEval +3.77%
Agentic / function callingBFCL overall accuracy +7.92%, BFCL AST summary +12.32%
Reduced hallucinationsTruthfulQA +9%

The most striking entries sit in function calling — the BFCL AST-summary gain of +12.32% and overall BFCL accuracy gain of +7.92%, which lines up with the model's stated goal of better agentic behaviour. The reasoning and QA gains are real in the reported numbers but smaller in absolute terms, which is what one should expect once a base instruct model is already reasonably strong on those suites.

Comparison With Hermes-3-Llama-3.1-8B

The team also benchmarked Storm against Hermes-3-Llama-3.1-8B, another fine-tune built on the same Llama-3.1-8B-Instruct base. Their reported result: Storm outperforms Hermes-3 on 7 out of 9 benchmarks. Two honest exceptions are retained in the source. Hermes-3 surpasses Storm on the MuSR reasoning benchmark, and the two models show comparable performance on BBH. Presenting the comparison this way matters, a fine-tune that beats a strong peer on most but not all tests is a more credible claim than a clean sweep, and the MuSR/BBH exceptions should be carried forward rather than dropped in any summary.

Variants and Practical Deployment

The collection publishes the model in several formats so it can be dropped into different stacks:

  • BF16, Llama-3.1-Storm-8B, the original checkpoint.
  • FP8-Dynamic, Llama-3.1-Storm-8B-FP8-Dynamic, for lower-precision serving; the broader economics and tradeoffs of pre-optimized quantized checkpoints are covered in our deep dive into pre-optimized model quantization.
  • GGUF, Llama-3.1-Storm-8B-GGUF, for llama.cpp, Ollama, LM Studio and other local apps.
  • Ollama, pulled with ollama run ajindal/llama3.1-storm:8b.

The Hugging Face transformers library loads the model in bfloat16 by default, which matches the checkpoint type and is the recommended way to run it for best results; the installation notes call for transformers>=4.43.2. Conversational use is shown both through the high-level pipeline() API and through model.generate() with the Llama 3.1 chat template applied explicitly. For serving, the model card documents vllm serve, SGLang's launch_server, and docker model run hf.co/akjindal53244/Llama-3.1-Storm-8B, all exposing an OpenAI-compatible chat-completions endpoint. The practical takeaway is that the model ships ready for the standard 2024-era open-model toolchain: a Transformers/vLLM path for GPU servers and a GGUF/Ollama path for laptops.

How to Read These Results

Three caveats keep the reported numbers in proportion.

The gains are self-reported. Every percentage in the strengths table comes from the model's own authors, evaluated largely on the Open LLM Leaderboard suite, and labelled explicitly as absolute gains over Llama-3.1-8B-Instruct. They are useful evidence that the recipe moved the model in the intended direction, but they are not third-party validation, and the model card's own "Eval Results (legacy)" tag signals that leaderboard methodology has shifted since. Benchmark claims within the Llama family deserve this kind of scrutiny in general: as our look at Llama 3.1 70B versus the 90B vision model shows, headline numbers can invert under a different evaluation task.

Part of the behaviour is merged in. Because Storm is a SLERP merge of the curated fine-tune with Llama-Spark, its strengths are an interpolation of two parent models, not a clean readout of the curation-and-fine-tuning step alone. That is a legitimate technique, but it means "the self-curation caused gain X" is a looser claim than the three-step recipe summary implies.

Family traits are not Storm-specific proof. Storm inherits the Llama 3.1 8B family's characteristics, the architecture, the Llama 3.1 license, multilingual coverage, and the broader base-model capabilities. Those are inherited properties of the underlying family rather than evidence that Storm's fine-tune itself improved a given dimension. When assessing what Storm "changes," it is worth separating inherited base-model behaviour from the benchmark deltas the team actually attributes to its work.

Bottom Line

Llama 3.1 Storm 8B is a compact, efficiency-driven fine-tune whose value proposition is a specific recipe, self-curation with the small model as its own annotator, Spectrum targeted fine-tuning with half the layers frozen, and a SLERP merge with Llama-Spark, reported to lift an already-strong instruct base across reasoning, conversation and especially function calling, while remaining a single-GPU model. The reported wins are concentrated in agentic benchmarks, the headline comparison against Hermes-3 is 7-of-9 with two named exceptions, and the model ships in BF16, FP8-Dynamic and GGUF formats for the standard serving stack. Treat the gains as the authors' reported evaluations and the recipe as the genuinely interesting contribution.

Sources

llama storm 8b

More blogs