ProBackend
open source ai models
3 hours ago5 min read

Migrating PyTorch Model Weights to Safetensors for Enhanced Security and Speed

A comprehensive, human-written guide on migrating PyTorch model weights from legacy pickle .bin files to secure, high-speed safetensors format using Hugging Face tools and local scripts.

The Hidden Danger in Standard PyTorch Checkpoints

For years, saving PyTorch model weights meant reaching for .bin files powered by Python's pickle utility. It was convenient, built right into the ecosystem, and completely standard. But convenience has a dark side. Pickle files are essentially serialized Python object streams capable of executing arbitrary code upon loading. If you download a malicious checkpoint or compromise a storage bucket, an attacker can execute arbitrary commands on your training cluster or inference servers just by loading a model weight.

That architectural vulnerability isn't theoretical. In production environments handling untrusted model artifacts from community hubs, loading raw pickle files represents a massive security surface area. Industry leaders realized we needed a better container—one that stores raw tensor data without executing code during deserialization.

What Makes Safetensors Different and Fast

Enter safetensors, an open-source tensor serialization format designed from the ground up for safety and speed. Instead of wrapping complex Python object graphs, safetensors stores a strict JSON header specifying tensor metadata (such as data types and exact byte offsets) followed immediately by contiguous raw binary data blocks.

Because deserialization requires no Python object instantiation or bytecode execution, loading models is fundamentally secure. Furthermore, safetensors achieves exceptional performance through zero-copy memory mapping (mmap). When multiple worker processes or GPUs load the same weight file, they can map the underlying physical memory pages directly without duplicating tensors in RAM, drastically cutting down startup times and memory overhead in large-scale inference deployments.

Installation and Getting Started

Getting started with safetensors in your Python environment is straightforward. You can install the package via standard pip or conda depending on your workflow infrastructure:

pip install safetensors

Or via conda-forge:

conda install -c conda-forge safetensors

Once installed, interacting with .safetensors files is designed to feel familiar to PyTorch developers while providing robust utility functions for both saving and selective loading.

Programmatic Usage: Saving and Loading Tensors

Migrating or working with .safetensors programmatically in PyTorch is handled seamlessly via safetensors.torch. You can easily serialize dictionary mappings of tensors using save_file:

import torch
from safetensors.torch import save_file

tensors = {
    "embedding": torch.zeros((2, 2)),
    "attention": torch.zeros((2, 3))
}
save_file(tensors, "model.safetensors")

When loading weights back into memory—especially across multi-GPU nodes or memory-constrained instances—using safe_open allows you to inspect and load tensors efficiently:

from safetensors import safe_open

tensors = {}
with safe_open("model.safetensors", framework="pt", device=0) as f:
    for k in f.keys():
        tensors[k] = f.get_tensor(k)

Loading Tensor Slices for Multi-GPU Environments

One of the standout features of safe_open is the ability to inspect shapes and load precise slices of tensors without reading the entire tensor into host memory first. This is particularly advantageous when initializing sharded models across multiple GPUs:

from safetensors import safe_open

tensors = {}
with safe_open("model.safetensors", framework="pt", device=0) as f:
    tensor_slice = f.get_slice("embedding")
    vocab_size, hidden_dim = tensor_slice.get_shape()
    # Load only a specific slice of the embedding tensor
    tensor = tensor_slice[:, :hidden_dim]

If you are choosing between multiple weight files in a repository before loading, a worked example of inspecting and selecting the right checkpoint is covered in our QuiltNet-B-32 weights file guide, which walks through OpenCLIP and Hugging Face loading paths.

Native Hugging Face Integration (transformers and diffusers)

Beyond low-level utility functions, modern Hugging Face libraries like transformers and diffusers feature native, out-of-the-box support for safetensors. When downloading or pushing models to the Hub, safetensors is now the default serialization format for supported architectures.

When saving a trained model using the Transformers library, you can enforce safetensors serialization directly through save_pretrained:

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("gpt2")
tokenizer = AutoTokenizer.from_pretrained("gpt2")

## Save model weights in secure safetensors format

model.save_pretrained("./my-gpt2-safetensors", safe_serialization=True)

Similarly, when loading weights into memory, Hugging Face automatically detects .safetensors files in the model repository. You can also explicitly enforce or verify safetensors usage:

model = AutoModelForCausalLM.from_pretrained(
    "./my-gpt2-safetensors", 
    use_safetensors=True
)

Methods for Converting Legacy PyTorch Weights

PyTorch model weights are commonly saved and stored as .bin files with Python’s pickle utility. To save and store your model weights in the more secure safetensors format, we recommend converting your existing .bin checkpoints to .safetensors.

Depending on your model's hosting and scale, you have two primary avenues for conversion:

1. Using the Hugging Face Convert Space

The easiest way to convert your model weights if they are already hosted on the Hugging Face Hub is to use the official Convert Space. The Convert Space automatically downloads the legacy pickled weights, converts them into secure .safetensors, and opens a Pull Request to upload the newly converted files directly to your model repository.

Note on scale: For extremely large language models or massive multi-gigabyte checkpoints, the hosted Space may experience queue delays because its computing resources are shared across community model conversions.

2. Running the Local Conversion Script

If you are working with private weights, air-gapped clusters, or want to bypass public queue times for large models, you can run the conversion utility locally. Running conversion locally gives you full control over memory limits, temporary storage, and output destination paths.

Handling Sharded Checkpoints and Large Models

For massive models containing tens or hundreds of billions of parameters, single-file .safetensors checkpoints become impractical due to file size limits and memory constraints. In these scenarios, libraries utilize sharded checkpoints accompanied by an index file named model.safetensors.index.json.

The index maps individual parameter names to their respective shard files (e.g., model-00001-of-00003.safetensors). When loading sharded models, transformers reads the index file and dynamically loads only the required shards into memory, preserving system stability and preventing out-of-memory (OOM) errors during initialization.

Ecosystem Adoption and Production Readiness

The shift toward safetensors has become an industry standard across major open-source AI platforms. Leading projects and organizations—including Hugging Face (transformers, diffusers), EleutherAI, StabilityAI (stable-diffusion-webui), ComfyUI, MLX (ml-explore), and many others—have embraced safetensors as the default distribution format for state-of-the-art models.

Securing weights is only half the longevity story: archiving and preserving master weights so models stay reproducible is equally important, a point argued in detail in Keeping Open AI Models Usable: The Case for Preserving the Master Weights.

By migrating legacy PyTorch .bin checkpoints to safetensors, engineering teams eliminate critical remote code execution vulnerabilities while simultaneously accelerating model loading times and optimizing memory utilization across production clusters.

the hidden danger in standard pytorch checkpoints

More blogs