Understanding the KeepEdit LoRA Repository
When working with specialized image-editing checkpoints, knowing exactly what is and isn't included in a model repository prevents endless debugging sessions. The Hugging Face repository Yitaallen/keepedit-release-weights serves a specific purpose: it houses the published Low-Rank Adaptation (LoRA) weight files for the KeepEdit project.
It is crucial to understand right from the start that this repository contains solely the LoRA adapter weights. It does not package or distribute the Qwen-Image-Edit-2511 base model. Engineers and researchers pulling these weights must supply their own compatible base model when initializing inference pipelines or evaluation harnesses.
The Apache 2.0 licensed repository provides three distinct training checkpoints, each reflecting different stages or architectural variations of the KeepEdit training process. By separating the adapters from the multi-billion parameter base model, the repository keeps download sizes manageable while offering flexibility for developers experimenting with distinct editing behaviors.
Directory Structure and Checkpoint Variants
Navigating the repository reveals three distinct subdirectories, each containing a specific safetensors checkpoint file. These checkpoints represent different training milestones or architectural paradigms explored during the development of KeepEdit.
The directory layout is organized as follows:
qwen_edit_2511_keepedit_gt_onestage/step-4404.safetensorsqwen_edit_2511_mtp_phasea/step-2269.safetensorsqwen_edit_2511_moe_teacher_onestage/step-2202.safetensors
The first variant, qwen_edit_2511_keepedit_gt_onestage, corresponds to step 4404 of a ground-truth one-stage training setup. The second folder, qwen_edit_2511_mtp_phasea, contains weights up to step 2269 under a multi-task or multi-phase training configuration. The third variant, qwen_edit_2511_moe_teacher_onestage, captures step 2202 of a mixture-of-experts teacher guided one-stage training run.
Each checkpoint targets specific nuances in image editing fidelity. Having multiple options allows teams to benchmark which training strategy performs best on their specific editing tasks, whether they prioritize structural preservation, text-instruction adherence, or compositional balance.
Inference Conditions and Minimal Inputs
One of the most practical aspects of the KeepEdit release is the consistency of its inference requirements across all three checkpoints. Regardless of whether you load the ground-truth one-stage weights or the MoE teacher variant, the operational contract remains identical.
The inference pipeline expects a straightforward input pair:
- A source image
- A textual instruction describing the desired change
This combination maps directly to an edited output image. Crucially, inference does not require target images, spatial masks, auxiliary expert candidate graphs, or MoE teacher maps. Because the adaptation layers have learned to interpret edit instructions directly from the source pixels and prompt text, runtime overhead is significantly reduced.
This simplicity makes integration into existing pipelines much easier. You don't need complex preprocessing pipelines to generate segmentation masks or ground-truth targets just to test an edit. You feed in your source visual, write your instruction, and let the adapted base model handle the rest.
Downloading Checkpoints via Hugging Face CLI
Before running any evaluations or integration scripts, you need to pull the weight files to your local environment. The recommended approach utilizes the official Hugging Face command-line interface.
You can fetch the complete set of weights using the following CLI command:
huggingface-cli download Yitaallen/keepedit-release-weights \
--repo-type model \
--local-dir checkpoints \
--local-dir-use-symlinks False
Setting --repo-type model ensures the CLI looks in the model namespace, while --local-dir checkpoints places the files neatly inside a local directory. Disabling symlinks (--local-dir-use-symlinks False) ensures that actual weight files are written to disk, which is often necessary when mounting volumes in containerized training or evaluation environments.
Once downloaded, the folder structure inside your local checkpoints/ directory mirrors the remote repository, placing each safetensors file in its respective experiment folder.
Running Evaluation Scripts
The repository documentation outlines a standardized evaluation workflow using shell scripts provided in the main project repository. To evaluate each checkpoint properly, you configure specific environment variables before invoking the evaluation script.
Here is how you can trigger evaluations for each of the three variants:
For the ground-truth one-stage model:
EXPERIMENT_NAME=qwen2511_gt_onestage \
LORA_PATH=checkpoints/qwen_edit_2511_keepedit_gt_onestage \
bash scripts/evaluate_qwen_edit_experiment.sh
For the MTP phase A checkpoint:
EXPERIMENT_NAME=qwen2511_mtp_phasea \
LORA_PATH=checkpoints/qwen_edit_2511_mtp_phasea \
bash scripts/evaluate_qwen_edit_experiment.sh
For the MoE teacher one-stage variant:
EXPERIMENT_NAME=qwen2511_moe_teacher_onestage \
LORA_PATH=checkpoints/qwen_edit_2511_moe_teacher_onestage \
bash scripts/evaluate_qwen_edit_experiment.sh
By passing the appropriate EXPERIMENT_NAME and pointing LORA_PATH to the downloaded local directory, the evaluation harness binds the correct adapter weights to the execution environment.
Integrating with Diffusers Pipelines
For developers building custom applications or notebooks outside of the default evaluation scripts, Hugging Face Diffusers provides a native pathway for loading LoRA weights.
A typical Python loading snippet looks like this:
import torch
from diffusers import DiffusionPipeline
## Initialize the pipeline with the base model
pipe = DiffusionPipeline.from_pretrained(
"fill-in-base-model",
dtype=torch.bfloat16,
device_map="cuda"
)
## Load the KeepEdit release LoRA weights
pipe.load_lora_weights("Yitaallen/keepedit-release-weights")
prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image = pipe(prompt).images[0]
Remember to replace "fill-in-base-model" with the actual path or identifier of your Qwen-Image-Edit-2511 base model. Utilizing torch.bfloat16 and device_map="cuda" ensures efficient memory utilization on supported GPU hardware.
Practical Considerations and Best Practices
When deploying or experimenting with these checkpoints, keeping a few practical tips in mind will save you time. First, verify that your base model version matches what KeepEdit expects; mismatched base architectures will lead to degraded editing results or shape mismatches during weight loading.
Second, ensure your environment has sufficient VRAM. Running multi-billion parameter vision-language editing models in bfloat16 requires adequate GPU memory, especially when batching inference requests.
By understanding the exact scope of Yitaallen/keepedit-release-weights—that it is strictly a collection of LoRA adapters—you can seamlessly integrate these powerful editing checkpoints into your existing workflows without second-guessing missing base files.