In building LLM applications, enterprises end up writing monstrous system prompts. They stuff them with company knowledge, internal preferences, tone rules, edge cases, application-specific instructions. The instinct is sound: in-context learning lets you steer a model at inference time without touching its weights. But that knowledge is transient. It does not survive a new conversation. So you paste the whole wall of text in front of the model on every single request, forever.
Do that at enterprise scale and two things break. Latency creeps past the thresholds your users will tolerate, and per-query cost climbs with every token you re-feed. This is the wall a lot of teams are quietly hitting in 2026.
Microsoft researchers put forward a technique that takes a different route. On-Policy Context Distillation (OPCD) bakes the knowledge and preferences baked into the prompt directly into the model's parameters. It does this by training on the model's own responses rather than on a teacher's polished answers. The result, per the published experiments, is a model that gets measurably better at the bespoke task without going stupid everywhere else.
The transience tax on in-context knowledge
In-context learning is cheap to set up and impossible to keep. You never have to retrain the base model, which is the slow, expensive path. The trade-off is that the knowledge lives only as long as the context window holds it. A support bot needs the company policy on every ticket. A clinical assistant needs the dense manual on every query. The model never learns it; it just re-reads it.
That repeated feeding is what drives the cost and latency curve upward. Long prompts are not just inelegant, they are a running bill. Microsoft's framing is that OPCD lets you pay that bill once, in a training run, and stop paying it per query.
On-policy distillation, in plain terms
The "on-policy" part is the whole trick, and it is worth understanding because it explains why this works when naive fine-tuning doesn't.
Traditional off-policy distillation takes demonstrations from a strong teacher and trains the student to imitate them token by token. The student learns to predict the next token given a perfect prefix. At inference time there are no perfect prefixes — the model is conditioning on its own imperfect generations. Errors compound, especially once you get reasoning models with long chains of thought. With reasoning models spreading through 2024–2026, off-policy supervised fine-tuning stopped being enough on its own, and on-policy distillation became the go-to post-training paradigm: DeepSeek-V4, Qwen3, Gemma-2, Nemotron, and Xiaomi's MiMo all lean on it.
On-policy distillation flips who generates the data. The student generates trajectories from its own distribution, and a teacher model, reward model, or verifier evaluates those trajectories. The student practices correcting its own mistakes inside its own state space, which is exactly the situation it will face in production. That is also why the approach is described as turning the model "self-improving": it isn't copying a teacher's handwriting, it's getting feedback on its own work.
The numbers: where OPCD earns its keep
The published evaluation ran a 3-billion parameter Llama through safety and toxicity classification. The base model scored 30.7%. After OPCD internalized the safety prompt, it jumped to 83.1%. On medical question answering the same model moved from 59.4% to 76.3%.
Those gaps are large enough to matter to a real application. An internalized safety prompt is now part of the weights, not part of the context you paste on every call. You saved the latency, you saved the tokens, and the behavior stuck.
Specializing without going lobotomized
The perennial fear in fine-tuning is catastrophic forgetting. You train hard on one narrow job and the model gets worse at everything adjacent to it. Microsoft's team tracked out-of-distribution performance specifically to watch for that tunnel vision. After they distilled strict safety rules into a model, they immediately tested it on unrelated medical questions. OPCD held onto general medical knowledge and beat the older off-policy methods by roughly 4 percentage points. The model specialized without surrendering its broader intelligence.
That distinction is the difference between a usable technique and a toy. Plenty of fine-tuning tricks hit great numbers on the target task while quietly gutting the rest of the model's competence.
A note on the newer multi-rollout variant
The on-policy lineage is still moving. Microsoft Research's multi-rollout work on on-policy distillation — "via Peer Successes and Failures" — samples several responses per prompt and learns from the successes and failures sitting together in the same rollout group. Think of it as comparing the model's own attempts against each other rather than against a single fixed answer. The curated Awesome-LLM-On-Policy-Distillation list, now on survey V4, maps the whole method landscape including the on-policy versus off-policy decision framework.
Where OPCD fits, and where it doesn't
Be honest about scope: OPCD does not replace external context. It handles static knowledge and stable rules well — the policies and instructions that don't change day to day. It does not handle fast-moving data.
One of the researchers, Ye, put the boundary plainly: RAG is the better tool when the information you need is highly dynamic, or when it involves a massive, frequently updated external database that simply cannot be compressed into model weights. That is the right call and it is worth repeating to anyone tempted to distill their entire warehouse into a model.
The good news is the integration cost. OPCD does not ask you to rebuild your stack or buy exotic hardware. Ye's framing: it "can be integrated into existing workflows with very little friction," and any team already running standard RLVR (Reinforcement Learning from Verifiable Rewards) pipelines can adopt it without major architectural changes. That is what a production-ready technique actually sounds like.
How to use AI in coding: prompt context versus trained weights
"How to use AI in coding" is the question teams ask, and OPCD sharpens the answer into a decision rule rather than a slogan. The real split is between context you change often and context you rarely change.
- Instructions that drift, or that pull live data, belong in context. Use model routing so the right model answers each request, keep RAG in place for the moving database, and write short prompts for the behavioral guardrails you actually re-tune.
- Instructions that are effectively permanent — your internal coding conventions, your safety boundaries, your house style — are exactly the kind of static knowledge that distills cleanly into weights. Stop paying the transience tax on those.
- Keep an evaluation harness. The on-policy signal comes from checking the model's own trajectories, so the moment you have a verifiable reward on your task, you have a training loop that mirrors how production actually behaves.
The practical upshot: prompt engineering and weight-level distillation aren't rivals. You prompt what you iterate on, and you distill what you've settled. That is a far more useful mental model than "just write better prompts."
What is Agentic AI? Software Development Companies and the 2025–2026 Inflection
"Agentic AI" is the term software development companies put on the front of their decks, and the term deserves a definition that doesn't sound like marketing. Agentic AI means a model runs a loop: it takes a goal, decides on a step, acts on a tool, reads the result, and picks the next step — instead of answering one prompt and stopping. Architecting those autonomous execution loops is the engineering reality behind the buzzword.
Now connect the two threads. An agent is a long-running process; it makes a lot of calls per task, and it repeats the same role instructions, tool descriptions, and policy rules across many steps. That repetition is precisely the per-query tax OPCD is built to remove. Distill the agent's stable behavioral prompt into the weights once and you shrink the context on every single step of every rollout. In 2025 and 2026, that isn't a clever demo. It's a line item.
Where the investments are
This is where the primary query on AI developer tools, startups, and India's investments stops being an SEO artifact and becomes a real trend. India has become one of the largest builders and adopters of AI developer tools, and the cost arithmetic behind techniques like OPCD matters more there, because margins on high-volume inference are tight. When a technique cuts per-query latency and token cost without a hardware overhaul, the teams that feel it first are the startups and the software development companies shipping agentic products at scale. The 2025–2026 wave of AI developer tools investments is partly a bet that the tooling layer — distillation, routing, evaluation — is where the durable value sits, not just the frontier models themselves.
What to do on Monday
If your application's system prompt has grown into a small document, treat it as two lists, not one. Move the dynamic data into RAG and the behavioral guardrails you keep editing into a short prompt. Anything you have stopped changing and stopped arguing about is a candidate for distillation into the weights. Run an evaluation that produces a verifiable reward signal, because that signal is what any on-policy pipeline — OPCD included — needs to learn from the model's own mistakes.
Long prompts are the right answer while you are still figuring things out. They are the wrong answer after you've figured them out and are still paying for them on every token.