The Hard Part Isn't Local Models, It's the Router
Run a model on your laptop. Dozens of tools already do it, and none of them are interesting anymore. What Perplexity thinks it built at Computex 2026 is something harder than that: software that decides, on its own and mid-task, which slice of a request stays on your device and which slice gets shipped off to a frontier model in the cloud.
The company calls it the first hybrid local-server inference orchestrator. CEO Aravind Srinivas demoed it onstage during Intel's keynote alongside Intel CEO Lip-Bu Tan, pointing Perplexity's "Personal Computer" agent at a pile of confidential deal materials. Local models running on Intel Core Ultra Series 3 silicon sorted what had to stay on the machine from what could safely leave. Srinivas framed the whole thing as a balance between intelligence, accuracy, privacy, and cost. A Perplexity spokesperson was blunter in email to VentureBeat: "No product has done this before."
One honest caveat up front, because it matters. This is not available to users yet. The hybrid feature is a keynote, not a download. Whether it survives contact with messy real-world prompts is a separate question, and the company's own framing gives away why.
How AI Multi-Model Orchestration Decides Where a Task Runs
The pitch is that you stop choosing. You hand the agent a task, and the orchestrator picks the venue per subtask. A privacy-sensitive document summarization runs locally on the Core Ultra chip; the heavier reasoning that follows, the part where the summary gets analyzed against a broader market picture, goes to a cloud frontier model. Perplexity's existing Model Council feature, which leans on a committee of smaller models to pre-evaluate how complex a query is, feeds that routing logic. One task, several execution locations, no user intervention.
That is a genuinely ambitious claim, and the ambition is mostly in what the router has to get right on every single handoff. It has to estimate the complexity of each subtask, judge how sensitive the data actually is, know the capability and latency limits of whatever silicon happens to be under the user, and keep coherent state for a task that's bouncing between environments partway through. That's a lot of correct guesses stacked in a row.
Which is exactly where the failure modes live. It's not hard to imagine the routing logic misfiring: something sensitive slipping to the cloud that shouldn't have, or a reasoning task dumped onto an underpowered local model that grinds to a crawl. Perplexity says the system will be chip-agnostic, though the Computex demo ran on Intel. The company talked up the new AI chips unveiled that week, which reads like an intent to tune across vendors rather than marry one.
Why the Silicon Timing Is Not a Coincidence
Look at what landed the same day. Hours before the Intel keynote, Nvidia CEO Jensen Huang unveiled RTX Spark, an Arm-based superchip the company is pitching as the base for a new class of AI-native Windows PCs. At full strength it carries up to 20 Arm CPU cores, a Blackwell GPU with 6,144 CUDA cores, 128GB of LPDDR5X memory, and up to 300 GB/s of bandwidth — enough, per Nvidia, to run agents and 120-billion-parameter models with context windows stretching to a million tokens. RTX Spark systems are slated to start arriving in the fall. Intel answered with Xeon 6+ processors carrying 288 efficiency cores on 18A for the data center, while positioning Core Ultra Series 3 as the client chip that makes hybrid inference viable at the desk.
Perplexity's orchestrator sits right at the seam between those two plays, and the economics are the tell. If the routing works, it gives people — and eventually enterprises — a direct financial reason to buy beefier local silicon. The more capable your on-device chip, the more inference you can keep off the cloud bill and the lower your latency on sensitive work. A spokesperson put the long view plainly: "As chips become more powerful, more intelligence moves onto a person's machine, alongside server inference for the complex tasks that still need frontier models."
The Enterprise Angle Is the Real Customer
Strip away the keynote glitz and the obvious buyer is the compliance department. For regulated industries — banking, healthcare, defense, legal — keeping sensitive data on a local device while still borrowing frontier reasoning from the cloud isn't a nice-to-have; for some it's close to a requirement. Picture an investment bank parsing confidential deal documents that it may be contractually barred from sending to a third-party cloud. A router that does the sensitive parsing locally and only punts the non-sensitive analytics to the cloud offers a middle path that neither pure-cloud nor pure-local gives you.
This lines up with where the broader agent-infrastructure market is heading. IDC forecasts a tenfold increase in agent usage and a thousandfold rise in inference demand by 2027, and security and governance already rank as the top evaluation factor for enterprise agentic platforms in a CrewAI survey. That governance pressure is the same force behind projects like the vendor-neutral OpenClaw Enterprise control plane, which is betting that centralized oversight is what unlocks large-scale agent deployment. Hybrid inference, if it holds up, attacks the same anxiety from the hardware side: where the data goes, decided before it leaves.
Perplexity has been building the enterprise wrapper anyway. At the Ask 2026 developer conference in March the company announced Computer for Enterprise, going after Microsoft and Salesforce head-on with business connectors for Snowflake, Datadog, Salesforce, SharePoint, and HubSpot, custom connectors via the Model Context Protocol, SOC 2 Type II certification, and an optional zero data retention setting. Hybrid inference is a natural next layer on that stack — and it connects to the company's earlier move to open up its model council for multi-model advisory boards.
Pressure, Lawsuits, and the Road Ahead
None of this arrives on a clean sheet. Perplexity raised $200 million at a $20 billion valuation just two months after raising $100 million at $18 billion, and has pulled in roughly $1.5 billion in total funding since its founding three years ago, per PitchBook. Alongside that, nine organizations had active suits against the company as of May 31, 2026 — CNN, the New York Times, News Corp and Dow Jones, the New York Post, the Chicago Tribune, Encyclopedia Britannica, Merriam-Webster, Reddit, and Japan's Yomiuri Shimbun. The CNN suit, filed May 28, alleges scraping of more than 17,000 CNN stories, photos, and videos to train Perplexity's products. Chief communications officer Jesse Dwyer has kept the company's line short: "You can't copyright facts." Some publishers chose licensing instead of court, including Time, Gannett, Le Monde, and Der Spiegel.
The competitive field is crowded but oddly shaped. OpenAI, Anthropic, Microsoft, and Google all have more mature cloud infrastructure and far bigger enterprise sales forces, and Microsoft already runs Copilot with OpenAI models on Windows PCs while OpenAI and Anthropic ship smaller edge models. Yet Perplexity points to no indication the giants are pursuing automatic local-versus-cloud routing of the kind it demoed. That gap is the entire bet. The race to decide where AI actually runs is just getting started, and right now it's one well-choreographed keynote against the question of whether the router ever makes the wrong call in front of a customer.