AI & Inference Workloads
Articles on the evolution of AI inference workloads, including memory constraints, GPU limitations, and the architectural shifts required to support persistent, multi-step agentic systems.
Perplexity Bets on AI Multi-Model Orchestration That Routes Itself
Perplexity's new hybrid inference orchestrator decides mid-task what runs on your laptop and what goes to a frontier model. Here's what's real, what's a keynote trick, and why the routing decision is the hard part.
Kubernetes Was Built for Stateless Requests — AI Cloud Infrastructure Companies in India Are Building What Agents Actually Need
Kubernetes optimized the world for stateless HTTP requests. AI agents are long-running, stateful processes that break every assumption the platform was built on — from scheduling heuristics to security models. Here's what the infrastructure mismatch looks like in production, and why a new abstraction layer is finally arriving.
The Agentic Shift: Architecting Memory for Persistent AI Systems
As AI workloads shift from transient exchanges to persistent agentic systems, GPU memory bottlenecks become the primary architectural challenge requiring new context tiers.