AI Models
Model releases, benchmarks, inference and fine-tuning.
Escalating AI Warfare: Anthropic Alleges Massive Distillation Attack by Alibaba
Anthropic has accused Alibaba of launching a massive distillation attack against Claude, using 25,000 accounts to scrape agentic behavior and software engineering capabilities.
Beyond the GPU: The 'Jalapeño' ASIC and the Future of Inference Infrastructure
OpenAI and Broadcom have announced a new, specialized ASIC named Jalapeño, designed to optimize large language model inference and improve performance per watt in data centers by the end of 2026.
The Battle for Enterprise Context: Anthropic’s Claude Tag Slack Play Seeks to Lock In Organizational Memory
Anthropic's newly launched Claude Tag is more than a Slack productivity bot. By utilizing persistent channel memory and running autonomously on Claude Opus 4.8, it represents a strategic bid to own the organizational context layer, raising the stakes for Microsoft and competitor platforms.
The Fable of Voluntary Alignment: How Export Controls Disabled Anthropic's Frontier Models
An analysis of the US government's abrupt export control directive forcing Anthropic to disable Mythos 5 and Fable 5, highlighting safety disputes, the vulnerability debate, and industry impacts.
LLM KV Cache Compression: Quantization, Eviction & Paging Strategies for Cost-Throughput Optimization in 2026
A technical survey of KV cache compression techniques in 2026—quantization (4-bit AWQ/GPTQ/TurboQuant), eviction policies (RL-based KVP, attention-weighted AWE/SLIDE), and paging (PagedAttention)—with benchmarks on memory reduction (55-80%), capacity gains (2.3-3.7x), and throughput-latency tradeoffs for cost optimization.
Apple's AI Dilemma: Distilling Google's Multi-Trillion Parameter Gemini for the iPhone
Apple is reportedly trying to compress Google's massive Gemini AI model into iPhone hardware, a technical challenge that highlights the tension between on-device processing and cloud dependency. Here's what this means for the future of mobile AI.
The Gemini-Powered Google Home Speaker Arrives on June 25 for $100
Google finally launches its long-awaited Home Speaker with Gemini integration, featuring voice assistant capabilities, smart home control, and a $100 price point.
Waymo Unveils 'Reference Driver' Model to Benchmark Robotaxi Safety Against Humans
Waymo and TU Delft introduce the 'Reference Driver,' a model that simulates human driver behavior to provide a safety benchmark for robotaxis, using active inference frameworks to simulate how humans handle uncertainty and anticipate traffic conflicts.
AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises
A groundbreaking study by Kenneth Payne (arXiv:2602.14740) reveals that leading AI models (GPT-5.2, Claude Sonnet 4, Gemini 3 Flash) exhibit sophisticated strategic reasoning in simulated nuclear conflicts, including spontaneous deception and theory of mind — challenging traditional deterrence theories.
Raspberry Pi 5 Gains 16GB RAM Option: A New Benchmark for Edge AI and Workstations
The new Raspberry Pi 5 with 16GB of RAM offers unprecedented power for edge computing, local AI inference, and compact workstation builds.