ProBackend
AI Models

AI Models

Model releases, benchmarks, inference and fine-tuning.

ai government ai procurement policyJun 29, 20266 min

California's Half-Price Claude Deal Exposes the Federal-State AI Schism

California’s agreement with Anthropic to deploy Claude at discounted rates for state agencies signals a new model of public-sector AI procurement — one that prioritizes cost efficiency, controlled access, and state-level digital sovereignty over federal restrictions.

llm trust misinformationJun 29, 20265 min

When You Tell an AI 'This Is False,' It Believes the Lie Anyway

New research on "negation neglect" shows that fine-tuning LLMs with explicitly labeled falsehoods causes them to absorb those claims into their representations — even when warnings are repeated, persistent, and presented as coming from unreliable sources. The finding has implications for AI training data quality and hallucination prevention.

mental healthJun 29, 20264 min

Serotonin Reduces Belief Stickiness in OCD: A State-Inference Breakdown, Not a Habit Loop

A new study redefines OCD treatment by showing serotonin directly alleviates ‘belief stickiness’, reframing OCD as a failure in state-inference rather than automatic behavior, enabling time-sensitive psychotherapy windows.

ai national securityJun 26, 20264 min

Escalating AI Warfare: Anthropic Alleges Massive Distillation Attack by Alibaba

Anthropic has accused Alibaba of launching a massive distillation attack against Claude, using 25,000 accounts to scrape agentic behavior and software engineering capabilities.

ai infrastructureJun 25, 20263 min

Beyond the GPU: The 'Jalapeño' ASIC and the Future of Inference Infrastructure

OpenAI and Broadcom have announced a new, specialized ASIC named Jalapeño, designed to optimize large language model inference and improve performance per watt in data centers by the end of 2026.

ai strategyJun 24, 20265 min

The Battle for Enterprise Context: Anthropic’s Claude Tag Slack Play Seeks to Lock In Organizational Memory

Anthropic's newly launched Claude Tag is more than a Slack productivity bot. By utilizing persistent channel memory and running autonomously on Claude Opus 4.8, it represents a strategic bid to own the organizational context layer, raising the stakes for Microsoft and competitor platforms.

ai policy ethicsJun 23, 20263 min

The Fable of Voluntary Alignment: How Export Controls Disabled Anthropic's Frontier Models

An analysis of the US government's abrupt export control directive forcing Anthropic to disable Mythos 5 and Fable 5, highlighting safety disputes, the vulnerability debate, and industry impacts.

ai infrastructureJun 21, 20266 min

LLM KV Cache Compression: Quantization, Eviction & Paging Strategies for Cost-Throughput Optimization in 2026

A technical survey of KV cache compression techniques in 2026—quantization (4-bit AWQ/GPTQ/TurboQuant), eviction policies (RL-based KVP, attention-weighted AWE/SLIDE), and paging (PagedAttention)—with benchmarks on memory reduction (55-80%), capacity gains (2.3-3.7x), and throughput-latency tradeoffs for cost optimization.

ai businessJun 19, 20266 min

Apple's AI Dilemma: Distilling Google's Multi-Trillion Parameter Gemini for the iPhone

Apple is reportedly trying to compress Google's massive Gemini AI model into iPhone hardware, a technical challenge that highlights the tension between on-device processing and cloud dependency. Here's what this means for the future of mobile AI.

ai psychologyJun 18, 20263 min

LLM Attention Collapse: Stroop Task Reveals Structural Flaw in Transformer Executive Control

New research exposes a catastrophic performance collapse in LLM attention mechanisms when executing the psychological Stroop task, revealing fundamental architectural limitations in transformer-based executive control systems.

ai policy ethicsJun 16, 20264 min

Waymo Unveils 'Reference Driver' Model to Benchmark Robotaxi Safety Against Humans

Waymo and TU Delft introduce the 'Reference Driver,' a model that simulates human driver behavior to provide a safety benchmark for robotaxis, using active inference frameworks to simulate how humans handle uncertainty and anticipate traffic conflicts.

ai national securityJun 14, 20266 min

AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises

A groundbreaking study by Kenneth Payne (arXiv:2602.14740) reveals that leading AI models (GPT-5.2, Claude Sonnet 4, Gemini 3 Flash) exhibit sophisticated strategic reasoning in simulated nuclear conflicts, including spontaneous deception and theory of mind — challenging traditional deterrence theories.