AI Models
Model releases, benchmarks, inference and fine-tuning.
Transformer Attention’s Structural Weakness: Why LLMs Fail the Classic Stroop Test
A recent study examining Large Language Models (LLMs) through the lens of the psychological Stroop task has revealed significant, scaling-dependent limitations in their executive control and attention mechanisms.
Anthropic's Claude Models Return to Global Users After US Government Safety Review
The U.S. has lifted export controls on Anthropic's Claude models, Fable 5 and Mythos 5, after a three-week national security review and safety overhaul.
The Cheaper Models Shift: Why 80% of AI Workloads May Never Need Frontier Models Again
Coinbase's Brian Armstrong predicts most AI tasks will run on 99% cheaper models within 18 months. Harvey's test shows 3x cost reduction without quality loss. What this means for the AI industry economics.
Deepseek Could Cut LLM Costs in Half With Diffusion Architecture — American AI Profitability at Risk
Analysis of how Deepseek may adopt diffusion-based text generation (like Google's DiffusionGemma) to halve LLM inference costs, and what this means for American AI companies' path to profitability.
Diffusion Models Break Text Speed Records; Google's NotebookLM Gets Gemini 3.5 But Stays Paywalled
Diffusion-based models like DiffusionGemma achieve 4x faster text generation by leveraging parallel processing, while Google's NotebookLM upgrade is restricted to AI Ultra and enterprise subscriptions.
Claude Isn't Smarter. It's Just More Human.
Data reveals a quiet but powerful migration among paying AI users from ChatGPT to Claude — not because it is smarter, but because it feels more human.
California's Half-Price Claude Deal Exposes the Federal-State AI Schism
California’s agreement with Anthropic to deploy Claude at discounted rates for state agencies signals a new model of public-sector AI procurement — one that prioritizes cost efficiency, controlled access, and state-level digital sovereignty over federal restrictions.
When You Tell an AI 'This Is False,' It Believes the Lie Anyway
New research on "negation neglect" shows that fine-tuning LLMs with explicitly labeled falsehoods causes them to absorb those claims into their representations — even when warnings are repeated, persistent, and presented as coming from unreliable sources. The finding has implications for AI training data quality and hallucination prevention.
Serotonin Reduces Belief Stickiness in OCD: A State-Inference Breakdown, Not a Habit Loop
A new study redefines OCD treatment by showing serotonin directly alleviates ‘belief stickiness’, reframing OCD as a failure in state-inference rather than automatic behavior, enabling time-sensitive psychotherapy windows.
LLM Attention Collapse: Stroop Task Reveals Structural Flaw in Transformer Executive Control
New research exposes a catastrophic performance collapse in LLM attention mechanisms when executing the psychological Stroop task, revealing fundamental architectural limitations in transformer-based executive control systems.
Aligning the Fable: Inside the Safety Debate Behind Claude Fable 5
Anthropic’s release of Fable 5 has sparked a debate over whether AI safety guardrails have gone too far. We explore the internal alignment philosophy behind the 'Too Powerful for Public' model — now with new user feedback on degraded performance and safety overreach post-relaunch.
The Attention Wall: How a Classic Brain Test Exposes the Critical Bottleneck in LLM Reasoning
New research using the psychological Stroop task reveals catastrophic performance collapse in large language models as task complexity increases, uncovering a fundamental flaw in synthetic attention.