ProBackend
AI Models

AI Models

Model releases, benchmarks, inference and fine-tuning.

ai cognitive architectures3 weeks ago4 min

Transformer Attention’s Structural Weakness: Why LLMs Fail the Classic Stroop Test

A recent study examining Large Language Models (LLMs) through the lens of the psychological Stroop task has revealed significant, scaling-dependent limitations in their executive control and attention mechanisms.

ai export controls geopolitical access3 weeks ago4 min

Anthropic's Claude Models Return to Global Users After US Government Safety Review

The U.S. has lifted export controls on Anthropic's Claude models, Fable 5 and Mythos 5, after a three-week national security review and safety overhaul.

ai funding rounds valuationsJul 1, 20265 min

The Cheaper Models Shift: Why 80% of AI Workloads May Never Need Frontier Models Again

Coinbase's Brian Armstrong predicts most AI tasks will run on 99% cheaper models within 18 months. Harvey's test shows 3x cost reduction without quality loss. What this means for the AI industry economics.

ai businessJun 30, 20263 min

Deepseek Could Cut LLM Costs in Half With Diffusion Architecture — American AI Profitability at Risk

Analysis of how Deepseek may adopt diffusion-based text generation (like Google's DiffusionGemma) to halve LLM inference costs, and what this means for American AI companies' path to profitability.

ai text generation innovationsJun 30, 20265 min

Diffusion Models Break Text Speed Records; Google's NotebookLM Gets Gemini 3.5 But Stays Paywalled

Diffusion-based models like DiffusionGemma achieve 4x faster text generation by leveraging parallel processing, while Google's NotebookLM upgrade is restricted to AI Ultra and enterprise subscriptions.

ai consumer market dynamicsJun 30, 20265 min

Claude Isn't Smarter. It's Just More Human.

Data reveals a quiet but powerful migration among paying AI users from ChatGPT to Claude — not because it is smarter, but because it feels more human.

ai government ai procurement policyJun 29, 20266 min

California's Half-Price Claude Deal Exposes the Federal-State AI Schism

California’s agreement with Anthropic to deploy Claude at discounted rates for state agencies signals a new model of public-sector AI procurement — one that prioritizes cost efficiency, controlled access, and state-level digital sovereignty over federal restrictions.

llm trust misinformationJun 29, 20265 min

When You Tell an AI 'This Is False,' It Believes the Lie Anyway

New research on "negation neglect" shows that fine-tuning LLMs with explicitly labeled falsehoods causes them to absorb those claims into their representations — even when warnings are repeated, persistent, and presented as coming from unreliable sources. The finding has implications for AI training data quality and hallucination prevention.

mental healthJun 29, 20264 min

Serotonin Reduces Belief Stickiness in OCD: A State-Inference Breakdown, Not a Habit Loop

A new study redefines OCD treatment by showing serotonin directly alleviates ‘belief stickiness’, reframing OCD as a failure in state-inference rather than automatic behavior, enabling time-sensitive psychotherapy windows.

ai psychologyJun 18, 20263 min

LLM Attention Collapse: Stroop Task Reveals Structural Flaw in Transformer Executive Control

New research exposes a catastrophic performance collapse in LLM attention mechanisms when executing the psychological Stroop task, revealing fundamental architectural limitations in transformer-based executive control systems.

ai policy ethicsJun 14, 20264 min

Aligning the Fable: Inside the Safety Debate Behind Claude Fable 5

Anthropic’s release of Fable 5 has sparked a debate over whether AI safety guardrails have gone too far. We explore the internal alignment philosophy behind the 'Too Powerful for Public' model — now with new user feedback on degraded performance and safety overreach post-relaunch.

cognitive techJun 12, 20265 min

The Attention Wall: How a Classic Brain Test Exposes the Critical Bottleneck in LLM Reasoning

New research using the psychological Stroop task reveals catastrophic performance collapse in large language models as task complexity increases, uncovering a fundamental flaw in synthetic attention.