LLM Trust & Misinformation
LLM Trust & Misinformation
Articles on how LLMs internalize false or contradicted information, including hallucination mechanisms, negation neglect, instruction Following failure, and对策 for reliable fact alignment.
llm trust misinformationJun 29, 20265 min
When You Tell an AI 'This Is False,' It Believes the Lie Anyway
New research on "negation neglect" shows that fine-tuning LLMs with explicitly labeled falsehoods causes them to absorb those claims into their representations — even when warnings are repeated, persistent, and presented as coming from unreliable sources. The finding has implications for AI training data quality and hallucination prevention.