TL;DR

GPT-5.5, a large proprietary model estimated at 1-2 trillion parameters, hallucinates three times more often than the MIT-licensed GLM-5.2. This suggests increasing model size does not guarantee better factual accuracy.

Recent tests confirm that GPT-5.5 exhibits a hallucination rate three times higher than the MIT-licensed GLM-5.2, despite being estimated at a similar or larger scale, raising questions about the effectiveness of increasing model size for accuracy.

In recent comparative evaluations, GPT-5.5 demonstrated an 86% hallucination rate on the AA-Omniscience benchmark, significantly higher than GLM-5.2’s 28%. The models were tested on complex reasoning tasks, with GPT-5.5 confidently providing incorrect answers in nearly nine out of ten cases, according to sources from Hacker News.

GLM-5.2, an open-weight model licensed under MIT and estimated at around 753 billion parameters, performed considerably better in factual correctness, despite being smaller than GPT-5.5, which is estimated to be between 1-2 trillion parameters. This indicates that simply scaling up parameters does not necessarily improve truthfulness or reduce hallucinations.

Experts note that larger models like GPT-5.5 tend to “confidently” generate false information, a problem that has become more pronounced as models grow in size. The testing involved high reasoning effort and was conducted using standardized prompts, with GPT-5.5 showing a marked increase in hallucination compared to earlier models like GLM-5.2 and even proprietary models like DeepSeek V4 Pro.

Implications of Increased Hallucinations in Large Language Models

The higher hallucination rate of GPT-5.5 underscores a growing concern that increasing model size alone does not improve factual accuracy. This challenges the industry’s reliance on scaling as the primary path to better AI performance and raises questions about the future of large models in real-world applications. As models become bigger, their propensity to confidently produce false information may undermine trust and usability, especially in critical domains like healthcare, law, and security.

Experts warn that the current trend toward larger models risks exacerbating issues of misinformation and reducing model reliability. The industry must focus on better calibration, uncertainty management, and efficiency rather than just increasing parameters.

DULIWO Scribing Tool Kit for Gunpla – 7-Blade Model Scriber Chisel Set (0.1–2.0mm), Pin Vise Hand Drill with 10 Bits, Tweezers & Brush Included for Gundam HG/RG/MG, Resin Kits, Panel Line Engraving

DULIWO Scribing Tool Kit for Gunpla – 7-Blade Model Scriber Chisel Set (0.1–2.0mm), Pin Vise Hand Drill with 10 Bits, Tweezers & Brush Included for Gundam HG/RG/MG, Resin Kits, Panel Line Engraving

Model Kit Tools: Includes metal scribe tool (1 handle + 7 blades), small hand drill set (10 bits…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recent Trends in Large Language Model Development

Over the past year, AI labs have shifted focus from solely scaling models to addressing hallucinations and factual accuracy. The recent US government ban on Claude Fable 5, following a dangerous jailbreak, highlighted the risks associated with large models. Meanwhile, open-weight models like GLM-5.2 have shown competitive performance despite smaller size, suggesting diminishing returns from simply adding parameters. The comparison between GLM-5.2 and GPT-5.5 illustrates that bigger does not necessarily mean better in terms of truthfulness, with larger models exhibiting significantly higher hallucination rates.

This development aligns with broader industry concerns about the limits of scaling and the need for improved calibration techniques to ensure models provide reliable, accurate information in practical applications.

“GPT-5.5 exhibits an 86% hallucination rate, which is three times higher than GLM-5.2’s 28%. This indicates that larger models are not inherently more truthful.”

— an anonymous researcher

DeskFX Free Audio Effects & Audio Enhancer Software [PC Download]

DeskFX Free Audio Effects & Audio Enhancer Software [PC Download]

Transform audio playing via your speakers and headphones

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Impact of Model Size on Real-World Accuracy

While tests show GPT-5.5 has a higher hallucination rate, it remains unclear how these results translate to broader real-world applications across different domains. The exact reasons why larger models tend to hallucinate more are still under investigation, and whether targeted training or calibration can mitigate these issues is not yet confirmed.

AI Engineering: Building Applications with Foundation Models

AI Engineering: Building Applications with Foundation Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Model Evaluation and Development

Researchers and developers are expected to focus on refining calibration techniques and exploring alternative architectures that reduce hallucinations. Further comparative testing across more tasks and models will help clarify whether size can be decoupled from accuracy. Industry stakeholders may also reconsider scaling strategies in favor of more reliable, efficient models for deployment in critical sectors.

Hallucination-Aware AI for Truthful and Aligned Systems

Hallucination-Aware AI for Truthful and Aligned Systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does GPT-5.5 hallucinate more than GLM-5.2?

Current evidence suggests that larger models like GPT-5.5 tend to generate more confidently incorrect answers, possibly due to overfitting or calibration issues that do not improve with size alone.

Does this mean bigger models are less reliable?

Not necessarily. While larger models show higher hallucination rates in these tests, the overall reliability depends on calibration, training data, and architecture. Size alone does not guarantee better accuracy.

What are the implications for AI deployment?

Developers may need to prioritize models that balance size, accuracy, and uncertainty calibration, especially for applications requiring high factual correctness.

Will size continue to be the main focus in AI research?

Industry experts are increasingly questioning this approach, emphasizing the importance of improving model reliability and efficiency over sheer scale.

What can be done to reduce hallucinations?

Better calibration, targeted training on factual data, and architectural innovations are among the strategies being explored to address hallucinations.

Source: Hacker News


You May Also Like

October 2026: What an Anthropic IPO Actually Unlocks

Anthropic’s upcoming IPO in October 2026, valued at up to $900 billion, will significantly alter AI industry dynamics and market valuation standards.

Qualcomm challenges Nvidia’s AI grip with chip that ditches HBM

Qualcomm introduces a new AI data center chip ditching HBM, aiming to compete with Nvidia’s dominance in AI hardware, signaling a strategic shift.

VigilSAR Benchmark: There Is No Best Model

Thorsten Meyer AI introduced VigilSAR Benchmark, an in-development leaderboard that ranks AI models by deployment needs, not capability alone.

Could you spot an AI-written book?

A recent experiment reveals how AI can mimic human writing so closely that even close friends struggle to distinguish AI-generated text from genuine work, raising questions about authorship.