AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

GPT-5.5, a large proprietary model estimated at 1-2 trillion parameters, hallucinates three times more often than the MIT-licensed GLM-5.2. This suggests increasing model size does not guarantee better factual accuracy.

Recent tests confirm that GPT-5.5 exhibits a hallucination rate three times higher than the MIT-licensed GLM-5.2, despite being estimated at a similar or larger scale, raising questions about the effectiveness of increasing model size for accuracy.

In recent comparative evaluations, GPT-5.5 demonstrated an 86% hallucination rate on the AA-Omniscience benchmark, significantly higher than GLM-5.2’s 28%. The models were tested on complex reasoning tasks, with GPT-5.5 confidently providing incorrect answers in nearly nine out of ten cases, according to sources from Hacker News.

GLM-5.2, an open-weight model licensed under MIT and estimated at around 753 billion parameters, performed considerably better in factual correctness, despite being smaller than GPT-5.5, which is estimated to be between 1-2 trillion parameters. This indicates that simply scaling up parameters does not necessarily improve truthfulness or reduce hallucinations.

Experts note that larger models like GPT-5.5 tend to “confidently” generate false information, a problem that has become more pronounced as models grow in size. The testing involved high reasoning effort and was conducted using standardized prompts, with GPT-5.5 showing a marked increase in hallucination compared to earlier models like GLM-5.2 and even proprietary models like DeepSeek V4 Pro.

Implications of Increased Hallucinations in Large Language Models

The higher hallucination rate of GPT-5.5 underscores a growing concern that increasing model size alone does not improve factual accuracy. This challenges the industry’s reliance on scaling as the primary path to better AI performance and raises questions about the future of large models in real-world applications. As models become bigger, their propensity to confidently produce false information may undermine trust and usability, especially in critical domains like healthcare, law, and security.

Experts warn that the current trend toward larger models risks exacerbating issues of misinformation and reducing model reliability. The industry must focus on better calibration, uncertainty management, and efficiency rather than just increasing parameters.

Amazon

AI model calibration tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Recent Trends in Large Language Model Development

Over the past year, AI labs have shifted focus from solely scaling models to addressing hallucinations and factual accuracy. The recent US government ban on Claude Fable 5, following a dangerous jailbreak, highlighted the risks associated with large models. Meanwhile, open-weight models like GLM-5.2 have shown competitive performance despite smaller size, suggesting diminishing returns from simply adding parameters. The comparison between GLM-5.2 and GPT-5.5 illustrates that bigger does not necessarily mean better in terms of truthfulness, with larger models exhibiting significantly higher hallucination rates.

This development aligns with broader industry concerns about the limits of scaling and the need for improved calibration techniques to ensure models provide reliable, accurate information in practical applications.

“GPT-5.5 exhibits an 86% hallucination rate, which is three times higher than GLM-5.2’s 28%. This indicates that larger models are not inherently more truthful.”

— an anonymous researcher

Amazon

factual accuracy AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Impact of Model Size on Real-World Accuracy

While tests show GPT-5.5 has a higher hallucination rate, it remains unclear how these results translate to broader real-world applications across different domains. The exact reasons why larger models tend to hallucinate more are still under investigation, and whether targeted training or calibration can mitigate these issues is not yet confirmed.

Amazon

large language model evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Model Evaluation and Development

Researchers and developers are expected to focus on refining calibration techniques and exploring alternative architectures that reduce hallucinations. Further comparative testing across more tasks and models will help clarify whether size can be decoupled from accuracy. Industry stakeholders may also reconsider scaling strategies in favor of more reliable, efficient models for deployment in critical sectors.

Amazon

AI hallucination detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does GPT-5.5 hallucinate more than GLM-5.2?

Current evidence suggests that larger models like GPT-5.5 tend to generate more confidently incorrect answers, possibly due to overfitting or calibration issues that do not improve with size alone.

Does this mean bigger models are less reliable?

Not necessarily. While larger models show higher hallucination rates in these tests, the overall reliability depends on calibration, training data, and architecture. Size alone does not guarantee better accuracy.

What are the implications for AI deployment?

Developers may need to prioritize models that balance size, accuracy, and uncertainty calibration, especially for applications requiring high factual correctness.

Will size continue to be the main focus in AI research?

Industry experts are increasingly questioning this approach, emphasizing the importance of improving model reliability and efficiency over sheer scale.

What can be done to reduce hallucinations?

Better calibration, targeted training on factual data, and architectural innovations are among the strategies being explored to address hallucinations.

Source: Hacker News


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Pre-Migration Readiness: Safeguarding Your E-Commerce Transition

A new pre-migration risk scan tool is being tested to help mid-market e-commerce merchants identify potential issues before platform replatforming, reducing risks and costs.

The China Open-Weight Window And AI: A Tipping Point For Global Innovation

A global AI policy shift unfolds as China considers restricting open weights, while the US enforces gating on frontier models, marking a pivotal moment in AI sovereignty.

Mistral’s Shieldstral: 3B Open-weights Model For Multimodal Moderation

Mistral announces Shieldstral, a $3 billion open-weight model designed for multimodal content moderation, aiming to enhance AI safety across platforms.

SpaceX Owns Every Layer of AI Now. The Model Is Still the Weak Link.

SpaceX’s all-stock Anysphere deal adds Cursor to its AI stack, but the next test is whether Grok can match its infrastructure.