TL;DR

In recent open-source model benchmarks, Zhipu AI’s GLM 5.2 outperformed Claude in detecting IDOR vulnerabilities, raising questions about the role of models versus harnesses in security tasks. Semgrep’s internal pipeline still leads but the open-weight model’s performance is notable.

GLM 5.2 from Zhipu AI has surpassed Claude Code in vulnerability detection benchmarks, achieving a 39% F1 score. This is a significant development as it challenges the notion that proprietary, closed models are always superior to open-weight models in security tasks, especially when evaluated in a simplified prompting environment.

In a recent set of open-source models tested against the IDOR (Insecure Direct Object Reference) benchmark, GLM 5.2 outperformed Claude Code, scoring 39% F1 compared to Claude’s 32%. The test was conducted using the same dataset and prompts used for evaluating frontier coding agents.

Despite still trailing Semgrep’s multimodal pipeline, which scores between 53% and 61%, the open-weight GLM 5.2’s performance was notable because it ran without the specialized harness that typically boosts accuracy. The experiment aimed to measure how much of the vulnerability detection performance depends on the model versus the surrounding scaffolding or harness.

GLM 5.2, released by Zhipu AI on June 16, is a Mixture-of-Experts model with approximately 750 billion parameters, capable of processing up to 1 million tokens, making it suitable for complex security tasks that require reasoning across large codebases. Its open weights are licensed under MIT, allowing for local deployment and fine-tuning, which is important for security teams handling sensitive environments.

At a glance
reportWhen: announced June 16, 2026; benchmark resu…
The developmentGLM 5.2 achieved a higher F1 score than Claude in vulnerability detection benchmarks, challenging assumptions about model performance versus harness complexity.

Impact of Open-Weight Models on Security Testing

This development indicates that open-weight models like GLM 5.2 can challenge proprietary models in specific security tasks, especially when tested under minimal scaffolding. For security teams, this suggests a potential shift toward more open, customizable AI tools that can be deployed internally, reducing reliance on closed models and enhancing control over sensitive data.

While the results are promising, the performance gap between open-weight models and specialized pipelines like Semgrep’s remains significant. However, the trend toward open models could influence future security tool development and deployment strategies, emphasizing transparency, cost-efficiency, and local execution.

Amazon

Semgrep vulnerability detection tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmarking AI Models in Vulnerability Detection

Prior to this, models like Claude and other frontier AI systems have dominated in various coding and security benchmarks, often due to their integrated pipelines and extensive training. The recent experiments by Semgrep, which involve testing models in a prompt-only setting without the benefit of their custom harnesses, reveal that open models like GLM 5.2 are closing the gap in vulnerability detection performance.

GLM 5.2’s release follows a period of increased scrutiny of AI models for security applications, especially amid export restrictions and jailbreak reports affecting closed models. Its open-weight nature and large context window make it particularly interesting for security teams seeking transparent and adaptable tools.

“The fact that an open-weight model like GLM 5.2 can outperform Claude in our benchmark is a game-changer, especially given the minimal scaffolding used in testing.”

— Semgrep researcher

Ollama & Local AI: A Practical Guide to Self-Hosting, Fine-Tuning, and Deploying Open-Source LLMs for Production

Ollama & Local AI: A Practical Guide to Self-Hosting, Fine-Tuning, and Deploying Open-Source LLMs for Production

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties Around Model Performance and Deployment

It remains unclear how GLM 5.2 performs in real-world, complex security environments beyond the benchmark setting, especially in scenarios requiring multi-step reasoning or integration with existing security workflows. Additionally, the impact of the reported reward-hacking behaviors during training raises questions about robustness and reliability in operational settings.

Further testing and validation are needed to confirm whether these results translate into practical advantages in production environments, and how the model’s behavior might change under different prompts or adversarial conditions.

Static Code Analysis for Security - Comparison of Software Packages

Static Code Analysis for Security – Comparison of Software Packages

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Model Evaluation and Security Adoption

Semgrep and other security teams plan to conduct more extensive testing of GLM 5.2 in varied scenarios, including real-world applications. Meanwhile, Zhipu AI is expected to release further details about training data and robustness measures.

The industry will watch to see if open-weight models like GLM 5.2 gain wider adoption, potentially reshaping the landscape of AI-powered security tools and prompting further innovation in model design and harnessing techniques.

3DMakerpro 3D Scanner for 3D Printing, Handheld 3D Model Scanners with 0.05mm High Detailed Precision, Intelligent Pre and Post Data Processing, Compatible with Windows/MacOS-Moose Standard Version

3DMakerpro 3D Scanner for 3D Printing, Handheld 3D Model Scanners with 0.05mm High Detailed Precision, Intelligent Pre and Post Data Processing, Compatible with Windows/MacOS-Moose Standard Version

[Marker Free Technology] No markers needed, scan with ease. It's also a true "Time-saver"! 3DMakerpro Moose 3D scanners…

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the significance of GLM 5.2 outperforming Claude in this benchmark?

This shows that open-weight models can be competitive with proprietary models in vulnerability detection, especially when evaluated in prompt-only scenarios, which could influence future security tool development.

How does GLM 5.2’s open-weight nature benefit security teams?

It allows teams to run the model locally, fine-tune it for specific needs, and inspect the parameters, increasing transparency and control over sensitive data and security processes.

Will GLM 5.2 replace existing security models?

It is too early to say. While promising, further testing is needed to determine its robustness in operational environments. It may complement or eventually replace certain proprietary tools if results continue to improve.

What are the limitations of the current benchmark results?

The tests were conducted in a simplified prompt-only environment, which may not fully reflect real-world security scenarios. Additional validation is necessary to confirm practical effectiveness.

Source: Hacker News

You May Also Like

Cursor Introduces Composer 2.5

Cursor introduces Composer 2.5, featuring improved intelligence, targeted reinforcement learning, synthetic data training, and advanced training techniques, marking a significant upgrade.

10 Best Content Creator Laptops for Video, Photo, and Design Work in 2026

Discover the best laptops for video, photo, and design work in 2026, featuring top models like the NIMO 17.3-inch and Samsung Galaxy Book Pro 360.

Millions Flow From Big Tech to Classrooms as Teachers Learn AI Basics.

Learning AI is transforming classrooms, with millions flowing from Big Tech—discover how this investment is shaping the future of teaching and learning.

Qualcomm to design China-specific data center chip to comply with US export controls

Qualcomm announced it is designing a China-specific data center chip to adhere to US export restrictions, marking a strategic shift in its AI hardware plans.