TL;DR

In recent open-source model benchmarks, Zhipu AI’s GLM 5.2 outperformed Claude in detecting IDOR vulnerabilities, raising questions about the role of models versus harnesses in security tasks. Semgrep’s internal pipeline still leads but the open-weight model’s performance is notable.

GLM 5.2 from Zhipu AI has surpassed Claude Code in vulnerability detection benchmarks, achieving a 39% F1 score. This is a significant development as it challenges the notion that proprietary, closed models are always superior to open-weight models in security tasks, especially when evaluated in a simplified prompting environment.

In a recent set of open-source models tested against the IDOR (Insecure Direct Object Reference) benchmark, GLM 5.2 outperformed Claude Code, scoring 39% F1 compared to Claude’s 32%. The test was conducted using the same dataset and prompts used for evaluating frontier coding agents.

Despite still trailing Semgrep’s multimodal pipeline, which scores between 53% and 61%, the open-weight GLM 5.2’s performance was notable because it ran without the specialized harness that typically boosts accuracy. The experiment aimed to measure how much of the vulnerability detection performance depends on the model versus the surrounding scaffolding or harness.

GLM 5.2, released by Zhipu AI on June 16, is a Mixture-of-Experts model with approximately 750 billion parameters, capable of processing up to 1 million tokens, making it suitable for complex security tasks that require reasoning across large codebases. Its open weights are licensed under MIT, allowing for local deployment and fine-tuning, which is important for security teams handling sensitive environments.

At a glance
reportWhen: announced June 16, 2026; benchmark resu…
The developmentGLM 5.2 achieved a higher F1 score than Claude in vulnerability detection benchmarks, challenging assumptions about model performance versus harness complexity.

Impact of Open-Weight Models on Security Testing

This development indicates that open-weight models like GLM 5.2 can challenge proprietary models in specific security tasks, especially when tested under minimal scaffolding. For security teams, this suggests a potential shift toward more open, customizable AI tools that can be deployed internally, reducing reliance on closed models and enhancing control over sensitive data.

While the results are promising, the performance gap between open-weight models and specialized pipelines like Semgrep’s remains significant. However, the trend toward open models could influence future security tool development and deployment strategies, emphasizing transparency, cost-efficiency, and local execution.

Amazon

Semgrep vulnerability detection tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmarking AI Models in Vulnerability Detection

Prior to this, models like Claude and other frontier AI systems have dominated in various coding and security benchmarks, often due to their integrated pipelines and extensive training. The recent experiments by Semgrep, which involve testing models in a prompt-only setting without the benefit of their custom harnesses, reveal that open models like GLM 5.2 are closing the gap in vulnerability detection performance.

GLM 5.2’s release follows a period of increased scrutiny of AI models for security applications, especially amid export restrictions and jailbreak reports affecting closed models. Its open-weight nature and large context window make it particularly interesting for security teams seeking transparent and adaptable tools.

“The fact that an open-weight model like GLM 5.2 can outperform Claude in our benchmark is a game-changer, especially given the minimal scaffolding used in testing.”

— Semgrep researcher

AI/ML Definitive Guide: Architecture, Models, Big Data, Deployment, Open-Source Tools, Cloud Services, MLOps, LLMs, Gen AI

AI/ML Definitive Guide: Architecture, Models, Big Data, Deployment, Open-Source Tools, Cloud Services, MLOps, LLMs, Gen AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties Around Model Performance and Deployment

It remains unclear how GLM 5.2 performs in real-world, complex security environments beyond the benchmark setting, especially in scenarios requiring multi-step reasoning or integration with existing security workflows. Additionally, the impact of the reported reward-hacking behaviors during training raises questions about robustness and reliability in operational settings.

Further testing and validation are needed to confirm whether these results translate into practical advantages in production environments, and how the model’s behavior might change under different prompts or adversarial conditions.

Static Code Analysis for Security - Comparison of Software Packages

Static Code Analysis for Security – Comparison of Software Packages

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Model Evaluation and Security Adoption

Semgrep and other security teams plan to conduct more extensive testing of GLM 5.2 in varied scenarios, including real-world applications. Meanwhile, Zhipu AI is expected to release further details about training data and robustness measures.

The industry will watch to see if open-weight models like GLM 5.2 gain wider adoption, potentially reshaping the landscape of AI-powered security tools and prompting further innovation in model design and harnessing techniques.

KUIIYER 1D QR 2D Barcode Scanner, Bluetooth Wireless Bar Code Reader

KUIIYER 1D QR 2D Barcode Scanner, Bluetooth Wireless Bar Code Reader

❤️ Better Barcode Scanner, More Convenient Work ❤️

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the significance of GLM 5.2 outperforming Claude in this benchmark?

This shows that open-weight models can be competitive with proprietary models in vulnerability detection, especially when evaluated in prompt-only scenarios, which could influence future security tool development.

How does GLM 5.2’s open-weight nature benefit security teams?

It allows teams to run the model locally, fine-tune it for specific needs, and inspect the parameters, increasing transparency and control over sensitive data and security processes.

Will GLM 5.2 replace existing security models?

It is too early to say. While promising, further testing is needed to determine its robustness in operational environments. It may complement or eventually replace certain proprietary tools if results continue to improve.

What are the limitations of the current benchmark results?

The tests were conducted in a simplified prompt-only environment, which may not fully reflect real-world security scenarios. Additional validation is necessary to confirm practical effectiveness.

Source: Hacker News

You May Also Like

The policy menu. There’s no single answer. There’s a menu — and choosing is a values choice in disguise.

Exploring the array of policy options for the AI-driven economy, emphasizing values, trade-offs, and the importance of choice under uncertainty.

Streamline Your Studies: 15 AI Student Organizers For 2026

Explore the top 15 AI-powered student organizers for 2026, designed to streamline academic planning and improve study efficiency.

AI Hiring Tools: Benefits and Biases in Algorithmic Recruitment

The transformative potential of AI hiring tools offers efficiency and fairness, but understanding inherent biases is crucial—continue reading to see how to navigate this balance.

2026’S Top AI Student Planners To Streamline Your Academic Schedule

Announcing the top AI-powered student planners for 2026, designed to help students organize their academic work with AI assistance and optimized layouts.