Discover your interests, together

Real deals, honest reviews and shopping stories from people who share your interests — every day on Milik.

Discover your interests, togetherReal deals, honest reviews and shopping stories from people who share your interests — every day on Milik.

How AI Red Teams Are Rewriting the Rules of Software Security

How AI Red Teams Are Rewriting the Rules of Software Security
Interest|High-Quality Software

AI Red Teams: From Niche Experiment to Security Shockwave

AI red team testing is the use of artificial intelligence systems to simulate attackers, scan large software codebases, and automatically identify, prioritize, and help reproduce security vulnerabilities at a pace and scale that far exceed manual reviews by human security engineers. That core shift—machines testing machines—explains why AI vulnerability detection is already reshaping how software is built and defended. A volunteer Bitcoin security group reviewed 390 open-source projects with AI assistance and submitted 4,962 findings in roughly 30 hours, proving that automated security audits can expose weaknesses that would have taken weeks of human analysis to surface. The uncomfortable truth is that defenders are not the only ones scaling up. A Chinese-speaking threat actor has already run an AI-enabled autonomous hacking campaign against real infrastructure, using multiple known vulnerabilities as entry points.

How AI Red Teams Are Rewriting the Rules of Software Security

Bitcoin Red Team: Thousands of Findings in Hours, Not Weeks

The Bitcoin Red Team shows what happens when a small, motivated group plugs AI into their workflow instead of treating it as a side experiment. This globally distributed team of 16 volunteers used AI models to examine open-source repositories, identify suspicious sections of code and help researchers test possible attack paths. In about 30 hours, they reviewed 390 projects and submitted 4,962 findings, including 85 potential critical vulnerabilities and another 635 high-severity issues—720 high and critical findings in total. Only 21.4% of those had been reproduced at the time of reporting, which underlines both the power and the noise of vulnerability discovery automation. Still, the output is staggering: they were averaging roughly 2.31 high or critical issues per person per hour during the campaign. That is not a marginal efficiency gain; it is a step change in what a small security team can audit.

Autonomous AI Offense: Proof That Attackers Won’t Wait

If you think AI red teams are a defensive luxury, consider the other side of the chessboard. A Chinese-speaking threat actor, operating under the aliases knaithe and KnYuan, has already run an AI-enabled autonomous hacking campaign targeting infrastructure via seven different vulnerabilities across Langflow, n8n, Citrix NetScaler, Apache Tomcat, Marimo Notebook, Palo Alto Networks PAN-OS, and Microsoft Windows IKE Extensions. DeepSeek, connected through the Hermes Agent framework, acted as an autonomous offensive operator—handling reasoning for code generation, vulnerability assessment, target selection and decision-making—while Telegram served as the command-and-control layer for enumerating targets, sourcing exploit tools, and launching attacks without human intervention. When initial exploitation failed because of restrictive configurations, the Hermes Agent did not stop; it automatically searched for critical CVEs, surveyed ten product families, scanned GitHub for trending proofs of concept, and prioritized vulnerabilities by attack surface.

From Coldcard Chaos to Always-On AI Security

The Bitcoin ecosystem learned the hard way why slow audits are no longer acceptable. A serious seed-generation vulnerability in Coldcard hardware wallet firmware weakened randomness and allowed sophisticated attackers to reconstruct private keys, leading to the theft of approximately 1,816 BTC, worth around USD 116 million (approx. RM534,400,000), from more than 5,200 addresses over four attack waves. Firmware updates alone were not enough; funds tied to vulnerable recovery seeds stayed exposed, forcing Coinkite to advise users to update devices, generate entirely new seeds, and move funds to fresh addresses. Against that backdrop, AI red team testing is less a novelty and more a survival strategy. The scale of Bitcoin Red Team’s results could drive urgent disclosures and updates across wallets, exchanges, Lightning applications, mining software and other infrastructure. Their custom AI-driven security harness is expected to be released as open source, so companies can test both public and proprietary code with automated security audits.

What AI Red Teams Change—and What They Don’t

AI vulnerability detection will not magically secure software, but it will change the tempo of the entire ecosystem. Systems that can scan for unsafe assumptions, faulty randomness, access-control mistakes, memory problems and odd interactions between components far faster than humans mean that thousands of potential issues can appear in a single reporting cycle. That flood forces teams to rethink triage, patching windows and how they budget for security work. It also exposes the limits of automation: AI reviews can generate false positives, duplicates and context-free findings that never become real-world exploits. Yet ignoring AI because it is noisy is worse than stubborn; it cedes the advantage to attackers already running autonomous campaigns. The real shift is cultural. Security reviews can no longer be a late-stage box-tick. They have to become a continuous process where AI red team testing runs alongside development, while humans decide what to fix, when, and why.

Milik earns a commission when you shop through our links, at no extra cost to you. This article was generated with AI from published sources and product data.

You May Also Like

Comments
Say something...
No comments yet. Be the first to share your thoughts!