Paul Miller, a developer who maintains cryptography libraries used in Bitcoin and Ethereum projects, tested 7 artificial intelligence (AI) agents to classify and correct 5 real security vulnerabilities reported by the Bitcoin Red Team. The results showed marked differences between the models, both in the dollar cost incurred to run each one and in the accuracy of their diagnoses.
Of the 7 agents evaluated, 3 did not complete the task. Qwen 3.8 from Alibaba got stuck during the process, with no reason stated by Miller. GPT-5.6 Sol from OpenAI and Fable from Anthropic outright refused to work on the code, according to Miller himself, despite the task being to identify and fix flaws, not to exploit them.
The 4 agents that did complete the work, Grok 4.6 (xAI), Kimi K3 (Moonshot AI), DeepSeek V4 Pro, and DeepSeek V4 Flash (DeepSeek), had very different dollar costs for token consumption (user instructions) when executed:
The time taken also varied considerably. DeepSeek V4 Pro completed its work in 30 minutes, Kimi K3 in 40 minutes, and Grok 4.6 in 1 hour, while DeepSeek V4 Flash took 3 hours, six times (500%) longer than DeepSeek V4 Pro for the same task.
According to Miller, only Grok 4.6 matched his own assessment of the severity of the 5 vulnerabilities.
The other 3 agents that completed the work (Kimi K3, DeepSeek V4 Pro, and DeepSeek V4 Flash) exaggerated the severity of the flaws found, something that, according to Miller himself, ultimately affected the overall quality of the results delivered by those 3 models.
Miller's test comes amid a wave of attacks that use artificial intelligence to accelerate the search for and exploitation of security flaws in the Bitcoin ecosystem.
In recent weeks, Coldcard, Boltz, ZEUS, BTCPay Server, and LNP2Pbot suffered security incidents linked to flaws that attackers identified and exploited with the help of AI models. None of these cases affected the Bitcoin protocol itself, but rather products and services operating outside the network.
This wave of incidents prompted the Bitcoin Red Team, a group of developers dedicated to finding security flaws in open-source Bitcoin projects, to launch a massive AI-assisted audit of the ecosystem.
The Bitcoin Red Team audit has already covered more than 300 open-source repositories, and from that work emerged the 5 vulnerabilities that Miller used as the basis for his own test on the 7 AI agents, among another 8,000 flaws found by that group of researchers, as reported by CriptoNoticias.
Of the 7 agents that Miller tested, 4 are open-source, Kimi K3, DeepSeek V4 Pro, DeepSeek V4 Flash, and Qwen 3.8. This means that the companies that trained them publicly released their parameters, allowing any user to download and run them on their own machines without relying on the approval of those companies.
The other 3 agents, Grok 4.6, GPT-5.6 Sol, and Fable, are closed, so they can only be used through the platform or the application programming interface (API) of the company that developed them.
This distinction between open-source and closed models is relevant to the security of the Bitcoin ecosystem. By running on the companies' own servers, closed models allow their creators to maintain active security filters, such as those that led GPT-5.6 Sol and Fable to refuse to work on the vulnerabilities in this test.
Open-source models, on the other hand, can be run on a local computer and modified by any user, making it possible to remove those security filters. An attacker could use an unrestricted version of one of these same models to assist real attacks against platforms in the Bitcoin ecosystem, something that no external filter could prevent.
Miller's test exposes that, beyond the advancement of artificial intelligence in detecting vulnerabilities, significant differences persist between the various models when it comes to classifying and correcting them accurately, both in the dollar cost incurred by each and in the criteria they apply.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





























