AkitaOnRails Blog
Featured
Jul 30— New LLM Benchmark: I Reran Every Test!Jul 20— What's New in My AI-MEMORY: Switch AI Agents Without Losing the SessionJul 12— Quantum News: Majorana 2 and Understanding ShorJul 12— Using AI to Solve My Little Day-to-Day ProblemsJul 11— How Do I Protect Myself From My Agents Deleting My Stuff?Jun 24— Why LLMs Will Fail at Your CompanyJun 05— AI Controversy in Open Source Project Contributions - My TakeMay 30— Open Source Best Practices with LLMs - The Bare MinimumApr 20— Clean Code for AI AgentsApr 11— VS Code Is the New Punch CardFeb 24— RANT: Did Akita Bend Over for AI??Feb 16— Vibe Code: From Zero to Production in 6 DAYS | The M.Akita ChroniclesFeb 08— RANT: Did AI Kill Programmers?Jun 18— AGI or Skynet Isn't Coming Anytime SoonMay 02— RANT - LLMs are LOOT BOXES!
New LLM Benchmark: I Reran Every Test!
I reran the LLM Coding Benchmark with a harder test, three phases, native harnesses, and new tiers. Fable led, Terra tied Sol as the best GPT, and 24 models proved suitable for programming.
What's New in My AI-MEMORY: Switch AI Agents Without Losing the Session
ai-memory run keeps the same programming session while switching among Claude Code, Codex, and other harnesses, with searchable workstreams and integration with ai-jail and ai-usagebar.
Quantum News: Majorana 2 and Understanding Shor
Microsoft announced Majorana 2 with 20-second qubits and physicists answering that nothing was resolved. I take the chance to properly explain Shor's algorithm, why factoring becomes period finding, and what quantum computers are actually good at.
Using AI to Solve My Little Day-to-Day Problems
A roundup of my small open source projects: a desk clock with widgets, manga readers, decent email, typing practice, YouTube karaoke, ComfyUI in Docker and more. All born from real little problems in my day-to-day.
How Do I Protect Myself From My Agents Deleting My Stuff?
LLMs deleting files from famous people made headlines this week. In five months of heavy use, in YOLO mode, it never happened to me. But I don't trust them either: BTRFS snapshots, restic backups, sandboxing and discipline.
Why LLMs Will Fail at Your Company
The deliberately speculative thesis is that LLMs amplify processes where people outsource decisions. The proposed alternative is daily cycles of decision, implementation, testing, and review close to the outcome.
AI Controversy in Open Source Project Contributions - My Take
AI in open source contributions is here to stay, even with AI slop, regressions, and false bugs. The practical answer is to automate triage and auditing without taking the final decision away from humans.
Open Source Best Practices with LLMs - The Bare Minimum
Open source projects built with LLMs are only ready with simple installation, reliable tests and CI, and documentation focused on the problem. Standardizing releases and deployments makes automation predictable.
Clean Code for AI Agents
Clean Code carries different weight when an agent is the primary reader: small code, greppable names, provenance context, headless tests, and explicit rules reduce navigation, cost, and errors.
VS Code Is the New Punch Card
In the era of coding agents, typing everything by hand in VS Code is becoming the modern equivalent of a punch card. Foundations did not become legacy.
RANT: Did Akita Bend Over for AI??
I have been 'bending over' for AI since 2023. If you only watch out-of-context podcast clips, let me spell it out for you.
Vibe Code: From Zero to Production in 6 DAYS | The M.Akita Chronicles
How I shipped a full newsletter, blog, podcast, and Discord bot to production in six days using vibe coding the right way.
RANT: Did AI Kill Programmers?
Akita's rant on why AI won't replace real programmers, why the programming bubble already popped, the S-curve ceiling of current LLM architecture, and what actually changes for engineers in the LLM era.
AGI or Skynet Isn't Coming Anytime Soon
Why the current LLM architecture cannot reach AGI, and why the hype from Big Tech CEOs and media-anointed experts is mostly backroom negotiation noise.
RANT - LLMs are LOOT BOXES!
Why LLMs for real coding behave like gacha loot boxes, and why every incentive in the industry pushes you to burn more tokens.
2026 - August
2026 - July
- URGENT - If You Keep Bitcoin on a ColdCard: MOVE EVERYTHING
- Removing DRM from Kindle Ebooks in 2026
- New LLM Benchmark: I Reran Every Test!
- AI-Jail: Security Update, Docker Goes Opt-In
- LLM Benchmark: Is Opus 5 Any Good?
- What's New in My AI-MEMORY: Switch AI Agents Without Losing the Session
- LLM Benchmark: Should I Use the Highest-Scoring Model?
- LLM Benchmark: Has Kimi K3 Reached Claude Opus Level?
- Quantum News: Majorana 2 and Understanding Shor
- Using AI to Solve My Little Day-to-Day Problems
- How Do I Protect Myself From My Agents Deleting My Stuff?
- The Bun in Rust Response Andrew Kelly Should Have Written
- LLM Benchmarks: Grok 4.5 and GPT 5.6 Sol
- I Had Fable 5 Analyze the Code of TikTok, Clash of Kings and Gov.br - Understanding Fingerprinting
- Frank GO: Playing Go with AI
- LLM Benchmark: Sonnet 5 Fails, Gemini Flash Surprises, Sakana Fugu Almost Reaches Tier A
- Off-Topic: Do Not Put Hope in Politics - How to Become Immune
2026 - June
- Why LLMs Will Fail at Your Company
- ai-memory: long-term memory (Karpathy Wiki) and self-improvement (Hermes) for your projects
- The Rio 3.5 LLM Controversy. Plagiarism?
- ai-memory: Emergent Architecture and Malleable Software
- LLM Benchmark: Kimi v2.7 Code, GLM 5.2, MiniMax M3 Local
- Bypassing the GitHub API Block in Brazil
- LLM Benchmark: Fable 5 and the Anthropic Soap Opera
- Playing with TUIs and LLMs - Ratatui+Bubbletea
- AI Controversy in Open Source Project Contributions - My Take
- LLM Benchmarks - An Update on Grok 4.3, MiniMax v3, and Opus 4.8
2026 - May
- Open Source Best Practices with LLMs - The Bare Minimum
- Complete Manga Solution: Frank Manga+, Frank Yomik, and the Prettify-Manga Extension
- Backing Up Gmail to Maildir on Linux
- First Impressions Using Oh-My-Pi and OpenCode
- Akita's AI Tips and Toolkit: ai-jail, ai-memory, ai-usagebar
- I Built a Waybar Widget for Omarchy to Monitor LLM Plan Usage: ai-usagebar
- I Built a CLI to Check My GitHub Pending Stuff: ghpending
- I Built a Memory System for Coding Agents: ai-memory
- AI Agent Memory: Karpathy LLM Wiki and agentmemory in Practice
- Wrapping Up My AI Marathon: Success or Failure?
- LLM Benchmarks: DeepSeek Unlocked! Use DeepClaude
- NW-Omarchy: Bringing Omarchy to X11 with XLibre
2026 - April
- LLM Benchmarks: Is It Worth ($$) Mixing 2 Models? (Planner + Executor)
- LLM Coding Benchmark (May 2026): DeepSeek v4, Kimi v2.6, Grok 4.3, GPT 5.5
- How DriveClub and shadPS4 Almost Defeated AI and Me: How to Learn
- Clean Code for AI Agents
- My Favorite Retro Racing Games Running on My Distrobox
- LLM Benchmarks Part 2: Is It Worth Combining Multiple Models in the Same Project? Claude + GLM??
- Omarchy on the Thinkpad T14 Gen 6: Mini-Review and Full Setup
- Why LLMs Aren't Giving You the Result You Expect | Why I Prefer Claude Code Today
- Seedance 2.0 Is Finally Out: First Impressions
- An Emulation Distrobox with Claude Code
- VS Code Is the New Punch Card
- How ElevenLabs Was Not Killed by Qwen3 TTS
- 20 Years of Blogging: Translating Everything to English
- Is RAG Dead? Long Context, Grep, and the End of the Mandatory Vector DB
- Testing Open Source and Commercial LLMs - Can Anyone Beat Claude Opus?
- Turning YouTube into a Karaoke App | Frank Karaoke
- Bitcoin on the Home Server: Sovereignty and Privacy with Coldcard, Sparrow and Fulcrum
- My Sim Racing Cockpit - Formula FX1
2026 - March
- Claude Code's Source Code Leaked. Here's What We Found Inside.
- Migrating my Home Server with Claude Code | openSUSE MicroOS
- Review: Minisforum MS-S1 Max | AMD AI Max+ 395 with 96GB of VRAM
- Teaching People to Question the News | Frank Investigator
- I Rewrote OpenClaw in Rust. Did It Work? | FrankClaw
- Going After Email Fraud | Frank FBI
- Porting 10K Lines of Python to Crystal with Claude: easy-subtitle
- Crystal and a Smart FFmpeg Wrapper Built in 3 Hours | easy-ffmpeg
- 37 Days of Vibe Coding Immersion: Conclusions on Business Models
- My First Vibe Code Failure and How I Fixed It | Frank Yomik
- I Built a Data Mining System for My Influencer Girlfriend — Tips and Tricks
- ai-jail: Sandbox for AI Agents — From Shell Script to Real Tool
- Software Is Never 'Done' — 4 Projects, Life After Deploy, and Why One-Shot Prompting Is a Myth
2026 - February
- RANT: Did Akita Bend Over for AI??
- Vibe Code: I Built a Smart Image Indexer with AI in 2 Days | Frank Sherlock
- Vibe Code: I Built a Mega Clone in Rails in 1 Day for My Home Server
- From Zero to Post-Production in 1 Week - How to Use AI on Real Projects | Behind The M.Akita Chronicles
- Integration Tests in a Monorepo | Behind The M.Akita Chronicles
- SQLite + Kamal: Rails Deploy Without Drama | Behind The M.Akita Chronicles
- Discord as an Admin Panel | Behind The M.Akita Chronicles
- Async Jobs That Survive Chaos | Behind the Scenes of The M.Akita Chronicles
- Frontend Without a Framework | Behind the Scenes of The M.Akita Chronicles
- Serving AI in the Cloud: My Personal TTS | Behind the Scenes of The M.Akita Chronicles
- Web Scraping in 2026 | Behind the Scenes of The M.Akita Chronicles
- Sending Emails Without Getting Flagged as Spam | Behind The M.Akita Chronicles
- Vibe Code: From Zero to Production in 6 DAYS | The M.Akita Chronicles
- AI Agents: What Would Be the Best Programming Language for LLMs?
- RANT: Did AI Kill Programmers?
- Vibe Code: I Built a Markdown Editor From Scratch With Claude Code (FrankMD) PART 2
- Vibe Code: I Built a Markdown Editor From Scratch With Claude Code (FrankMD) PART 1
2026 - January
- Vibe Code: Which LLM Is the BEST?? Let's Talk for REAL
- Vibe Code: I Built a Little App 100% with GLM 4.7 (TV Clipboard)
- AI Agents: Which One Is Best? OpenCode, Crush, Claude Code, GPT Codex, Copilot, Cursor, Windsurf, Antigravity?
- AI 3D: Can You Actually Model 3D With Prompts Now?
- AI Agents: Is GLM 4.7 Flash really that good?
- Omarchy 3: Dual GPU Setup with AMD and NVIDIA
- AI Agents: Installing LSPs for Crush
- AI Agents: Comparing the Top LLMs of 2026 on the Zig Challenge
- AI Agents: Locking Down Your System
- Omarchy 3 - One of the Best Coding Agents Out There: Crush
2025 - September
- Omarchy 2.0 - LazyVim Basics
- Omarchy 2.0 - Recommended for Beginners?
- Omarchy 2.0 - Install with the Omarchy ISO
- Omarchy 2.0 - Bitwarden Self-Hosted / VaultWarden
- Installing Grafana on My Home Server
- Protecting Your Home Server with Cloudflare Zero Trust
- Accessing My Home Server With a Real Domain
- Omarchy 2.0 - Understanding SSH and Yubikeys
- Omarchy 2.0 - TUIs (Terminal User Interface Apps)
- Omarchy 2.0 - Mise for Organizing Development Environments
- Omarchy 2.0 - LazyVim - LazyExtras
- Omarchy 2.0 - ZSH Configs
2025 - August
- How to Contribute to the AkitaOnRails Blog Using Docker
- Installing Omarchy 2.0 from Scratch - Personal Notes
2025 - June
2025 - May
- Your Windows May Be Crippled Without You Knowing. Check This!!
- Computing History and Retro Dev on YouTube
- Final Attempt to Train an LLM with LoRA. Cannon Shot, But Missing the Fly.
- Teaching the Latest Zig to Your LLM - Training LoRAs (Sort Of)
- RANT - LLMs are LOOT BOXES!
- When Do LLMs Fail at Programming? A More Realistic Use Case.
- Rant - Will LLMs Evolve Forever? Demystifying LLMs in Programming
2025 - April
- Dissecting an Ollama Modelfile - Tuning Qwen3 for Code
- Testing the Newly Released Open Source LLM - Qwen3 (with Aider and Ollama)
- Destroying the ChatGPT 4o "Personality"
- Testing LLMs with Aider on RunPod - which one to use for code?
- Your Own Free Universal Co-Pilot Running Local: AIDER-OLLAMA-QWEN
- LLM Hello World: Building Your Own Local AI Chat
- Accessing Your NAS Using iSCSI Instead of SMB
- Changing Clothes Using A.I. (ComfyUI)
- Using A.I. (ComfyUI) to Generate NPCs in Game Development
- Understanding the Basics of ComfyUI to Generate AI Images
- Generating AI Images - even Ghibli style 😂 - with Docker and CUDA
- Generating Up to 2-Minute Videos from a Photo with A.I.
- Upscaling Old Anime to 4K with AI
- Colorizing Black and White Images with A.I.
- Configuring My Synology NAS with NFS on Linux
- NVIDIA and Wayland - Problems for PCI Passthrough in VMs
- BIOS Configuration of My PC - X670E Aorus Xtreme
2026 - August
1 postExploiting Coinkite's RNG Egregious Problem
How an attacker enumerates the ColdCard's reduced key space, finds vulnerable wallets on the public blockchain, and moves the funds. Real data from the ongoing theft and step-by-step didactic code.
2026 - July
17 postsURGENT - If You Keep Bitcoin on a ColdCard: MOVE EVERYTHING
An entropy flaw left seeds generated by ColdCard firmware far below the promised security level, with estimated losses above 1,000 BTC. Understand the bug and migrate without repeating the mistake.
Removing DRM from Kindle Ebooks in 2026
I used Calibre, DeDRM 10.0.28, KFX Input, and the Microsoft Store Kindle app in an Omarchy VM to archive 106 books and validate an EPUB conversion workflow I control.
New LLM Benchmark: I Reran Every Test!
I reran the LLM Coding Benchmark with a harder test, three phases, native harnesses, and new tiers. Fable led, Terra tied Sol as the best GPT, and 24 models proved suitable for programming.
AI-Jail: Security Update, Docker Goes Opt-In
Issue #88 proved the Docker socket inside ai-jail gave any agent root on the host. In v1.16.0 the passthrough went opt-in. The flaw, a hands-on demo, best practices, and why Podman was born from this criticism.
LLM Benchmark: Is Opus 5 Any Good?
Opus 5 scored 95/100 on the Rails benchmark, tying Opus 4.8 and edging Fable 5 by one point. Its API rate is half Fable’s, but this narrow test does not define the best LLM.
What's New in My AI-MEMORY: Switch AI Agents Without Losing the Session
ai-memory run keeps the same programming session while switching among Claude Code, Codex, and other harnesses, with searchable workstreams and integration with ai-jail and ai-usagebar.
LLM Benchmark: Should I Use the Highest-Scoring Model?
Why the top entry in an LLM ranking is not necessarily the best model, why 90+ is a cluster, and how to use benchmarks without outsourcing your technical judgment.
LLM Benchmark: Has Kimi K3 Reached Claude Opus Level?
I had Kimi K3 build a chat in Rails 8 on its own: it scored 89/A, beat Opus 4.6, and was a cheaper alternative to Opus 4.8. It still lagged behind in architecture, hardening, and tests.
Quantum News: Majorana 2 and Understanding Shor
Microsoft announced Majorana 2 with 20-second qubits and physicists answering that nothing was resolved. I take the chance to properly explain Shor's algorithm, why factoring becomes period finding, and what quantum computers are actually good at.
Using AI to Solve My Little Day-to-Day Problems
A roundup of my small open source projects: a desk clock with widgets, manga readers, decent email, typing practice, YouTube karaoke, ComfyUI in Docker and more. All born from real little problems in my day-to-day.
How Do I Protect Myself From My Agents Deleting My Stuff?
LLMs deleting files from famous people made headlines this week. In five months of heavy use, in YOLO mode, it never happened to me. But I don't trust them either: BTRFS snapshots, restic backups, sandboxing and discipline.
The Bun in Rust Response Andrew Kelly Should Have Written
I used Claude Fable 5 and GPT 5.6 Sol as judges to compare Rust and Zig in Bun’s rewrite. Rust helps reduce ownership bugs and maintenance costs, but the agent-based method mattered just as much as the language.
LLM Benchmarks: Grok 4.5 and GPT 5.6 Sol
Grok 4.5 entered Tier A with 87, and GPT 5.6 Sol scored 92. Sol beat GPT 5.5 by 92 to 81 in the blind A/B test, but re-auditing brought it down to 85. With cache-read fixed, Opus 4.8 remains the best API balance.
I Had Fable 5 Analyze the Code of TikTok, Clash of Kings and Gov.br - Understanding Fingerprinting
A static analysis with Fable 5 compares TikTok’s opaque fingerprint, Clash of Kings’ persistent UUID, and gov.br’s privacy precautions, with caveats about the method.
Frank GO: Playing Go with AI
I built Frank GO as a Sabaki fork with KataGo, bringing together more than 4,700 tsumego, historical games, and 15 of Hikaru’s Go games. Overlays make territory and scoring visible to beginners.
LLM Benchmark: Sonnet 5 Fails, Gemini Flash Surprises, Sakana Fugu Almost Reaches Tier A
In the coding benchmark, Sonnet 5 scored 58/100 after calling a nonexistent API, while Gemini 3.5 Flash scored 93/100. Sakana Fugu Ultra landed at 79/100, close to Tier A.
Off-Topic: Do Not Put Hope in Politics - How to Become Immune
I argue for treating politics as environmental risk, not religion: an election isn’t a life plan. The proposed strategy is to increase income in hard currency, understand taxes, and organize assets within the law.
2026 - June
10 postsWhy LLMs Will Fail at Your Company
The deliberately speculative thesis is that LLMs amplify processes where people outsource decisions. The proposed alternative is daily cycles of decision, implementation, testing, and review close to the outcome.
ai-memory: long-term memory (Karpathy Wiki) and self-improvement (Hermes) for your projects
ai-memory turns agent sessions into a searchable Markdown wiki with hooks, MCP, and handoffs between tools. A Hermes-inspired loop promotes memories with validation, evidence, and optional review.
The Rio 3.5 LLM Controversy. Plagiarism?
Public evidence indicates that Rio 3.5’s initial checkpoint was a blend of Nex-N2-Pro and Qwen, without credit to Nex. The case points to an attribution failure, not legal plagiarism.
ai-memory: Emergent Architecture and Malleable Software
In 24 days, contributors took ai-memory from a personal MVP to a cross-platform, multi-user system. The architecture emerged from use and was only consolidated afterward.
LLM Benchmark: Kimi v2.7 Code, GLM 5.2, MiniMax M3 Local
In the benchmark, GLM 5.2 scored 87 and Kimi K2.7 Code scored 86. The open MiniMax M3 doesn’t fit in 128 GB, while serious programming still calls for Opus 4.8 or GPT 5.5.
Bypassing the GitHub API Block in Brazil
I tested GitHub’s API block in Brazil: changing DNS didn’t fix it on my connection, but a SOCKS5h wrapper with Tor made gh and ghpending work. It’s a workaround for an opaque block.
LLM Benchmark: Fable 5 and the Anthropic Soap Opera
Fable 5 scored 94/100, practically tied with Opus 4.8, but costs ten times more and redirects sensitive tasks. Since Mythos isn’t auditable, I’m sticking with 4.8.
Playing with TUIs and LLMs - Ratatui+Bubbletea
I built ratatui-bubbletea to give Ratatui Bubble Tea-inspired themes and components using Rust and LLMs. The toolkit already powers real TUIs like ai-usagebar without replacing the base library.
AI Controversy in Open Source Project Contributions - My Take
AI in open source contributions is here to stay, even with AI slop, regressions, and false bugs. The practical answer is to automate triage and auditing without taking the final decision away from humans.
LLM Benchmarks - An Update on Grok 4.3, MiniMax v3, and Opus 4.8
In the Rails 8 benchmark, Opus 4.8 keeps the lead at 95/100, while Grok 4.3 and MiniMax M3 finally become usable but remain in Tier B, behind GPT and Opus.
2026 - May
12 postsOpen Source Best Practices with LLMs - The Bare Minimum
Open source projects built with LLMs are only ready with simple installation, reliable tests and CI, and documentation focused on the problem. Standardizing releases and deployments makes automation predictable.
Complete Manga Solution: Frank Manga+, Frank Yomik, and the Prettify-Manga Extension
Frank Manga+ brings the MANGA Plus catalog to the desktop, while Prettify improves reading on websites and Kindle, and Frank Yomik adds furigana or translation to Japanese and Korean pages.
Backing Up Gmail to Maildir on Linux
I replaced manual Thunderbird backups with mbsync in pull mode, Maildir, and a systemd timer. Restic replicates the archive to NAS and off-site storage, keeping emails local, open, and accessible to multiple clients.
First Impressions Using Oh-My-Pi and OpenCode
Oh-My-Pi is more flexible for projects with varied files and artifacts, while OpenCode is more polished for ordinary code. The test showed that no harness audits everything, and Codex remains behind.
Akita's AI Tips and Toolkit: ai-jail, ai-memory, ai-usagebar
After more than 600 hours with coding agents, Akita brings together tools like ai-jail, ai-memory, ai-usagebar, and ghpending, and argues for testing, backups, refactoring, and disciplined engineering.
I Built a Waybar Widget for Omarchy to Monitor LLM Plan Usage: ai-usagebar
ai-usagebar, a Rust port of claudebar, brings Claude, Codex, Z.AI, and OpenRouter into a Waybar widget or standalone TUI that works in any terminal, including over SSH.
I Built a CLI to Check My GitHub Pending Stuff: ghpending
Akita built ghpending to query issues and PRs across multiple repositories in a Rust digest, fetching everything in parallel and raising GitHub’s limit from 60 to 5,000 requests per hour with a token.
I Built a Memory System for Coding Agents: ai-memory
After finding reindexing, data loss, and broken hooks in agentmemory, Akita built ai-memory in Rust with Markdown, SQLite/FTS5, automatic capture, and agent handoff. It’s still beta.
AI Agent Memory: Karpathy LLM Wiki and agentmemory in Practice
Akita compares compaction in Claude Code, Codex, and opencode, explains Karpathy’s LLM Wiki, then tests agentmemory before dropping it over structural bugs and pointing to ai-memory.
Wrapping Up My AI Marathon: Success or Failure?
After more than 500 hours across 24 repositories, Akita ranks his experiments and concludes Codex and Claude Code support daily coding, but without review and XP, productivity turns into slop.
LLM Benchmarks: DeepSeek Unlocked! Use DeepClaude
I swapped opencode’s harness for Claude Code’s loop through DeepClaude, and DeepSeek V4 Pro rose from 69/B to 89/A in 18 minutes and $3.14. The controller’s multi-turn bug remains.
NW-Omarchy: Bringing Omarchy to X11 with XLibre
I built NW-Omarchy as an X11 session alongside Omarchy, using XLibre, bspwm, and equivalents to preserve themes and 70-plus shortcuts. HDR and fractional scaling stay in Hyprland.
2026 - April
18 postsLLM Benchmarks: Is It Worth ($$) Mixing 2 Models? (Planner + Executor)
Three rounds show multi-agent doesn’t beat solo Opus 4.7 in opencode, which scored 97/100 in 18 minutes for about $4. GPT 5.4 xHigh with a medium executor saves money but loses 3 points.
LLM Coding Benchmark (May 2026): DeepSeek v4, Kimi v2.6, Grok 4.3, GPT 5.5
I reaudited 24 LLMs on the same Rails app with RubyLLM: Opus 4.7 and GPT 5.4 tie at 97, GPT 5.5 reaches 96, Kimi and Gemini become Tier A, and DeepSeek V4 Pro reaches Tier A only through DeepClaude.
How DriveClub and shadPS4 Almost Defeated AI and Me: How to Learn
After 31 phases and 44 commits, I found DriveClub’s nighttime blackout in shadPS4 came from unsynchronized GPU→CPU feedback. `readbacks_mode: 2` fixed the bug, but the investigation taught me the path.
Clean Code for AI Agents
Clean Code carries different weight when an agent is the primary reader: small code, greppable names, provenance context, headless tests, and explicit rules reduce navigation, cost, and errors.
My Favorite Retro Racing Games Running on My Distrobox
I consolidated a collection of classic racing games on Linux in Distrobox. Driveclub runs at 30 FPS with a shadPS4 fork, and FM4 works in Xenia after extensive tweaking.
LLM Benchmarks Part 2: Is It Worth Combining Multiple Models in the Same Project? Claude + GLM??
I tested seven combinations in Claude Code, opencode, and Codex, but no sub-agent was called. Variance across runs and harnesses keeps solo Opus as the default for greenfield projects.
Omarchy on the Thinkpad T14 Gen 6: Mini-Review and Full Setup
I installed Omarchy on a ThinkPad T14 Gen 6 and turned the laptop into a Linux companion for SSH, Claude Code, and NAS. The screen feels dated, but ports, durability, and support make up for it.
Why LLMs Aren't Giving You the Result You Expect | Why I Prefer Claude Code Today
Every week somebody tells me 'I canceled my Claude plan, it just doesn't perform as well as GPT for me.' I've got 500+ hours in Claude Code and Codex, 400k lines generated, and both deliver. The difference isn't the model. It's how people are talking to it.
Seedance 2.0 Is Finally Out: First Impressions
ByteDance opened Seedance 2.0 to the public today after months of restricted access. I tested audio-driven lip sync with my anime avatar, and a Blender render fed in as a video reference. There's real work you can do here, but it's nowhere near professional production yet. Also: deepfakes just stopped being hypothetical.
An Emulation Distrobox with Claude Code
I built an Arch Distrobox for 12 emulation platforms and Steam, organized into 17 Ansible roles with Claude Code’s help. The automation makes setup reproducible and keeps decisions visible.
VS Code Is the New Punch Card
In the era of coding agents, typing everything by hand in VS Code is becoming the modern equivalent of a punch card. Foundations did not become legacy.
How ElevenLabs Was Not Killed by Qwen3 TTS
When Qwen3 TTS dropped, half the internet called it an 'ElevenLabs killer'. I spent weeks trying to run Qwen3 in production on my podcast. Yesterday I finally switched to ElevenLabs v3. Less than a day later, I can tell you: open source is still miles behind.
20 Years of Blogging: Translating Everything to English
Four days ago I hit 20 years of blogging. When I sat down to write the anniversary post, I ended up doing something I never had the bandwidth for: translating the whole blog to English. With Claude Code, over a weekend.
Is RAG Dead? Long Context, Grep, and the End of the Mandatory Vector DB
Frontier models went from 200k to 1M tokens of context. Does it still make sense to wire up a whole vector DB stack to do RAG, when you can just grep and dump the entire document into the window?
Testing Open Source and Commercial LLMs - Can Anyone Beat Claude Opus?
This historical benchmark compared 33 LLMs on a Rails app: Opus, Sonnet, and GLM 5 worked, while Qwen 3.6 35B came close on an RTX 5090 after a fix. The rankings were later revised.
Turning YouTube into a Karaoke App | Frank Karaoke
I built a Flutter Android app that overlays live scores on YouTube karaoke videos, using audio filtering and YIN in four modes. It’s fun, but scoring is still experimental.
Bitcoin on the Home Server: Sovereignty and Privacy with Coldcard, Sparrow and Fulcrum
I built a home Bitcoin stack with Coldcard, Sparrow, Fulcrum, and bitcoind: offline keys, private queries, and my own broadcast. It takes work, but gives me more control over custody and operations.
My Sim Racing Cockpit - Formula FX1
I replaced years of unstable stands with a Formula FX1 cockpit featuring Fanatec direct drive, PS5, RTX 4090, and a 4K OLED. It doesn’t wobble and is ready in 30 seconds for my single-player use.
2026 - March
13 postsClaude Code's Source Code Leaked. Here's What We Found Inside.
I analyzed Claude Code’s leaked source map and found hidden features, layered memory, multi-agents, and DRM based on xxHash64. The code also exposed a product that’s difficult to maintain.
Migrating my Home Server with Claude Code | openSUSE MicroOS
I migrated 49 containers from Ubuntu to openSUSE MicroOS with Claude Code’s help, reorganizing Docker, NFS, ROCm, backups, and SELinux. Automation sped things up, but architecture and validation stayed human.
Review: Minisforum MS-S1 Max | AMD AI Max+ 395 with 96GB of VRAM
In my tests, the RTX 5090 was up to 7 times faster on models that fit within 32GB. The Minisforum, however, ran 50-to-81GB models the card can’t handle.
Teaching People to Question the News | Frank Investigator
I built Frank Investigator to compare coverage, track omitted sources, and analyze news in 15 stages with three LLMs. The report flags omissions and reframing, but doesn’t issue final verdicts.
I Rewrote OpenClaw in Rust. Did It Work? | FrankClaw
I rewrote OpenClaw’s core in Rust as FrankClaw, with 56,586 lines, 7 channels, and fixes for the 7 audited critical vulnerabilities. It handles simple conversations, but complex workflows remain untested.
Going After Email Fraud | Frank FBI
I built Frank FBI to analyze forwarded emails in six layers, combining authentication, reputation, OSINT, and three LLMs. The result is a self-hosted risk report, with the data under your control.
Porting 10K Lines of Python to Crystal with Claude: easy-subtitle
I asked Claude to port Subservient to Crystal with feature parity and got easy-subtitle in under 40 minutes: 2,516 lines, 76 tests, and a 6 MB static binary with no Python runtime.
Crystal and a Smart FFmpeg Wrapper Built in 3 Hours | easy-ffmpeg
I wanted to convert videos without memorizing FFmpeg flags. Out came a smart CLI in Crystal with presets, interactive mode and static compilation for Linux and macOS, in 3 hours of vibe coding.
37 Days of Vibe Coding Immersion: Conclusions on Business Models
After 37 days, 650+ commits, ~144K lines of code and nearly 10 published projects, my conclusion on what vibe coding means for the future of startups and business models.
My First Vibe Code Failure and How I Fixed It | Frank Yomik
How I spent days building a manga speech bubble detection system with OpenCV, threw it all out and rebuilt it in hours with a pretrained model. Lessons on vibe coding, productive failure, and knowing when to change course.
I Built a Data Mining System for My Influencer Girlfriend — Tips and Tricks
How I built a Rails 8 + SQLite + Discord bot data mining system for my influencer girlfriend in 3 days, 58 commits, with 40+ tools accessible via LLM tool calling.
ai-jail: Sandbox for AI Agents — From Shell Script to Real Tool
ai-jail is a Rust tool that wraps bubblewrap (Linux) and sandbox-exec (macOS) to safely run AI coding agents like Claude Code, Codex, OpenCode and Crush in a sandbox.
Software Is Never 'Done' — 4 Projects, Life After Deploy, and Why One-Shot Prompting Is a Myth
125 post-production commits across 4 projects in 10 days. Real bugs, real users, real iteration. Why one-shot prompting is a myth and software is never 'done'.
2026 - February
17 postsRANT: Did Akita Bend Over for AI??
I have been 'bending over' for AI since 2023. If you only watch out-of-context podcast clips, let me spell it out for you.
Vibe Code: I Built a Smart Image Indexer with AI in 2 Days | Frank Sherlock
How I built Frank Sherlock, a local AI-powered image catalog desktop app, in a weekend using Agile Vibe Coding with Claude Code.
Vibe Code: I Built a Mega Clone in Rails in 1 Day for My Home Server
Building FrankMega, a self-hosted Mega.nz clone in Rails 8, in a single day with Claude Code, and why experience still matters in vibe coding.
From Zero to Post-Production in 1 Week - How to Use AI on Real Projects | Behind The M.Akita Chronicles
How 274 commits in 8 days shipped a full newsletter, podcast, blog, and Discord bot to production using Extreme Programming with Claude Code as the pair.
Integration Tests in a Monorepo | Behind The M.Akita Chronicles
How a monorepo running three apps that share a filesystem uses a dedicated integration environment, DevCache, rsynced production data, and preflight checks to catch the bugs unit tests never see.
SQLite + Kamal: Rails Deploy Without Drama | Behind The M.Akita Chronicles
How Rails 8 with SQLite and Kamal lets you run a full production app on a $12/month VPS with zero external services.
Discord as an Admin Panel | Behind The M.Akita Chronicles
How I replaced the classic Rails admin panel with a Discord bot: parser/dispatcher patterns, embeds, reactions, tool-calling LLMs, and the lessons from failing silently in production.
Async Jobs That Survive Chaos | Behind the Scenes of The M.Akita Chronicles
Production-grade patterns for Rails 8 ActiveJob and SolidQueue: retries, distributed locks, atomic email claiming, orchestrators, cron safety nets, and status notifications.
Frontend Without a Framework | Behind the Scenes of The M.Akita Chronicles
How I built a modern, responsive, dark-mode, email-compatible, multi-platform frontend for The M.Akita Chronicles without a single line of JavaScript framework.
Serving AI in the Cloud: My Personal TTS | Behind the Scenes of The M.Akita Chronicles
Behind the scenes of putting Qwen3-TTS in production on a serverless GPU to auto-generate a weekly podcast: cold starts, voice cloning, sampling, loudness, and prompt engineering.
Web Scraping in 2026 | Behind the Scenes of The M.Akita Chronicles
How I went from a four-line HTTP GET to packing a full Chromium into Docker just to read the news for my newsletter in 2026.
Sending Emails Without Getting Flagged as Spam | Behind The M.Akita Chronicles
How to build a reliable newsletter sending pipeline on Amazon SES, covering atomic claiming, terminal states, DKIM/SPF/DMARC, List-Unsubscribe and the silent SES suppression list.
Vibe Code: From Zero to Production in 6 DAYS | The M.Akita Chronicles
How I shipped a full newsletter, blog, podcast, and Discord bot to production in six days using vibe coding the right way.
AI Agents: What Would Be the Best Programming Language for LLMs?
A thought experiment on what a programming language designed for LLM coding agents, rather than human programmers, would look like.
RANT: Did AI Kill Programmers?
Akita's rant on why AI won't replace real programmers, why the programming bubble already popped, the S-curve ceiling of current LLM architecture, and what actually changes for engineers in the LLM era.
Vibe Code: I Built a Markdown Editor From Scratch With Claude Code (FrankMD) PART 2
Part 2 of the FrankMD story: refactoring, i18n, syntax highlighting, tests, and the hard-earned lessons from 30 hours of vibe coding with Claude.
Vibe Code: I Built a Markdown Editor From Scratch With Claude Code (FrankMD) PART 1
How I vibe-coded FrankMD, a self-hosted Markdown editor tailored to my blogging workflow, in three days with Claude Code.
2026 - January
10 postsVibe Code: Which LLM Is the BEST?? Let's Talk for REAL
I ran the same extensive analysis pass through Opus 4.5, GPT 5.2 Codex, Kimi 2.5, Gemini 3 Pro and MiniMax v2.1 on a tiny app, and every single one still left things on the table.
Vibe Code: I Built a Little App 100% with GLM 4.7 (TV Clipboard)
A full walkthrough of vibe coding a tiny Go WebSocket app with GLM 4.7, showing in practice what the 80/20 rule really costs in time, tokens and intervention.
AI Agents: Which One Is Best? OpenCode, Crush, Claude Code, GPT Codex, Copilot, Cursor, Windsurf, Antigravity?
A practical take on the current state of AI coding agents, why proprietary harnesses matter, and why I am sticking with Crush.
AI 3D: Can You Actually Model 3D With Prompts Now?
A practical walkthrough of generating high-quality 3D models from prompts and images using Nano Banana, Hunyuan, Hitem3D, and Blender, all the way to a physical Bambulab print.
AI Agents: Is GLM 4.7 Flash really that good?
Running GLM 4.7 Flash locally on an RTX 5090 via Ollama to see if it can finally match commercial LLMs on a hard coding challenge.
Omarchy 3: Dual GPU Setup with AMD and NVIDIA
How I configured Hyprland on Omarchy to use an AMD iGPU as primary and offload 3D apps to an NVIDIA RTX 5090 via Prime.
AI Agents: Installing LSPs for Crush
My Crush config with LSPs wired in, plus an Arch Linux script that auto-detects installed languages and installs the matching language servers.
AI Agents: Comparing the Top LLMs of 2026 on the Zig Challenge
I put the top commercial and open source LLMs of early 2026 through a nasty Zig + llama.cpp migration challenge to see which ones actually deliver.
AI Agents: Locking Down Your System
I show how to isolate coding agents with Bubblewrap: the jail leaves only the project writable, hides the rest of the system, and mounts only what is needed, reducing the risk of destructive commands.
Omarchy 3 - One of the Best Coding Agents Out There: Crush
I used Crush with Claude Opus 4.5 on a Hugo blog and a Zig chat using llama.cpp, CUDA, and Qwen3. It refactored scripts, updated APIs, fixed a segmentation fault, and delivered a working binary for nearly $8.
2025 - September
12 postsOmarchy 2.0 - LazyVim Basics
A hands-on introduction to LazyVim for people coming from VS Code or Sublime Text, covering the essential Vim navigation, editing, and block commands to get productive fast.
Omarchy 2.0 - Recommended for Beginners?
Why Omarchy is my daily driver, why there is no such thing as a "distro for beginners", and how to use Ventoy and Distrobox to test distros yourself instead of outsourcing your decisions.
Omarchy 2.0 - Install with the Omarchy ISO
Why the official Omarchy ISO is the easiest way to install Omarchy, with LUKS encryption, Limine+Snapper rollback, and Hyprland+UWSM auto-login out of the box.
Omarchy 2.0 - Bitwarden Self-Hosted / VaultWarden
Self-hosting a Bitwarden-compatible password manager with VaultWarden on a home server, routed through a Cloudflare tunnel, with 2FA via Aegis.
Installing Grafana on My Home Server
How I set up Grafana, Prometheus, and cAdvisor with Docker Compose to monitor my Intel NUC home server.
Protecting Your Home Server with Cloudflare Zero Trust
How I added a Cloudflare Zero Trust login layer on top of my home server services, using Google OAuth as the identity provider.
Accessing My Home Server With a Real Domain
An advanced guide to exposing a Docker-based home server through Cloudflare Tunnels with a real domain and proper HTTPS.
Omarchy 2.0 - Understanding SSH and Yubikeys
A practical guide to SSH keys, Yubikey 5 hardware-backed keys, and ssh-agent on Arch Linux and Omarchy.
Omarchy 2.0 - TUIs (Terminal User Interface Apps)
Why TUIs deserve a spot on your desktop in 2025, plus a tour of my favorite terminal apps and how Omarchy makes them feel native.
Omarchy 2.0 - Mise for Organizing Development Environments
How to use Mise on Omarchy (or any Linux distro) to pin language versions per project and stop breaking your apps on every system update.
Omarchy 2.0 - LazyVim - LazyExtras
A practical walkthrough of LazyVim and the :LazyExtras command that ships pre-installed with Omarchy, covering the basic hotkeys and how to enable language support.
Omarchy 2.0 - ZSH Configs
How I swapped Omarchy's default Bash for ZSH and wired up Atuin, Starship, secrets, and handy aliases to match my workflow.
2025 - August
2 postsHow to Contribute to the AkitaOnRails Blog Using Docker
A quick walkthrough on using Docker to spin up the AkitaOnRails blog development environment and start contributing without installing Hugo, Go, or Ruby locally.
Installing Omarchy 2.0 from Scratch - Personal Notes
My personal notes walking through a clean install of Arch Linux with Omarchy 2.0, covering BTRFS, Timeshift snapshots, separate HOME, NFS, and Hyprland customization.
2025 - June
1 postAGI or Skynet Isn't Coming Anytime Soon
Why the current LLM architecture cannot reach AGI, and why the hype from Big Tech CEOs and media-anointed experts is mostly backroom negotiation noise.
2025 - May
7 postsYour Windows May Be Crippled Without You Knowing. Check This!!
How a hidden Windows Power Saver profile was silently throttling a high-end desktop to a fraction of its real speed, and how to fix it.
Computing History and Retro Dev on YouTube
A curated list of YouTube channels on computing history and retro development that every dev should watch.
Final Attempt to Train an LLM with LoRA. Cannon Shot, But Missing the Fly.
Renting an 80GB H100 on RunPod to fine-tune Qwen3-32B with a LoRA on Zig 0.14, and discovering that even cannon-class hardware can still miss the target.
Teaching the Latest Zig to Your LLM - Training LoRAs (Sort Of)
A hands-on walkthrough of training a LoRA adapter on Qwen3-8B to teach it about Zig 0.14, covering datasets, fine-tuning, vLLM serving, and the limits of the experiment.
RANT - LLMs are LOOT BOXES!
Why LLMs for real coding behave like gacha loot boxes, and why every incentive in the industry pushes you to burn more tokens.
When Do LLMs Fail at Programming? A More Realistic Use Case.
A real-world attempt to build a Zig program using llama.cpp with Gemini 2.5 Pro and Aider, showing where LLMs crumble on new languages and libraries.
Rant - Will LLMs Evolve Forever? Demystifying LLMs in Programming
A hype-busting rant on LLM benchmarks, the myth of exponential AI progress, and my hands-on experience showing why Vibe Coding does not work.
2025 - April
17 postsDissecting an Ollama Modelfile - Tuning Qwen3 for Code
A walkthrough of sampling parameters (temperature, top_p, top_k, min_p, repeat_penalty) and how to build a custom Ollama Modelfile to tune Qwen3 for software development tasks.
Testing the Newly Released Open Source LLM - Qwen3 (with Aider and Ollama)
First hands-on impressions of Qwen3 running locally via Ollama and remotely via OpenRouter, tested with Aider on real refactoring tasks.
Destroying the ChatGPT 4o "Personality"
How to disable ChatGPT's pseudo-human chatter and get straight, dry technical answers by tweaking the custom traits settings.
Testing LLMs with Aider on RunPod - which one to use for code?
Hands-on testing of Qwen2.5 Coder, DeepSeek, Codestral and others with Aider, locally on an RTX 4090 and on RunPod, comparing them against Claude and Gemini.
Your Own Free Universal Co-Pilot Running Local: AIDER-OLLAMA-QWEN
How to run Aider with Ollama and Qwen 2.5 Coder locally to build a free, universal Co-Pilot alternative that works with any editor.
LLM Hello World: Building Your Own Local AI Chat
A hands-on Hello World experiment explaining how LLMs actually work under the hood, building a tiny local chat CLI with Qwen 2.5 Coder running on a single GPU.
Accessing Your NAS Using iSCSI Instead of SMB
Why I moved my Docker data-root off NFS and onto an iSCSI virtual drive served from my Synology NAS, and how block-level storage crushes file-based protocols for high-churn workloads.
Changing Clothes Using A.I. (ComfyUI)
A walkthrough of a ComfyUI workflow using IDM-VTON and IPAdapter to swap clothing on photos while preserving the original face and pose.
Using A.I. (ComfyUI) to Generate NPCs in Game Development
A walkthrough of Mickmumpitz's ComfyUI workflow for generating coherent NPC character sheets from scratch, running on my Docker Compose setup.
Understanding the Basics of ComfyUI to Generate AI Images
A hands-on, no-nonsense walkthrough of ComfyUI's core concepts (checkpoints, LoRAs, ControlNet, VAEs, samplers) and a real workflow for turning a photo into anime-style art.
Generating AI Images - even Ghibli style 😂 - with Docker and CUDA
A technical walkthrough of running ComfyUI locally with Docker and CUDA to generate AI images, including anime and Ghibli styles, with full control over models, extensions, and workflows.
Generating Up to 2-Minute Videos from a Photo with A.I.
Using FramePack with Tencent's HunyuanVideo in Docker on an RTX 4090 to animate photos of action figures and drawings into videos up to two minutes long.
Upscaling Old Anime to 4K with AI
Using Real-ESRGAN and Video2X inside a CUDA Docker container to upscale old anime files to 4K on an NVIDIA RTX 4090.
Colorizing Black and White Images with A.I.
Using DDColor inside Docker to colorize old black and white photos, plus a Reinhard global color transfer hack to bias the result toward a reference image.
Configuring My Synology NAS with NFS on Linux
How I switched from SMB to NFS to mount my Synology DS1821+ shares on Linux with proper ownership and permissions, plus the UID/GID gotcha along the way.
NVIDIA and Wayland - Problems for PCI Passthrough in VMs
How GNOME on Wayland breaks live NVIDIA GPU passthrough to QEMU/Libvirt VMs, and the workaround scripts I use to make it work again.
BIOS Configuration of My PC - X670E Aorus Xtreme
My BIOS tuning notes for a Ryzen 9 7950X3D on a Gigabyte X670E Aorus Xtreme, plus Linux-Zen kernel and an automated BIOS update checker.