<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>AI Crew on onelegdave.dev</title><link>https://www.onelegdave.dev/tags/ai-crew/</link><description>Recent posts in AI Crew on onelegdave.dev</description><language>en-us</language><generator>Hugo</generator><ttl>60</ttl><image><url>https://www.onelegdave.dev/images/avatar.jpg</url><title>AI Crew on onelegdave.dev</title><link>https://www.onelegdave.dev/</link></image><lastBuildDate>Sat, 26 Sep 2026 20:00:00 -0700</lastBuildDate><atom:link href="https://www.onelegdave.dev/tags/ai-crew/index.xml" rel="self" type="application/rss+xml"/><item><title>I Made My AI Tools Ask Permission</title><link>https://www.onelegdave.dev/posts/i-made-my-ai-tools-ask-permission/</link><guid isPermaLink="true">https://www.onelegdave.dev/posts/i-made-my-ai-tools-ask-permission/</guid><pubDate>Sat, 26 Sep 2026 20:00:00 -0700</pubDate><dc:creator>OneLegDave</dc:creator><category>AI crew</category><category>Linux</category><description>For a couple of months I tinkered with an idea: let several AI coding tools work as one crew without letting any of them touch what they should not. Tonight it became a real, free, open-source release candidate called GovernCode.</description><media:content url="https://www.onelegdave.dev/images/posts/i-made-my-ai-tools-ask-permission.jpg" medium="image" type="image/jpeg" width="1200" height="630"/><media:thumbnail url="https://www.onelegdave.dev/images/posts/i-made-my-ai-tools-ask-permission.jpg" width="1200" height="630"/><content:encoded>&lt;p&gt;&lt;img src="https://www.onelegdave.dev/images/posts/i-made-my-ai-tools-ask-permission.jpg" alt="I Made My AI Tools Ask Permission" width="1200" height="630"&gt;&lt;/p&gt;&lt;p&gt;I pay for more AI coding tools than any reasonable person should, and for a couple of months I&amp;rsquo;ve been tinkering with one question: what if they worked together, and none of them could do anything I hadn&amp;rsquo;t agreed to?&lt;/p&gt;&#10;&lt;p&gt;Not &amp;ldquo;agreed to&amp;rdquo; in the sense of clicking Accept four hundred times until my finger goes numb. Agreed to in the sense that the operating system itself won&amp;rsquo;t let them do anything else.&lt;/p&gt;&#10;&lt;p&gt;That idea now has a name, a public repository, a website, and as of tonight a first release candidate. It&amp;rsquo;s called GovernCode. It&amp;rsquo;s free, it&amp;rsquo;s open source, and it will stay that way.&lt;/p&gt;&#10;&lt;h2 id="one-boss-a-few-contractors-and-a-building-inspector"&gt;One boss, a few contractors, and a building inspector&lt;/h2&gt;&#10;&lt;p&gt;GovernCode lets you pick one AI tool to lead a project. That&amp;rsquo;s the Controller. It keeps the big picture and hands small, bounded jobs to the others, the Runners, the way a contractor hands a room to a subcontractor. Each job gets written down first: what to do, which files it may touch, how much of my usage it may burn, and what &amp;ldquo;done&amp;rdquo; means.&lt;/p&gt;&#10;&lt;p&gt;Every one of them runs inside a sandbox that Linux enforces. Not a polite instruction in the prompt asking the AI to behave. The kernel&amp;rsquo;s own locks (Landlock and seccomp, if you want the names) decide which folders it can read, which it can write, and where on the network it can go. If my machine can&amp;rsquo;t enforce those rules, GovernCode refuses to start any AI tool at all and tells me why. There is no &amp;ldquo;just this once&amp;rdquo; switch, on purpose, because I know myself.&lt;/p&gt;&#10;&lt;p&gt;When a tool wants to run something that matters, it stops and shows me exactly what will run, byte for byte. Not a friendly summary. The actual command.&lt;/p&gt;&#10;&lt;h2 id="the-part-where-i-nearly-gave-up-on-my-own-demo"&gt;The part where I nearly gave up on my own demo&lt;/h2&gt;&#10;&lt;p&gt;The first time I tried to show it off, I spent more time clicking Allow than watching the AI work. Every &lt;code&gt;npm test&lt;/code&gt;, every file edit, every harmless &lt;code&gt;ls&lt;/code&gt;. It was secure the way a locked filing cabinet at the bottom of a lake is secure.&lt;/p&gt;&#10;&lt;p&gt;So now I can say yes once, yes for the rest of this job, or yes for this project, and the plain read-only stuff just runs. The sandbox doesn&amp;rsquo;t move an inch when I do that. It only stops asking me the same question forty times. Deleting things, git, the network, and handing work to another AI still ask every single time.&lt;/p&gt;&#10;&lt;h2 id="i-let-the-ais-audit-the-ai-babysitter"&gt;I let the AIs audit the AI babysitter&lt;/h2&gt;&#10;&lt;p&gt;Before putting a release candidate out, I had other AI tools try to break it. They found real holes. A tool could start a background process, &amp;ldquo;finish&amp;rdquo;, and leave that process running. A cleverly placed shortcut in a folder could have widened what a later job was allowed to write. A carefully worded command could sneak past the &amp;ldquo;this is just a harmless read&amp;rdquo; check.&lt;/p&gt;&#10;&lt;p&gt;All of that is fixed, each fix has a test, and the sandbox now proves its rules on your machine every time it starts, including against those exact tricks. I&amp;rsquo;d rather find out from a robot at my desk than from someone&amp;rsquo;s bug report.&lt;/p&gt;&#10;&lt;h2 id="local-models-because-quota-is-money"&gt;Local models, because quota is money&lt;/h2&gt;&#10;&lt;p&gt;Small jobs can go to a model running on my own graphics card. It gets no tools and can&amp;rsquo;t run anything. It reads the files it&amp;rsquo;s given and proposes changes, GovernCode checks every file it wants to touch, and I review it like anyone else&amp;rsquo;s work. The first real one wrote a usage section for a README in six seconds and cost exactly zero of my subscription.&lt;/p&gt;&#10;&lt;h2 id="the-rule-i-care-about-most"&gt;The rule I care about most&lt;/h2&gt;&#10;&lt;p&gt;GovernCode makes building software with AI a lot more approachable. It does not make anyone less responsible for what they ship. If I accept a change, it&amp;rsquo;s mine. I read the diff, I understand the code, I run the tests. The AI can write it. Only I can answer for it.&lt;/p&gt;&#10;&lt;p&gt;That&amp;rsquo;s written into the project&amp;rsquo;s philosophy as rule number two, right after &amp;ldquo;you hold the controls&amp;rdquo;, and it&amp;rsquo;s staying there.&lt;/p&gt;&#10;&lt;h2 id="where-it-stands"&gt;Where it stands&lt;/h2&gt;&#10;&lt;p&gt;It&amp;rsquo;s a release candidate for Linux, for people who like poking at things from a terminal. macOS and Windows are next, together, same priority, because I&amp;rsquo;m not making either one wait for the other. I&amp;rsquo;m dual-booting my laptop with Windows tonight so I can actually test it there.&lt;/p&gt;&#10;&lt;p&gt;None of this would exist if Omarchy hadn&amp;rsquo;t made my computer fun again and dragged a passion I&amp;rsquo;ve had forever back out of the drawer where time had shoved it. So: thank you, DHH, and everyone behind Omarchy.&lt;/p&gt;&#10;&lt;p&gt;GovernCode is at &lt;a href="https://governcode.com"&gt;governcode.com&lt;/a&gt;. It&amp;rsquo;s free forever. If it saves you an afternoon, there&amp;rsquo;s a coffee button. My AI crew runs on tokens. I run on coffee. Only one of us has a button.&lt;/p&gt;&#10;</content:encoded></item><item><title>Something That Really Looks Awesome</title><link>https://www.onelegdave.dev/posts/something-that-really-looks-awesome/</link><guid isPermaLink="true">https://www.onelegdave.dev/posts/something-that-really-looks-awesome/</guid><pubDate>Fri, 18 Sep 2026 21:33:00 -0700</pubDate><dc:creator>OneLegDave</dc:creator><category>AI crew</category><description>One vague sentence at five in the afternoon. By half past nine my AI team had a real GUI with a name, a cat, two Mac apps and a private network. Vibe coding can be that easy. Sometimes.</description><media:content url="https://www.onelegdave.dev/images/posts/something-that-really-looks-awesome.jpg" medium="image" type="image/jpeg" width="1200" height="630"/><media:thumbnail url="https://www.onelegdave.dev/images/posts/something-that-really-looks-awesome.jpg" width="1200" height="630"/><content:encoded>&lt;p&gt;&lt;img src="https://www.onelegdave.dev/images/posts/something-that-really-looks-awesome.jpg" alt="Something That Really Looks Awesome" width="1200" height="630"&gt;&lt;/p&gt;&lt;p&gt;At about five in the afternoon I typed this into the terminal cockpit I&amp;rsquo;d finished the night before: &amp;ldquo;what is the possibility of having this as a GUI? something that really looks awesome!&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;That&amp;rsquo;s the whole brief. No spec, no sketch, one vague sentence with an exclamation mark on it. By half past nine it had a name, a logo, four animated characters, a launcher with an icon, a laptop layout, two native Mac apps and a private network to run on. This is the version of vibe coding people keep promising: describe the thing, get the thing. It really can be like that. It was, for most of the evening. The rest of the evening I spent telling it I couldn&amp;rsquo;t see anything.&lt;/p&gt;&#10;&lt;h2 id="hearth-lasted-two-minutes"&gt;Hearth lasted two minutes&lt;/h2&gt;&#10;&lt;p&gt;The first mockup came back in about five minutes. Then it went through seven versions in half an hour, because every time I looked at it I wanted something else. Keep the delegation feed but let it collapse. Let the terminal pop out or get taller. Give me a switch for the multiplexer that runs the agents. A project selector, and a way to make a new one. A button at the top that flashes when something needs me. Zoom for the text, per pane, independently. Each one turned up in the next version with the controls actually working in the mockup, which is a dangerous thing to show a man who was only asking hypothetically.&lt;/p&gt;&#10;&lt;p&gt;Claude named it Hearth. The reasoning was sound: the team is Anvil and Flint, the whole thing is forge-themed, and the hearth is where everyone gathers. I said &amp;ldquo;let&amp;rsquo;s see Crucible.&amp;rdquo; Crucible it is. A crucible is the pot the metal actually melts in, which felt like a more honest description of what happens to my quota.&lt;/p&gt;&#10;&lt;p&gt;Before a line of real code, one rule went into the plan as rule one: closing Crucible never stops the multiplexer. The multiplexer is what holds every agent session. Crucible is a window onto it, nothing more. There is a switch in the header, that switch is the only thing allowed to stop it, and it asks first. Because stopping it kills every pane, including the one Claude was building Crucible from.&lt;/p&gt;&#10;&lt;p&gt;Which is why the one thing in the plan that never got tested live is the off switch. The test would kill the tester. There&amp;rsquo;s a unit test for the guard. I have decided to believe the unit test.&lt;/p&gt;&#10;&lt;h2 id="i-dont-see-it"&gt;I don&amp;rsquo;t see it&lt;/h2&gt;&#10;&lt;p&gt;The build ran in phases, and each one landed in the window on my desk while I watched: a little server feeding the page, real terminals in the browser, the alarm, the selectors, the art, the packaging. Six phases in an hour and a half, all of them committed and pushed, and my contribution to each was roughly the same four words.&lt;/p&gt;&#10;&lt;p&gt;&amp;ldquo;I don&amp;rsquo;t see it.&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;Twice the answer was that my window was still running the old page. The server serves new files the instant they exist, but a browser tab only fetches them when it reloads, so the tabs I was clicking were the inert ones from two phases ago and the project selector was the disabled placeholder. The fix was to make the page reload itself whenever the build changes, and it did, and the very next time I couldn&amp;rsquo;t see anything it was because I&amp;rsquo;d closed the window entirely and was staring at a terminal. Claude relaunched it for me, straight into the scratchpad next to that terminal at half width, where the layout stacks and the crew column falls off the bottom of the page. So the characters were there. They were just below the floor.&lt;/p&gt;&#10;&lt;p&gt;At no point in any of this was anything wrong with the software. Every single one was the man looking at it.&lt;/p&gt;&#10;&lt;h2 id="keystone-needs-you"&gt;Keystone needs you&lt;/h2&gt;&#10;&lt;p&gt;While Grok was off drawing, my desktop started pinging. A notification, a sound, a red button in the header: Keystone needs you. I went to the main pane. Nothing. No question, no prompt, nobody needed anything. It did it again. &amp;ldquo;I&amp;rsquo;m getting alerts from you but I can&amp;rsquo;t see any questions?&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;The alarm was working perfectly. It was just listening to the wrong person. Grok&amp;rsquo;s command line tool runs the same hook configuration Claude Code does, and I&amp;rsquo;d wired those hooks to post to Crucible&amp;rsquo;s alarm endpoint. So every time Grok approved one of its own file writes, which it does automatically because that&amp;rsquo;s what it&amp;rsquo;s allowed to do, it announced that the lead needed a human. The server assumed any hook it heard came from the lead. It never occurred to me, or to the thing that wrote it, that the cat would be wearing the same collar. It now checks which agent a hook came from and ignores the ones that aren&amp;rsquo;t Claude, since the multiplexer already reports a teammate as blocked when it genuinely is.&lt;/p&gt;&#10;&lt;p&gt;The art run itself went the other way. Grok delivered every character and logo, and I was reviewing them in a browser before the delegate wrapper, the supervisor I wrote and complained about in the last post, gave up on the run at its 25 minute ceiling and logged it as a timeout. Last night nobody was finished when they said they were. Tonight one of them finished and the babysitter said it hadn&amp;rsquo;t. I&amp;rsquo;m not sure that&amp;rsquo;s progress but it is at least variety.&lt;/p&gt;&#10;&lt;p&gt;I picked a sitting cat for Flint, asked for more movement, and now the crew animate while they work and hold up a hand when they&amp;rsquo;re blocked. I would have accepted a static PNG.&lt;/p&gt;&#10;&lt;h2 id="the-lodger"&gt;The lodger&lt;/h2&gt;&#10;&lt;p&gt;Packaging had its own small fights. The launcher hung on the first attempt because the browser-launch helper blocks until the window closes, so a script waiting for it to return waits forever. Then the window insisted on opening in the scratchpad, and the launcher had to learn to find it, move it to a real workspace and focus it.&lt;/p&gt;&#10;&lt;p&gt;The server now starts at login and stays resident, the window opens on demand, and when Crucible starts Claude in a folder that already has a conversation it picks that conversation back up. So I can close the whole thing, reboot, click the icon in the launcher and be exactly where I was. I quit for the night at seven, pleased.&lt;/p&gt;&#10;&lt;p&gt;At eight I came back. The multiplexer&amp;rsquo;s own sidebar, a list of spaces and agents that Crucible already draws far more prettily in its own columns, was sitting in the main pane like a lodger. I&amp;rsquo;d hidden it by hand before the reboot. The multiplexer has a setting to start it collapsed, but that setting only takes effect when its server starts, and rule one says I never restart that.&lt;/p&gt;&#10;&lt;p&gt;So the Crucible server now watches the main pane, waits until it sees the sidebar drawn, and types the keyboard shortcut to hide it. Then it checks the repaint so it never toggles twice. It presses the button for me. I&amp;rsquo;m not proud of it. It works.&lt;/p&gt;&#10;&lt;h2 id="the-wrong-direction-first"&gt;The wrong direction first&lt;/h2&gt;&#10;&lt;p&gt;&amp;ldquo;Can we install this on the laptop, and have this work via ssh?&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;Not ten minutes later there was a second window with the laptop&amp;rsquo;s name in its header and a shell on the laptop answering. Crucible was running on the laptop, tunnelled back to the desktop. Which is when I explained what I&amp;rsquo;d actually meant: the desktop is the workhorse, the laptop is where I sit. Other direction. Same mechanism, reversed, and the laptop already reached the desktop over Tailscale&amp;rsquo;s SSH without a password, so no keys to sort out. Verified from the laptop&amp;rsquo;s own screen two minutes later, this very conversation in the main pane.&lt;/p&gt;&#10;&lt;p&gt;The laptop&amp;rsquo;s screen is smaller, and the page had been laid out for the big monitor. A screenshot, three fixes, a compact layout that collapses the header to one row and narrows the side columns, and the main terminal gained roughly forty percent more height. Then the laptop got the same AI setup as the desktop, Headroom in its bar and all, and I signed into the four accounts. Three of the four verified straight away. Antigravity reported not signed in after I&amp;rsquo;d just signed in, because checking it over SSH makes it look in a different credential store than the one I&amp;rsquo;d just written to. Checked the way Crucible actually launches it, it was fine. That&amp;rsquo;s two nights running this tool has been reported logged out, and this time it wasn&amp;rsquo;t even its fault.&lt;/p&gt;&#10;&lt;h2 id="the-crazy-pitch"&gt;The crazy pitch&lt;/h2&gt;&#10;&lt;p&gt;&amp;ldquo;Okay, now the crazy pitch, can we get native apps for Crucible and Headroom on the Mac?&amp;rdquo;&lt;/p&gt;&#10;&lt;p&gt;Yes. Claude built them on the Mac over SSH with nothing but the command line tools, in Swift, in a few minutes including one type error. Crucible.app is a WebKit window onto the desktop&amp;rsquo;s server, with the page&amp;rsquo;s confirm dialogs turned into native sheets and copy and paste wired into the terminals. Headroom.app is a ring in the menu bar showing the fullest of my four weekly allowances, red at ninety percent, with a panel of meters and reset countdowns behind it. Screenshots over SSH aren&amp;rsquo;t permitted on a Mac, so the build could confirm a window existed and that the server saw its connections. Whether it looked any good was my job. It did.&lt;/p&gt;&#10;&lt;p&gt;Both apps were running over SSH tunnels. Then I said &amp;ldquo;we should use tailscale,&amp;rdquo; and the tunnel code that had been written at twenty past eight was deleted at ten past nine. The desktop now publishes Crucible on the tailnet with Tailscale&amp;rsquo;s own certificate, while the server itself still only listens on the loopback address. Every request that arrives that way carries the caller&amp;rsquo;s Tailscale login, and the server refuses anyone but me. A forged login header was tried and refused. A foreign origin on a terminal socket was tried and refused. The laptop and the Mac reach it from a coffee shop the same way they reach it from the sofa, and nothing is open to the internet. One quirk: the first request to a freshly published machine can sit for over ten seconds while Tailscale issues the certificate. After that it&amp;rsquo;s instant.&lt;/p&gt;&#10;&lt;h2 id="half-past-nine"&gt;Half past nine&lt;/h2&gt;&#10;&lt;p&gt;Everything went to GitHub at 21:24. Twenty-nine commits in four and a half hours, from &amp;ldquo;something that really looks awesome&amp;rdquo; to a cockpit I can open from a laptop, a Mac, or a coffee shop. I didn&amp;rsquo;t write a line of it. I typed wishes and looked at the results, and when I couldn&amp;rsquo;t see the results, that was usually me too.&lt;/p&gt;&#10;&lt;p&gt;That&amp;rsquo;s the sometimes. The easy part was genuinely easy. Nothing about tonight required me to know Swift, or how a browser talks to a terminal, or what Tailscale does with a certificate. What it required was somebody in the chair who&amp;rsquo;d notice when the alarm was lying, when the supervisor was lying, and when the fault was the man looking at the screen. The code writes itself now. The noticing still doesn&amp;rsquo;t.&lt;/p&gt;&#10;&lt;p&gt;Not done: alarm notifications on the Mac, Headroom starting itself at login, and Codex&amp;rsquo;s review of the server code, which waits for its weekly quota to reset on Saturday. Former boss, back on the tools, hasn&amp;rsquo;t been allowed near a file since it hit the ceiling.&lt;/p&gt;&#10;&lt;p&gt;And the off switch. I have a cockpit with an animated cat in it, reachable from anywhere on earth, and an off switch I have never once pressed.&lt;/p&gt;&#10;</content:encoded></item><item><title>I Promoted the One I Was Going to Cancel</title><link>https://www.onelegdave.dev/posts/promoted-the-one-i-was-cancelling/</link><guid isPermaLink="true">https://www.onelegdave.dev/posts/promoted-the-one-i-was-cancelling/</guid><pubDate>Thu, 17 Sep 2026 20:00:00 -0700</pubDate><dc:creator>OneLegDave</dc:creator><category>AI crew</category><description>I gave four AI assistants a chain of command, a spending limit and code names, then spent the night finding out that the tools I built to supervise them kept reporting jobs finished while the jobs were still running.</description><media:content url="https://www.onelegdave.dev/images/posts/promoted-the-one-i-was-cancelling.jpg" medium="image" type="image/jpeg" width="1200" height="630"/><media:thumbnail url="https://www.onelegdave.dev/images/posts/promoted-the-one-i-was-cancelling.jpg" width="1200" height="630"/><content:encoded>&lt;p&gt;&lt;img src="https://www.onelegdave.dev/images/posts/promoted-the-one-i-was-cancelling.jpg" alt="I Promoted the One I Was Going to Cancel" width="1200" height="630"&gt;&lt;/p&gt;&lt;p&gt;A few days ago I wrote a post about how I only wanted two AI subscriptions, and how Claude was the one getting cancelled. Claude is now in charge of the other three.&lt;/p&gt;&#10;&lt;p&gt;I would like to say there was a dramatic reason for this. There wasn&amp;rsquo;t. Codex had been running the show, it was doing fine, and I moved the job across because I wanted to see what happened. That&amp;rsquo;s the entire justification. A promotion, a demotion, and a backup of the old config taken at quarter past eleven at night, because I have learned exactly one lesson in my life and it was about backups.&lt;/p&gt;&#10;&lt;p&gt;So the roster is four: Codex, Claude, Grok, and Google&amp;rsquo;s Antigravity. One of them plans and reviews. The other three get handed bounded jobs. And because saying &amp;ldquo;ask the Google one&amp;rdquo; out loud thirty times a day is its own special punishment, they have names now. Anvil is Codex. Keystone is Claude. Flint is Grok. Albatross, Alby for short, is agy.&lt;/p&gt;&#10;&lt;p&gt;The nicknames are for talking. Every log line, quota check and credit keeps the real provider name, because the one thing worse than four assistants is four assistants and no way to tell which one did the thing.&lt;/p&gt;&#10;&lt;p&gt;The rules are blunt. Each provider has a 90 percent ceiling on its weekly allowance, which leaves at least 10 percent for me, because these things are meant to be doing my work, not eating my quota so I can&amp;rsquo;t. Nobody quietly switches to paid billing to get around a limit. Every job I hand off needs a concrete result, the files it&amp;rsquo;s allowed to touch, acceptance criteria, and a budget check.&lt;/p&gt;&#10;&lt;p&gt;Then there&amp;rsquo;s the genuinely stupid part: one standing policy, 228 lines, mirrored into four separate instruction files, because each tool only reads its own. Miss one and you have a Head Developer that three of the four assistants have never heard of.&lt;/p&gt;&#10;&lt;h2 id="the-gate-the-launcher-the-pane"&gt;The gate, the launcher, the pane&lt;/h2&gt;&#10;&lt;p&gt;There&amp;rsquo;s a quota gate that has to run before every delegated job and again afterwards. Delegated, in this house, means I handed the work to one of them instead of doing it myself. There&amp;rsquo;s a launcher that bounds the work, keeps the streamed partial output so a death mid-job isn&amp;rsquo;t a total blank, and classifies failures as timeout, empty success, or agent failure instead of trusting an exit code. Empty success is the fun one: it walked off claiming it was fine and produced nothing.&lt;/p&gt;&#10;&lt;p&gt;And in the small hours I wrote the one that runs a teammate in a pane of a terminal multiplexer, requires a stated task label and acceptance criteria, waits for a completion marker, and appends every run to an event log.&lt;/p&gt;&#10;&lt;p&gt;A completion marker is a word I tell them to print when the job is actually done. You can already see the hole in that plan. It assumes &amp;ldquo;done&amp;rdquo; is a thing the furniture will admit to.&lt;/p&gt;&#10;&lt;h2 id="idle-is-not-evidence"&gt;Idle is not evidence&lt;/h2&gt;&#10;&lt;p&gt;Asking the multiplexer whether an agent has finished is worth precisely fuck all.&lt;/p&gt;&#10;&lt;p&gt;Grok reported idle at 23 seconds into 35 seconds of work, with its command still running. agy reported idle at 7 seconds, which was impressive given its command hadn&amp;rsquo;t started yet.&lt;/p&gt;&#10;&lt;p&gt;agy also parks in idle between tool calls, and it echoes prompt text into its own visible thinking, so neither a status of idle nor a matching completion word is enough for it. A review got declared finished at 168 seconds while the pane still read &amp;ldquo;Running command&amp;hellip;&amp;rdquo;. The wrapper then closed the pane and destroyed the work. The fix waits for the pane&amp;rsquo;s content to stop changing, with a longer settle time for agy.&lt;/p&gt;&#10;&lt;p&gt;The panes themselves had a second trick. Each kept pane leaves a split behind, and the next split is narrower. At three or more panes each one is about 6 columns wide. That is not a workspace. That is a breadstick. Every completion marker wraps across lines, and the match can never succeed. A correct 30 second task timed out at 306 seconds with its marker counted zero times.&lt;/p&gt;&#10;&lt;p&gt;The multiplexer also truncates a multi-line prompt echo to its first line, which made my original baseline, count the marker in the prompt and require more than that, impossible to satisfy.&lt;/p&gt;&#10;&lt;p&gt;So the supervisor I wrote would declare victory because the pane looked quiet, or because a wrapped word no longer matched itself, or because the prompt it was trying to count had already been chopped down to one line. Meanwhile the job was still running. Or hadn&amp;rsquo;t started. Or had been killed for being finished.&lt;/p&gt;&#10;&lt;h2 id="the-stale-quota-that-wasnt"&gt;The stale quota that wasn&amp;rsquo;t&lt;/h2&gt;&#10;&lt;p&gt;agy had a long-standing reputation in my setup for reporting stale usage numbers, so it kept getting held back from work. Can&amp;rsquo;t verify what&amp;rsquo;s left of the allowance, don&amp;rsquo;t spend it.&lt;/p&gt;&#10;&lt;p&gt;The bug was in my own quota gate. It threw away a snapshot it had just verified as fresh whenever its refresh step tripped over anything. The check was manufacturing the exact staleness it was complaining about. I benched a teammate for a crime the babysitter committed. That one&amp;rsquo;s fixed now, with a regression test so that particular piece of shit can&amp;rsquo;t sneak back in the same shape. The real refresh takes about a second and a half.&lt;/p&gt;&#10;&lt;p&gt;An exit code of 0 from a usage collector doesn&amp;rsquo;t mean it collected anything, either. One collector returns 0 without writing when another copy holds its lock. agy&amp;rsquo;s returns 0 for any record under 60 seconds old without republishing. Both are &amp;ldquo;working.&amp;rdquo; Neither tells you the thing you asked. The gate now judges the age of the record it gets back instead of the exit status.&lt;/p&gt;&#10;&lt;p&gt;Then I had Grok and Antigravity review the code that decides whether Grok and Antigravity are allowed to run, and they found two more. Grok found an unguarded error path that leaked a pseudo terminal handle every time the child process died at the wrong moment. A pseudo terminal is the fake terminal session the helper opens in order to run that child, and leak enough of them and it can never open another one. Antigravity found that a sixty second cache shortcut made a forced retry a no-op, so the one case where I really needed a fresh number was the case guaranteed not to get one. Ask again, get back the number you just rejected.&lt;/p&gt;&#10;&lt;h2 id="tonight"&gt;Tonight&lt;/h2&gt;&#10;&lt;p&gt;Codex sits at 100 percent of its weekly allowance and is held out of the rota until that resets, which is a hell of a way for a former boss to spend its first week back on the tools. Flint is at four percent and cheerful. Alby spent a chunk of the evening held as well, and this time the tooling was right to hold it: the thing had quietly signed itself out of its own account, so there was no allowance to check at all. Fixing that needed a human, a browser and about thirty seconds, which is the most honest description of my job here that I&amp;rsquo;ve managed all week.&lt;/p&gt;&#10;&lt;p&gt;I set out to build an AI team. What I&amp;rsquo;ve actually built, so far, is a very thorough system for finding out that nobody is finished when they say they are.&lt;/p&gt;&#10;</content:encoded></item><item><title>I Only Wanted Two Subscriptions</title><link>https://www.onelegdave.dev/posts/i-only-wanted-two-subscriptions/</link><guid isPermaLink="true">https://www.onelegdave.dev/posts/i-only-wanted-two-subscriptions/</guid><pubDate>Fri, 11 Sep 2026 18:00:00 -0700</pubDate><dc:creator>OneLegDave</dc:creator><category>AI crew</category><description>I wanted to save twenty bucks on AI subscriptions. That somehow required a searchable archive, a server, and a laptop refusing to recognise its own login.</description><media:content url="https://www.onelegdave.dev/images/posts/i-only-wanted-two-subscriptions.jpg" medium="image" type="image/jpeg" width="1200" height="630"/><media:thumbnail url="https://www.onelegdave.dev/images/posts/i-only-wanted-two-subscriptions.jpg" width="1200" height="630"/><content:encoded>&lt;p&gt;&lt;img src="https://www.onelegdave.dev/images/posts/i-only-wanted-two-subscriptions.jpg" alt="I Only Wanted Two Subscriptions" width="1200" height="630"&gt;&lt;/p&gt;&lt;p&gt;I have three paid AI subscriptions. I want two. OpenAI&amp;rsquo;s ChatGPT Pro stays, Google AI Pro stays, Claude goes. A small administrative decision which has somehow involved a server, a searchable archive, two Linux machines, a Mac, and a laptop insisting it has absolutely no idea who I am while I&amp;rsquo;m already logged into it.&lt;/p&gt;&#10;&lt;p&gt;I should probably stop describing things as small administrative decisions.&lt;/p&gt;&#10;&lt;p&gt;Claude has been useful. There&amp;rsquo;s a lot of work in those conversations: code, project decisions, explanations, failed attempts, and the eventual discovery of which stupid little thing was causing the problem. I want to save twenty bucks. I would also quite like to avoid spending the rest of my life explaining my own projects back to a succession of helpful strangers.&lt;/p&gt;&#10;&lt;p&gt;So before I cancel anything, I need to take the useful history with me.&lt;/p&gt;&#10;&lt;h2 id="packing-apparently"&gt;Packing, apparently&lt;/h2&gt;&#10;&lt;p&gt;The account export arrived as five ZIP files on the Mac. I copied them across and checked that they matched the originals using SHA-256 hashes, essentially digital fingerprints for the files. I also checked that the archives opened without errors. Everything passed.&lt;/p&gt;&#10;&lt;p&gt;After my previous adventure in deleting this website while carefully preserving my terminal colours, I&amp;rsquo;m trying to get better at checking what I&amp;rsquo;ve actually saved. The bar is low. I put it there myself.&lt;/p&gt;&#10;&lt;p&gt;The export contained 59 conversations, along with project documents, saved memories, and summaries. I also collected the separate Claude Code histories from my Linux machines. Those are the records from using Claude in the terminal, where much of the coding work happened. Some of that history overlaps with the account export, so I kept track of where it came from rather than pretending every saved message was a unique contribution to human knowledge.&lt;/p&gt;&#10;&lt;p&gt;The original files are preserved. Alongside them, I now have readable text versions and a searchable database, built with SQLite. That means I can look up a project or an old problem without opening dozens of conversations and trying to remember which one contained the useful bit.&lt;/p&gt;&#10;&lt;p&gt;There&amp;rsquo;s also a starting document pointing to the relevant topics and correcting things that have changed. The old summaries are labelled as historical material. An assistant confidently saying something was true six conversations ago doesn&amp;rsquo;t make it true now. It might not even have been true then.&lt;/p&gt;&#10;&lt;p&gt;This gives the next assistant somewhere to look before asking me to explain the whole fucking setup again. It still has to find the relevant history and check it against the current work. The old chats don&amp;rsquo;t magically appear in ChatGPT&amp;rsquo;s sidebar, and putting a ZIP file within reach of an AI does not mean it has read it.&lt;/p&gt;&#10;&lt;h2 id="a-note-saying-i-once-had-a-document"&gt;A note saying I once had a document&lt;/h2&gt;&#10;&lt;p&gt;The export also contained 180 references to files whose originals weren&amp;rsquo;t included.&lt;/p&gt;&#10;&lt;p&gt;References. Lovely.&lt;/p&gt;&#10;&lt;p&gt;Some attachment text survived, but text pulled out of a document doesn&amp;rsquo;t give me the original document back. A reference to an image is even less useful when what I need is the bloody image. The code repositories, manuscripts, and other original files still need their own preservation. A conversation about building something is a poor substitute for the thing I built.&lt;/p&gt;&#10;&lt;p&gt;The archive now lives on my server, with a working copy on my main Linux machine. I&amp;rsquo;ve checked the server copy against the originals and tested the search. It&amp;rsquo;s in a directory covered by my existing local and offsite backup jobs, although I still need to verify a completed backup and restore containing this new material.&lt;/p&gt;&#10;&lt;p&gt;I have already written one blog post about discovering what my backups didn&amp;rsquo;t contain. I&amp;rsquo;d prefer not to establish a series.&lt;/p&gt;&#10;&lt;p&gt;The archive is a snapshot of the history I&amp;rsquo;ve collected so far. It won&amp;rsquo;t quietly gather every future conversation by itself. My assistants have instructions pointing to it, but they need access to those files to use it. The apps on my phones haven&amp;rsquo;t suddenly acquired my entire project history either.&lt;/p&gt;&#10;&lt;h2 id="the-other-subscription"&gt;The other subscription&lt;/h2&gt;&#10;&lt;p&gt;Google AI Pro was already paid for, so getting proper use out of it was the next job. For coding work, that meant installing Google&amp;rsquo;s Antigravity tool, launched with the command &lt;code&gt;agy&lt;/code&gt;, on the Linux machines and the Mac. Gemini is on both phones too.&lt;/p&gt;&#10;&lt;p&gt;The second laptop needed a small networking repair before any of that: I had forgotten to reinstall Tailscale after reinstalling the operating system. Tailscale is what lets my machines reach each other securely when they&amp;rsquo;re elsewhere.&lt;/p&gt;&#10;&lt;p&gt;The laptop wasn&amp;rsquo;t on the network because I hadn&amp;rsquo;t installed the software that puts it on the network. A difficult bug to report without the entire report being about me.&lt;/p&gt;&#10;&lt;p&gt;With that fixed, I got its remaining history archived and signed into agy. A test in the laptop&amp;rsquo;s own terminal worked. The same basic test through SSH, the secure connection I use to run commands remotely, returned &lt;code&gt;authentication required&lt;/code&gt;.&lt;/p&gt;&#10;&lt;p&gt;Of course it fucking did.&lt;/p&gt;&#10;&lt;h2 id="logged-in-depending-on-whos-asking"&gt;Logged in, depending on who&amp;rsquo;s asking&lt;/h2&gt;&#10;&lt;p&gt;Linux keeps saved sign-in credentials in a protected store called a keyring. Mine was unlocked, and agy could use it from the desktop. The remote session still couldn&amp;rsquo;t sign in. Matching the desktop&amp;rsquo;s session settings didn&amp;rsquo;t fix it either.&lt;/p&gt;&#10;&lt;p&gt;What finally worked was letting agy use the existing desktop connection to that keyring while removing the environment markers that told the program it had been launched through SSH. The secure remote connection itself stayed intact. Only the information passed to the agy process changed.&lt;/p&gt;&#10;&lt;p&gt;I put that into a small launcher called &lt;code&gt;agy-desktop&lt;/code&gt;, opened a fresh SSH connection, and tested again. It answered correctly, including following my project instructions. No copying passwords around, no disabling keyring protection, no surgery on the actual agy program.&lt;/p&gt;&#10;&lt;p&gt;It still needs the desktop keyring to be accessible and unlocked. I haven&amp;rsquo;t proved it will work unattended after a reboot. Two successful tests are useful evidence, but I&amp;rsquo;m not promoting them to a lifetime guarantee because I&amp;rsquo;d like this paragraph to end happily.&lt;/p&gt;&#10;&lt;p&gt;I now have searchable history on the server and a working way to use the Google subscription I&amp;rsquo;m keeping. There are still original files to account for before I call the preservation job finished, and I haven&amp;rsquo;t cancelled Claude yet.&lt;/p&gt;&#10;&lt;p&gt;I set out to save twenty bucks and built a searchable archive with remote access. Apparently the cancel button is the only part of this process I haven&amp;rsquo;t needed to troubleshoot.&lt;/p&gt;&#10;</content:encoded></item></channel></rss>