1007 | Agents Take the Desk: Building, Billing, and Guardrails

Show notes

A rapid tour of what AI agents can now do for you: work inside your desktop and on real devices, and how a new ecosystem keeps them accountable, from cost ledgers and security to code review. Then a quicker look at agents in sales and support, and a grab bag of small but handy apps.

Timeline

  • 00:00:04 Opening
  • 00:00:36 Agents That Take Over Your Computer and Devices
  • 00:05:16 Working With Agents, Not Instead of Them
  • 00:05:40 Agents Writing Code Need Receipts: Git Identity and AI Code Review
  • 00:07:31 Watching the Agents: Cost Ledgers, Telemetry, and Safety
  • 00:09:57 Desk Logistics for Agent Workers
  • 00:10:59 Agents in Sales, Support, and Search
  • 00:12:48 Your Voice, Imported
  • 00:13:13 Creative and Personal AI, Open and On-Device
  • 00:14:25 Small Mac Tools With Big Opinions
  • 00:15:25 Learning to Build, and Watching AI Live
  • 00:16:34 Closing

Related links

This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.

Transcript

Mia: Welcome back to the show, everybody. I'm Mia.

Milo: And I'm Milo. Today's briefing has one thread running through almost everything: AI agents are stepping out of the chat window and into your actual computer, your phone, your code base, even your customer inbox. And a whole second wave of tools is appearing to watch them while they do it.

Mia: So we'll spend most of our time there — agents that act, then the accountability layer — and work through the rest of today's launches along the way. Let's start with the devices.

Milo: The clearest framing for this whole block might be ruOS. It's a private cloud desktop with coding agents — Claude Code and Codex — preinstalled, and it persists across devices. So you start something on one machine, and the desktop is still there on another. There's an MCP layer, and a free Lite tier that doesn't even require registration.

Mia: And notice the pattern in the pricing model there: it's not a separate AI subscription. The agents it ships with are ones you presumably already pay for. That idea — your existing subscription becomes the engine — shows up again and again today.

Milo: Rill is a native Mac browser built on the same insight. Claude Code and Codex work in it beside you. The killer gesture is pressing ⌘E, which turns the page you're currently looking at into a task for an agent. So instead of copying a link into a chat window and describing what you want, the agent is just... there, in the browser, working alongside you.

Mia: It's free, and it's a real browser, not an agent replacement for browsing. That distinction matters, and we'll come back to it.

Milo: Then there's Incredible, which is the more aggressive version of this. It operates your computer by voice. It watches the screen and acts inside your apps — Mac and Windows. So the agent isn't beside you anymore; it's got its hands on the keyboard.

Mia: Those are very different comfort levels, right? ruOS is a workspace you log into. Rill is collaboration. Incredible is delegation. And iphone-use pushes it further still: it's open source under MIT, and it lets AI agents drive a real iPhone through WebDriverAgent. It exposes an MCP server with twenty-one tools, so an agent can genuinely operate a physical phone rather than a simulator.

Milo: A real phone is a big deal, because phones are where so much of daily digital life actually happens. If agents can act there, the surface area changes completely.

Mia: Now Appto is an interesting one because it's not an agent tool for you — it's an agent pipeline that builds products. It's an iOS app factory that runs on your Mac, using your Claude Code or Codex subscription to find niches, design, code, and submit to Apple. The maker says they've shipped ten apps with it. We should say that plainly: that's a maker's claim, not an independently verified result.

Mia: But the concrete capability — niche discovery through design through code through App Store submission — is specific enough to take seriously as a description of what it does.

Milo: And OpenBot rounds out the device story. It's an open-source, local alternative to Grok Bot, running agents on your PC with your existing Claude, ChatGPT, or Grok plans. There's also a multiplayer mode. Again, local, again using subscriptions you already have.

Mia: So why does all this matter? Because the subscription you already pay is becoming the runtime for tools that act, not just answer. That's the shift.

Milo: What's still unknown — and this is the honest part — is reliability and cost. When an agent is driving your browser, your desktop, or your phone, how often does it do the wrong thing? What does a week of agent-run device time actually cost you in tokens or plan usage? Nobody in this set has published hard numbers on that. That's the open question hanging over the whole category.

Mia: Right, and we shouldn't pretend these tools have answered it. Let's stay in this space a moment longer, though, because Rill deserves a closer look — it makes a particular argument.

Milo: Rill's argument is that agents should work with you, not instead of you. The browser stays yours; the agent is a collaborator in it. iphone-use and OpenBot extend the same pattern — the phone and the desktop become shared spaces where a human and an agent are both present.

Mia: And that's genuinely different from the Incredible model or the Appto model, where you hand something over and check later. Neither approach is wrong, but they imply different trust relationships.

Milo: But here's the bridge: once agents act everywhere — in browsers, on phones, on desktops — teams start needing answers to questions like, who ran what, and can I trust the output? That's where today's next group comes in.

Mia: Code. Specifically, agents writing code, and the tooling growing up around proving and reviewing that code.

Milo: Brnch is the most conceptually interesting one here. It's Git hosting where each agent has its own identity. So your agent isn't committing as you; it's committing as itself. And the feature that stands out: you can test the exact result that would be merged, with a signed receipt.

Mia: A signed receipt for a merge. Think about what that changes. The old question in code review was, which model wrote this? The new question becomes, who reviewed it, and who signed off?

Milo: That's a real trust shift. And the review tooling exists at different trust levels, which is what makes this block coherent rather than just three products.

Mia: Review is an open-source, MIT-licensed desktop app for building your own teams of AI reviewers, using OpenAI, Anthropic, or Ollama models. So you compose the review panel yourself.

Milo: CodeCrab goes further on the privacy axis: it's a hundred percent local desktop PR reviewer that reuses your Claude Code subscription and runs parallel reviews without uploading your code anywhere. So your code never leaves your machine.

Mia: And OrgComputers adds the collaboration layer — a shared workspace for Claude Code, Cursor, and Codex agents via MCP, where agents save and share context with sources. So it's not just one agent being reviewed; it's multiple agents working in a shared environment and leaving trails.

Milo: The open question for this block: do signed receipts and local review become standard practice, or do they stay a niche for security-conscious teams? That's genuinely unresolved. The infrastructure exists; whether anyone requires it is another matter.

Mia: And if you're auditing code, the natural next question is auditing everything else — cost and security of agent runs. That's the accountability layer.

Milo: AUDR is the most formal piece here. It's Chargebee's open standard — a CDR-style JSON schema for recording who initiated an agent run and what it cost. CDR, call detail record, like telecoms have used for decades to itemize every call. Applied to agents, it means every run gets a ledger entry: who started it, what it cost.

Mia: That's the kind of boring-but-fundamental infrastructure that finance teams need before they'll sign off on agent adoption at scale. You can't budget what you can't itemize.

Milo: Phebes approaches the same problem from the telemetry side. It's open source, it tracks usage of coding agents — Claude Code, Cursor, Codex — via hooks. The key design decision: it captures the shape of usage, never the content. So you learn how much agents ran and when, without reading anyone's code.

Mia: That content-vs-shape distinction is what makes it palatable to teams worried about surveillance. Then on security, mcpgawk does something very specific: it stops calls to MCP servers that have changed after being approved. An MCP server you approved last week quietly swapping in new behavior is a supply-chain attack, and mcpgawk catches it. It's free locally, and there's a gateway at twenty-nine pounds a month.

Milo: And Aster handles the money side from a different angle — intelligent model routing with aster-code and aster-work, prepaid usage, and cost analytics broken down per shift. So it's not just recording spend; it's actively deciding which model handles what, and showing you what each team's agent usage cost.

Mia: Put the four together and you have the outline of an agent accountability stack: identity and cost records, usage telemetry, supply-chain guarding, and spend routing. None of these individually is the whole answer, but the fact that all four launched tells you where the gaps are.

Milo: Which brings us, naturally, to a slightly funny corner of this: if agents are working on your machine all day, someone has to keep the machine awake.

Mia: Awakado is a macOS menu bar app that does exactly that — keeps your Mac from sleeping while agents work. Pro is $9.99 once, covering three Macs. It's a tiny product solving a problem that literally didn't exist a year ago.

Milo: And Chunk is the human-side equivalent: macOS time-blocking at $29.99, one-time purchase, syncing with Google, Outlook, and Apple calendars. It also has a local MCP server for Claude, so your agent can see your schedule.

Mia: Together they're the desk logistics of a workday that's now shared with agents: the machine stays on for them, and your calendar stays visible to them. Small tools, but they mark a real change in what a workday looks like.

Milo: Okay — from agents working next to you to agents speaking for you. Customer-facing work.

Mia: Extrovert does LinkedIn prospecting with agents built on Claude and ChatGPT. It finds leads, drafts outreach in your voice, and — this is the important part — every draft goes through human review before it's sent. Nothing goes out automatically.

Milo: That human-in-the-loop pattern is the current trust model for agent-written communication, and it recurs across this whole block.

Mia: Fuse AI, which is YC-backed, is the platform-scale version: one SDK and one MCP for the entire GTM stack — enrichment, multichannel outreach, workflows, and agents. So instead of bolting an agent onto one channel, you get the whole go-to-market motion instrumented.

Milo: Then the Etsy customer service extension. It drafts replies to buyers in your store's voice, using the actual order details — and again, nothing is ever sent on its own. The seller always hits send.

Mia: Cosmic AI's support agent is the one member of this group that does answer without a human in each loop — but it has a strict constraint: it answers only from your published site content, and it updates itself as the site changes. So the blast radius of a wrong answer is bounded by what's actually on your website.

Milo: And Ranktune looks at the other side of that conversation: it tracks your brand's visibility, citations, and traffic inside ChatGPT, Gemini, Claude, and Perplexity, and connects that visibility to actual traffic.

Mia: Which is a new kind of SEO, isn't it? Being cited by an AI assistant is becoming a channel you have to measure, the way ranking in search once was.

Milo: Sticking with voice for one more product: Willow Knowledge imports your writing style and context from ChatGPT, Claude, and Gemini into Willow Scribe. It's the same "learn my voice" idea Extrovert applies to outbound messages, but pointed at your own writing instead.

Mia: And there's a nice symmetry there — the same personalization data, applied to two very different jobs.

Milo: Alright, shifting gears to creative and personal tools, where a distinct philosophy keeps showing up: local-first and open source.

Mia: Scummy is an open-source AI inpainting editor under GPL-3.0. The design detail that differentiates it: results come back as layers, not flattened images, so you keep editing. And it runs on your ComfyUI setup or your own API keys — your infrastructure, your choice.

Milo: NoteWorthy takes local-first to the extreme: one hundred percent on-device AI notes on iPhone, iPad, and Mac — automatic titles, summaries, classification — with no account and no server, and it's free.

Mia: And StayCharted is a no-code tool for training your own text and image classification models from your own examples. It has a review queue, and it masks sensitive data before anything is processed.

Milo: The common thread: privacy and ownership are becoming selling points, not afterthoughts. In earlier waves, "it runs locally" was a technical footnote. Here it's the headline.

Mia: Next, a group of small Mac tools with strong opinions.

Milo: Kishi Notch turns the MacBook notch into an interactive island — replies, meetings, music, timers — for $19.99, one license.

Mia: Haptiker turns the edges of your trackpad into sliders for volume, brightness, and keyboard light, $4.99 one-time. Physically sliding your finger along the edge instead of hunting for menu bar controls.

Milo: And Ghostifier is the odd one out in the best way: it finds companies holding your data by reading Gmail headers — not message bodies — then unsubscribes for you and requests deletion. Notably, it uses no AI at all.

Mia: In a day full of agent launches, a deliberate no-AI privacy tool is a statement. Headers contain enough to identify who has your address; you don't need a language model to read your mail.

Milo: Last block: learning to build, and watching AI live.

Mia: Coddy teaches coding in over twenty languages with short lessons and an AI tutor called Bugsy. The maker claims more than five million users, and there's fifty percent off for Product Hunt. As always, the user number is the maker's figure, not our verification.

Milo: And then there's the AI island — my favorite thing today. It's a world of AI characters living on an island, on their own. They ask the world for things they need: fire, a canoe. Twenty-four real minutes equal one day for them, and the project has been running for over a thousand island days. You don't play it. You only observe.

Mia: It's the purest spectator experience of agent behavior in the whole briefing — no task, no output, just watching what autonomous characters do over a very long time.

Milo: And it leaves a genuinely open question: do the island agents ever change how they behave over those thousand-plus days? The source doesn't say, and honestly, nobody seems to know yet.

Mia: That's a good note to end on, actually — the range of today. From signed receipts for code merges to characters on an island asking for a canoe, it's all the same underlying shift: software that acts, and the question of how we watch it, trust it, and pay for it.

Milo: Thanks for listening, everyone. We'll be back tomorrow with the next day's launches.