
0930 | AI Agents Everywhere: This Week's Launch Roundup
Show notes
A fast tour of this week's launches: AI agents at work, model infrastructure, local-first Mac and iOS tools, AI research, smart home, and finance apps.
Timeline
- 00:00:04 Opening
- 00:00:46 Model layer and agent runtimes
- 00:03:31 MCP agents for media and memory
- 00:07:53 Agents in business workflows
- 00:10:57 AI research and voice input
- 00:12:44 AI in finance and validation
- 00:15:15 Local-first Mac, iOS, and hobby tools
- 00:18:14 Smart home and hobbies stay local
- 00:19:32 Trust and auditing AI agents
- 00:21:11 Closing
Related links
- Claude Sonnet 5.5
- Hopscotch AI
- Codex Remote
- Arsaze
- Imejis.io
- LUCI Desktop
- Timeless Code
- Szept
- Paste 7
- Jotform Sign for ChatGPT and Claude
- ZenABM
- Semos.ai Manager Agents
- ColdIQ x Slack
- Engine Room Media
- Curie
- Supertake
- Would you pay?
- ShipHappens:
- Clink
- Flotnote
- GhostDeck
- Declutr
- Ricly
- GroupShelf
- Enter Space 7
- Pokébinder
- Gladys Assistant 5
- iFixAi
This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.
Transcript
Mia: Welcome back to the show, everyone. I'm Mia.
Milo: And I'm Milo. Today we've got a big pile of launches to get through, and the thread running through almost all of them is the same: AI is moving out of the chat window and into actual work — agents editing video, signing contracts, running ad campaigns, managing your team — while at the same time there's this counter-current of tools that want nothing to do with the cloud at all.
Mia: Right, and we want to be careful here. Everything we're going to say about these products comes from how the makers describe them. These are claims, not verified results, and we'll flag where things are open questions rather than pretending everything works perfectly.
Milo: So let's start at the foundation, because everything else builds on it: the model layer. The headline here is Claude Sonnet 5.5. Compared to Sonnet 5, it's supposedly more than thirty percent faster and up to thirty percent cheaper.
Mia: And the benchmark numbers are where the interesting claim sits. On Terminal-Bench 4.0, it scores 70.6 percent, which the maker describes as a big improvement. Terminal-Bench is about real terminal work, not trivia, so if that holds up in practice, that's meaningful for anyone running agents that actually execute code and commands.
Milo: The thing to keep in mind, though, is that a benchmark score is still the maker's own framing. Seventy percent on a hard benchmark also means thirty percent of the time it doesn't finish the job, which matters a lot when you're letting an agent loose on your machine.
Mia: Now, most people aren't picking one model in isolation — they're juggling dozens. That's where Hopscotch AI comes in. The pitch is a single API that gives you access to more than five hundred models.
Milo: And the pricing detail is actually the differentiator: they charge the provider's list price with no markup. So it's not "we're cheaper," it's "we're the same price as going direct, but you get one integration."
Mia: On top of that, they bake in fallback — so if one provider is down or rate-limited, you route to another — and spend management, which, if you've ever run a burst of agent jobs and watched the bill spike, you know is not a nice-to-have.
Milo: So Hopscotch is the plumbing layer. But how do you actually run these agents somewhere useful? That's the third piece: Codex Remote. It's an open-source Mac menu bar app, MIT licensed, and what it does is deploy Codex or Claude Code to cloud virtual machines using OpenTofu.
Mia: So instead of running your coding agent on your laptop where it can touch your actual files and your actual credentials, you spin up a disposable VM in the cloud and let the agent work there.
Milo: And the privacy story matters here: your authentication credentials stay stored locally. The tool doesn't hold them. For an open-source infrastructure tool, that's the kind of detail people will actually check.
Mia: So that's the stack: faster, cheaper models at the bottom; a routing and cost layer in the middle; and a runtime layer that keeps agents off your personal machine. Which raises the obvious next question — what are people actually doing with all this?
Milo: The flashiest answer is video editing. There's a product called Arsaze, and it's a multitrack video editor that AI operates through MCP — that's Model Context Protocol, the emerging standard for how agents talk to tools.
Mia: The claim they put forward is striking: in a demonstration, an edit that would take three hours and twelve minutes of human work was completed by Claude in twenty-three minutes. That's roughly a nine-to-one speedup, if the demo is representative.
Milo: And I want to be careful with that number, because it is a demonstration, not an independent benchmark. We don't know how complex that edit was, how much was template work versus creative judgment, or whether the output needed human cleanup afterward. But even with heavy caveats, the direction is clear — if an agent can drive a multitrack timeline, editing stops being a manual craft for the mechanical parts.
Mia: It also connects back to what we just said: this is exactly the kind of workload you'd want running against a faster, cheaper model, possibly on a cloud VM rather than your laptop, because video tooling is heavy.
Milo: Same MCP story, different domain: Imejis.io gives AI agents a design studio over MCP. It's template-based with what they call deterministic rendering — meaning the same input produces the same output, no surprises — and it supports twenty-seven components.
Mia: Deterministic is doing a lot of work in that sentence. Design tools are usually where agents produce beautiful chaos; a template approach with fixed rendering is basically saying "the agent picks and fills, but the result is predictable." For things like social graphics or slides, that predictability might be the whole point.
Milo: So Arsaze and Imejis are both "agent drives a creative tool over MCP," one for motion, one for static design. Now let's move from creation to something quieter but maybe more personal: memory.
Mia: There's a tool called LUCI Desktop — free, on Mac and Windows — that keeps a history of your screen and your meeting notes stored locally on your machine, and then exposes that memory to agents like Claude Code through MCP.
Milo: So imagine you're working with a coding agent and you want it to know what you were doing last Tuesday, what was said in the standup, what you had on screen. LUCI is the bridge that gives the agent that context — locally, not uploaded somewhere.
Mia: The obvious question is whether agentic memory is actually reliable. An agent with access to your screen history could be incredibly useful, or it could confidently misremember things. Nobody's answered that yet, and it's the open question hanging over this whole category.
Milo: A related take on the same problem is Timeless, which runs inside Claude Code itself. Claude sets it up — you don't configure it — and it turns recordings into meeting notes, and you can ask questions about the meetings and get answers with citations pointing back into the recording.
Mia: Citations are the key word there. Instead of trusting the agent's summary, you can check where a claim came from. That's a small design decision that addresses a big trust problem.
Milo: Two quick adjacent ones worth folding in here. Szept is a Mac voice-input tool for AI agents — it does dictation in over sixty languages, and it can pass the path of a screenshot you take into the agent, so the agent can look at what you just captured. It uses the Soniox API and ships with the Swift source code included, which is a nice transparency signal.
Mia: And Paste announced an Intelligent Clipboard feature using Apple Intelligence — it watches what you're working on and privately predicts what you'll want to paste next, on the Mac itself. It's a smaller idea than LUCI or Timeless, but it's the same theme: the machine anticipates from context rather than waiting for explicit commands.
Milo: So we've seen agents making things — video, designs — and agents remembering things. The next wave is agents embedded in actual business processes, and this is where the stakes get higher.
Mia: Let's start with signatures, because it's a very concrete workflow. Jotform Sign is an e-signature integration that lives inside ChatGPT and Claude. You can create a document, send it, sign it, and track its progress — all without leaving the conversation.
Milo: The significance is that the workflow stops being "switch to the signing app, upload, configure, wait." The agent is the interface. Whether businesses actually trust legally binding documents to flow through a chatbot is the open question, though.
Mia: From documents to ad spend: ZenABM is an MCP server plus an agent called Zena that automates LinkedIn advertising. You can create ads, optimize them, and pull reports, all from Claude or similar tools. Importantly, drafts are approval-gated — nothing goes live without a human sign-off.
Milo: That approval gate is the design pattern I keep noticing. These tools are all flirting with full autonomy, and ZenABM is the one explicitly saying "no, a human clicks the button."
Mia: Which brings us to Semos.ai Manager Agents — a group of AI agents aimed at managers. They learn from your meetings and then proactively nudge: hey, this feedback is overdue, this approval is stuck because nobody realizes it's waiting on them.
Milo: That's the proactive piece — it's not a dashboard you check, it's an agent that comes to you. Of course, an agent that chases your team members about overdue work is walking a fine line between helpful and annoying, and no source tells us where that line lands in practice.
Mia: Sales and marketing next. ColdIQ partnered with Slack on an integration where you tag a bot inside Slack and build out your go-to-market system — they're offering more than forty data providers and over a hundred GTM skills.
Milo: So the pitch is that prospecting and enrichment, which usually means tab-switching between data vendors, becomes a Slack message. The breadth claim — forty providers, a hundred skills — is the differentiator, but breadth is easy to claim and hard to verify.
Mia: And then there's the measurement side: Engine Room Media unifies analytics across multiple platforms and adds what they call an AI Producer that suggests recommended actions. They also offer media kits and conversion tracking.
Milo: So you can trace an arc here: agent creates the document, agent runs the ads, agent chases the approvals, agent reads the analytics and tells you what to do next. The through-line question — and I think the honest one — is how much of this actually runs without human review, and should any of it?
Mia: From business workflows, let's move to research, where the trust problem gets even sharper, because in research a wrong claim is worse than a slow result.
Milo: Curie is an AI research assistant built specifically for scientific literature. It searches multiple databases in parallel, and — this is the part I want to highlight — it verifies each claim against the original source text.
Mia: That claim-by-claim verification is the differentiator. Most research tools summarize; Curie is designed to check. It also supports PRISMA-style systematic reviews, which is the rigorous methodology used in medical and academic literature reviews.
Milo: If you know PRISMA, you know that's a high bar — systematic reviews have strict protocols about what counts as evidence. Claiming support for that workflow is a strong statement about who this is for: actual researchers, not casual users.
Mia: Notice the echo of Timeless: citations, verification, traceability. Across editing, meetings, and now science, the products that feel most serious are the ones building in ways to check the machine's work.
Milo: There's also a physical-input angle here. Szept, which we mentioned — voice input in sixty-plus languages, screenshot hand-off to the agent — is essentially about lowering the friction of talking to these research and work agents. Speaking instead of typing, and pointing at things instead of describing them.
Mia: And both Curie and Szept stress verifiability in their own way — Curie through source checking, Szept by shipping the source code. Trust is becoming a feature you can point at.
Milo: Now let's go somewhere spicier: money. Supertake is an AI investing platform, and notably it's a spinout from USV — that's Union Square Ventures, a serious venture firm, which is itself an interesting signal.
Mia: The mechanics: you write your investment thesis in natural language — your "view" on the market or a sector — and Supertake generates a portfolio from it. Then you can actually place the orders through Robinhood or Coinbase.
Milo: So the full loop is: idea, generated portfolio, executed trades. And I want to be very clear that the risk questions here are wide open. Translating a vague sentence into actual positions with real money, with no source telling us anything about risk controls, backtesting, or disclosure — that's a leap of faith we can't evaluate from the description.
Mia: Right. It's the most consequential product we'll talk about today in terms of personal stakes, and the one where "the maker says so" carries the least weight. If your thesis is wrong, the AI is just a fast way to be wrong.
Milo: From investing to validation: there's a free tool called "Would you pay?" that shows you startups one at a time and you swipe — would you pay for this or not. After ten items, it shows you how people who actually pay differ in behavior, and the conversion rates.
Mia: So it's a consumer research toy with a methodology angle — segmenting the people who swipe "yes" and actually would pay versus the tire-kickers. It's free, it's lightweight, and the honest read is that swipes are a weak signal compared to actual purchases. But as a quick gut-check mechanism, it's a clever format.
Milo: And staying on the business-tools side, there's ShipHappens, a Mac app for ASO — App Store Optimization. It uses Apple's own popularity data for keyword suggestions, generates app screenshots with AI, translates everything into thirty-nine languages in one batch, and handles replying to app reviews.
Mia: That's a nice bundle for indie developers, because ASO is one of those jobs that's critical but tedious — keywords, screenshots, localization, review responses. Whether Apple's popularity data plus AI screenshots actually moves rankings is the unverified part, but the workflow coverage is concrete.
Milo: Okay, big tonal shift now. We've spent all this time on cloud AI, and the other half of today's launches want your data to never leave your device. Local-first, offline, one-time pricing — that's the theme.
Mia: Start with Clink, an iOS custom keyboard. Fully themeable, custom layouts, and the key claim: normal input happens entirely on the device. No account required. And an Android version is planned.
Milo: A keyboard is the most intimate piece of software on a phone — it sees everything you type. So "on-device and no account" isn't a nice feature for a keyboard, it's the entire pitch.
Mia: On the Mac side, Flotnote is a floating Markdown note app. You hit Command-period and a note appears instantly, anywhere. Your notes are plain .md files, so they're compatible with Obsidian. Five dollars, one-time purchase.
Milo: And GhostDeck goes even further — it's an offline music player for the Mac with a classic deck interface and a ten-band EQ. It has these little low-poly 3D worlds that react to your music, it's twelve megabytes, and it's completely network-blocked. Twelve megabytes! Most apps use that for their splash screen.
Mia: File management next. Declutr organizes your Mac folders with one click, sorted by file type, and — importantly — it's rule-based and doesn't read the contents of your files. It has monitored folders and an undo feature, and the Pro version is eight dollars ninety-nine, one-time.
Milo: That "doesn't read your contents" line matters: it's sorting by type and rules, not by AI analyzing your documents. Ricly goes in the same drawer: seventeen file tools baked into the Finder right-click menu — large file sharing, PDF merging, background removal — all processed on-device.
Mia: Then there's GroupShelf: hold Shift, drag, and any app's windows turn into browser-style tabs. Nine ninety-nine, one-time, requires macOS 15 or later. Simple idea, and the kind of thing people either instantly need or never will.
Milo: And for storage, Enter Space 7 takes rclone — the tool that connects to over seventy cloud storage services — and makes them usable natively on Apple devices. It mounts them through Apple's FSKit, adds a photo timeline view, and point-in-time restore, which means rolling your storage back to an earlier moment.
Mia: So the unifying idea across Clink, Flotnote, GhostDeck, Declutr, Ricly, GroupShelf, and Enter Space 7 is: privacy, offline capability, and one-time pricing instead of subscriptions. It's almost a small anti-industry movement.
Milo: And that philosophy extends beyond work into the home and hobbies. Gladys Assistant 5 is a self-hosted smart home platform, and version 5 has Zigbee, Z-Wave, and Matter support built in — those are the major smart home radio standards — plus ninety-two integrations.
Mia: And notably, no YAML. If you've ever wrestled with Home Assistant configuration files at midnight, "no YAML" is a real selling point. Everything runs locally — your lights and sensors don't phone home.
Milo: The open question is polish. Self-hosted hubs historically trade control for friction, and we don't know whether Gladys 5 closes that gap with the mainstream commercial hubs.
Mia: And one pure hobby item to end the local section on a light note: Pokébinder. It's free, no account, and it plans your Pokémon card binder — it prints only the cards you don't own yet, at real size, six point three by eight point eight centimeters.
Milo: Which is a genuinely thoughtful detail — printing the missing cards at actual binder-pocket size so you can see exactly what your collection needs. Free, account-less, delightfully niche.
Mia: Which brings us to the last thread, and honestly the right note to end on: trust and auditing. Because everything we've described today — agents editing video, chasing approvals, executing trades — assumes someone can check the machine.
Milo: iFixAi is an independent AI audit service. It runs inspections across sixty-nine categories with two hundred fifty items, and its job is to detect, explain, and help fix inconsistencies in AI agents.
Mia: So if Arsaze's nine-to-one editing speedup or Codex Remote's cloud agents are the acceleration, iFixAi is the brakes and the inspection lane. As adoption grows, this counterpart role only gets more important.
Milo: The open question is pace. Agents change monthly, benchmarks change quarterly — can an audit framework with 250 fixed items keep up with tools that mutate that fast? Nobody's answered that.
Mia: So here's the shape of today, without making it a list: the model layer got faster and cheaper, agents got hands — in editors, in Slack, in your meetings — the business world is negotiating how much runs without a human, researchers are building verification into the loop, a whole parallel universe of tools wants to stay on your device, and auditing is emerging to keep the whole thing honest.
Milo: And the honest through-line for you as a listener: every claim we relayed today is the maker's claim. The demos are demos, the benchmarks are self-reported, and the most interesting products are the ones that already anticipate that skepticism — with citations, local storage, source code, approval gates, and audits.
Mia: That's the episode. Thanks for listening — we'll catch you next time.
Milo: Bye, everyone.