1010 | Agents, Open Source, and New Screens

||Download

Show notes

A fast tour of this week's launches: autonomous agents, open-source alternatives, coding and editing tools, voice apps, and new consumer hardware.

Timeline

  • 00:00:04 Opening
  • 00:00:32 Agents that act on their own
  • 00:06:40 Open-source takes on SaaS stacks
  • 00:09:20 Coding agents and agent plumbing
  • 00:12:43 Voice-first and media tools
  • 00:16:13 Hardware, worlds, and cloud computers
  • 00:18:56 Workspace bundles, play, and everyday tools
  • 00:20:51 Closing

Related links

This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.

Transcript

Mia: Welcome back to the show, everyone. I'm Mia.

Milo: And I'm Milo. Today we've got a big stack of launches to get through, and the thread that runs through basically all of them is that agents — AI agents — are suddenly acting like independent coworkers rather than features bolted onto an app.

Mia: Yeah, they've got their own identities, their own email addresses, even their own faces. So let's just start right there, because there's a meeting tool that's doing something pretty unusual.

Milo: Alright, so HeyPi. It's an AI meeting platform, and the claim is that it follows the conversation live while the meeting is happening, and it turns the ideas from that meeting into apps, decks, and websites in real time.

Mia: In real time. So you're not leaving the meeting, opening a design tool, pulling out your notes — you're literally in the call and the deliverable is being built while people are still talking.

Milo: That's the pitch, anyway. And you know, I think the interesting way to frame it is: who is this actually for? It's for people whose main job is talking — founders, consultants, product people — where the meeting itself is the work, and everything after the meeting is just manual transcription of what was said.

Mia: And versus existing options, that's the real difference. A normal meeting tool records, maybe transcribes, maybe gives you a summary. HeyPi's claim is it skips past the summary and goes straight to the artifact. That's a big jump.

Milo: It is a big jump, and that's exactly why we should be careful with how we describe it. Everything we know here comes from the maker's own description. Whether it can actually produce a usable deck or a working app from a messy, rambling conversation — that's the open question. Meetings are messy. The interesting ideas in a call are often half-formed sentences.

Mia: Right, and a tool that grabs half-formed sentences and turns them into an app could either be magic or produce very plausible-looking junk. We don't have independent testing on that. So treat it as a claim.

Milo: Okay, staying in that same world of agents that act on their own — let's talk about the one with literally its own email address.

Mia: Gemini Agent for Google Cloud. The description is a universal agent that plans on its own, and it works across Google Workspace, Microsoft 365, and MCP — that's the standard connector protocol lots of these tools are adopting. And it's described as a "KI colleague with its own email."

Milo: Own email is the part that sticks with me. Because an agent having its own email means it can send things, receive things, follow up — it's not just operating inside your account, it has a separate identity in your organization. That's a meaningful step from "tool" to "colleague."

Mia: It is, and it also raises the obvious questions. If an agent has its own mailbox and autonomous planning, who reviews what it sends? What's the audit trail? The source describes the capability, but the governance questions are wide open. For a company deploying this, that's not a footnote — that's the whole decision.

Milo: Agreed. And there's a companion piece here that's a bit lighter but in the same family: OpenCharm. This is an open-source companion that lives in the Mac notch — that little cutout at the top of the screen — and it gives your AI agents a face and a voice.

Mia: So the idea is that when an agent is working on your machine, you can actually see its status, and there's a face associated with it, and you control it with a key. It's making the agent visible and, in a sense, giving it a presence at the top of your screen.

Milo: I find this one interesting as a design statement more than a product. If agents are going to feel like colleagues, people apparently want them to feel like something — a face, a voice, a status you can glance at. Versus just a spinning progress bar in some terminal.

Mia: And it being open source matters too — you can see what it's doing. Which, given everything we just said about agent trust, is the right instinct.

Milo: Now, another desktop agent in this same group: OpenPilot. Free, MIT licensed, a desktop AI agent that reads and edits files, and you can plug in your own models through the OpenAI API. macOS support is coming soon, per the listing.

Mia: MIT license means you can take it, fork it, do whatever. And "bring your own model" means you're not locked into someone's subscription — you point it at whatever model you want to pay for.

Milo: So if you're keeping score in this group: Gemini Agent is the big-company, everything-connected, own-inbox version. OpenPilot is the free, local, files-on-your-machine version. Melete is the personal one — and Melete is worth slowing down on.

Mia: Melete. The pitch: a personal AI agent with its own computer. Not running on your laptop — it has a dedicated computer of its own, it has persistent memory, meaning it remembers things across sessions indefinitely, and you can swap which model powers it.

Milo: That combination is the point. Persistent memory plus always-on compute. Most chat assistants forget everything the moment you close the window. Melete's claim is that your agent accumulates context about you over weeks and months, running on hardware that's always available.

Mia: And swap-able model is important because right now everyone's worried about picking the wrong horse. If your agent's brain can be swapped out later, you're betting on the agent itself, not on one model provider.

Milo: Availability is a waitlist — the cloud version has a waitlist open. So again, this is described, not yet something we can verify at scale. But conceptually it completes the picture: identity, memory, persistent presence. The agents are moving from tools to colleagues, and all five of these — HeyPi, Gemini Agent, OpenCharm, OpenPilot, Melete — are different angles on exactly that.

Mia: Okay, let's shift. Same agentic world, but now the angle is open source going after paid software.

Milo: This is a theme I really like, because it's about who pays and who controls. Start with AgentDR. It's an open-source AI SDR — a sales development rep — and the claim is it replaces the Clay, Smartlead, HubSpot stack.

Mia: That's a notable claim, because those are three separate products most sales teams pay real money for. AgentDR covers outreach across email, LinkedIn, and WhatsApp, it's self-hosted, and you supply your own AI keys.

Milo: Self-hosted plus your own keys — that means your prospect data and your message generation never have to touch someone else's cloud, and you pay the model providers directly instead of a markup on top.

Mia: For the buyer, the trade-off is clear: you get control and presumably lower costs, but you take on the running of it. Self-hosted means someone on your team maintains it. That's the price of ownership.

Milo: Right. And in the same spirit, BrightBean Studio. It's an open-source alternative to Buffer — the social media scheduling tool — under the AGPL license, and it schedules to more than ten platforms, includes an MCP server, and explicitly has no per-seat pricing.

Mia: No seat pricing is the headline for teams. With SaaS schedulers, your bill grows with headcount. With BrightBean, it doesn't. And AGPL is a strong copyleft license — it's genuinely open source, not "source available."

Milo: And the MCP server is a nice connective detail — it means an AI agent could drive your social scheduling. Which brings us to the third one, and this is where the two halves of our conversation meet: Busabase.

Mia: Busabase is an open-source database and workspace built for AI agents. The point is shared storage where agents can work with data, but with access controls and change auditing.

Milo: That last part is the crucial one. If you're going to let autonomous agents write to your actual business data — the stuff in AgentDR's outreach pipeline or BrightBean's scheduled posts — you need to know who changed what. Busabase is answering the governance question we raised earlier with Gemini Agent, but in open-source form.

Mia: So the throughline of this whole group: cost and control shifting back to users. Instead of renting seats in someone else's cloud, you run the stack yourself and keep the audit trail. The counterweight is that you're now the ops team.

Milo: Let's move to coding agents — arguably where agents first became genuinely useful, and now the interesting stuff is the plumbing around them.

Mia: Together Link is a good place to start. It connects coding agents — specifically Claude Code and Codex — to open models running on Together AI. The pitch is lower cost, and it's free.

Milo: So the coding agent harnesses that people already use can now be pointed at open-weight models instead of only the frontier closed ones. For a team burning through tokens on big refactors, model choice becomes a cost lever.

Mia: And that's the real change versus the status quo: before, the agent and the model were effectively a bundle. Now they're separable.

Milo: Next, Refs. This one's fun. It's a library of over three thousand films, treated as editing blueprints, and it works through MCP with Claude Code, Codex, and Cursor.

Mia: So if you're building a video editing tool or an agentic video workflow, you can pull in "how a real film is structured" as a reference — cuts, pacing, scene structure — and the agent uses that as scaffolding for its own edit.

Milo: And Opposable takes agents somewhere genuinely new: it lets Claude Code or Codex actually type on real iPhones and Android phones. Physical devices. And payments are gated — they wait for user approval.

Mia: That approval gate is doing a lot of work. You're letting an agent drive your actual phone, so the one category of action where mistakes cost real money — payments — requires a human sign-off. That's the right shape for agent autonomy: broad capability, hard stops on the dangerous parts.

Milo: And then Zernio, which connects to what we said about Busabase and MCP. Zernio is marketing infrastructure: one API that reaches sixteen social platforms, plus ads, WhatsApp, iMessage, and telephony, and it exposes MCP so agents can use it.

Mia: So if Opposable gives agents hands on a phone, Zernio gives them a switchboard to the entire messaging and social world. Together they sketch out an agent that can actually do outbound work end to end — which loops right back to AgentDR from our open-source segment.

Milo: It does. And one more connection worth making here: Pine Computer, which is nominally in a different category but fits this discussion. It's a cloud computer built for AI agents, and instead of reading the screen pixel by pixel, it reads app structures directly.

Mia: That's a big deal if it works as described. Most computer-using agents are essentially looking at screenshots and moving a mouse. If the agent can read the actual structure of the app, the claim is it's two to five times faster at one twenty-fifth of the cost. It's in invite beta.

Milo: Those numbers are the maker's claims — we should say that plainly. But the approach, structural access instead of pixel-peeping, is a real architectural insight, and it's the same instinct we keep seeing: give the agent a cleaner interface to the world, and everything downstream gets better.

Mia: Alright, let's shift to voice and media — where all of this agentic energy meets the most human interfaces we have: talking and watching.

Milo: Murmur first. It's a private, voice-first note-taking app. The idea: you talk, and your speech becomes searchable memory. It has semantic search — meaning it finds things by meaning, not just keywords — and it extracts tasks from what you said.

Mia: The "private" part matters. Voice notes are the most personal data there is. And versus typing-based notes apps, the change is that capturing a thought costs zero friction — you just speak — and the retrieval side is handled by the AI.

Milo: The open question, as always with these, is how well semantic search actually works across months of rambling voice notes. But the concept — your spoken words as a queryable personal database — is a genuinely useful framing.

Mia: Phonable attacks a different speech annoyance: voicemail. It replaces iPhone voicemail. Callers get a text message instead, and you get a transcript plus an AI summary of the call.

Milo: So instead of listening to a two-minute voicemail to extract "call the dentist Tuesday," you read a one-line summary. The caller isn't ignored either — they get an SMS confirming their message landed. It's a polite swap for both sides.

Mia: And going a bit more infrastructure-flavored: Comcent. This runs voice infrastructure on your own SIP trunk — that's your own telephony line — with AI transcription and summaries, and it outputs vCons, which is a standard format for call records. Twenty dollars a month flat, unlimited users, and it's open source under AGPL.

Milo: Notice the pattern again — this is the second AGPL project today, after BrightBean. Flat pricing instead of per seat, open source, and you own the telephony line. For a small business, that's the same cost-and-control argument we made in the SaaS segment, just applied to phone calls.

Mia: Then OpenVids, which brings agents into video editing. It's an open-source, agentic video editor for Mac and Windows. You give instructions in chat, it edits on a timeline, and it has render QA — quality checks — plus checkpoints.

Milo: The checkpoints and QA are the parts that show this was designed for agent use, not just chat convenience. If an agent makes a long series of edits, you want to be able to roll back and to verify the render didn't come out broken. Those are the safety rails.

Mia: And it pairs conceptually with Refs from earlier — Refs gives the agent the film grammar, OpenVids gives it the hands to apply it. Again, plumbing and capability meeting.

Milo: Okay, our last group: hardware, worlds, and cloud computers — the physical and platform layer under all of this.

Mia: Start with Amazon. New Android tablets, starting at 229 dollars, with full Google Play — which is notable because Amazon tablets historically walled off Google's app store — plus integrated Alexa+, their newer agent, and a Kindle reading mode.

Milo: So the pitch is: a cheap Android tablet that's both a normal Android device and an Alexa+ agent surface, and it doubles as a Kindle. At 229 dollars, that's an aggressive entry point for putting an agent in more living rooms.

Mia: And Odyssey 3 is on the completely opposite end of the scale. It's a world model from Foundation, with real-time physics simulation. It scored 66.1 on Physics-IQ, which the source calls a best-in-class result, and it can be used to steer robots and cars.

Milo: Let's unpack why a "world model" matters. Most AI models predict text or pixels. A world model predicts how the physical world evolves — if the robot arm moves left, what happens next. Real-time physics means it can do that fast enough to actually control something, like a robot or a vehicle.

Mia: That benchmark number is again a maker-reported figure — a benchmark we can't independently verify here — but if the direction is right, world models are the piece that lets agents act in physical reality rather than only on screens and keyboards.

Milo: We've already touched Pine Computer — the cloud computer for AI with structural access instead of pixels, two to five times faster at one twenty-fifth the cost, invite beta. So the last one in this group is Thravik.

Mia: Thravik is a native Mac browser that is not Chromium-based — which is rare; almost every browser you can name is Chromium underneath — and it has a feature called Private Memory: pages you've read are remembered encrypted, on your device, and you can search them with Cmd+K.

Milo: Non-Chromium matters for a few reasons — independence from one company's engine, and a different performance and security profile. And Private Memory is a privacy-first answer to a problem AI browsers are currently solving badly: remembering what you read without shipping your browsing history to a server. Encryption on-device is the promise; exactly how the search works over encrypted data would be the thing to scrutinize.

Mia: Right — the claim is there, the mechanism deserves questions. Okay, let's land the plane with the quick hits that round out the week, and these are genuinely useful ones.

Milo: Ambiguous Workspace first — it's a bundle of eighteen AI-native productivity apps, free for up to five users. Eighteen apps in one package is a lot to evaluate, but as a zero-cost way for a small team to try AI-native tools across the board, it lowers the barrier to basically nothing.

Mia: Playground by Google Labs — you describe a game in text, no code, and it gets built. You can share it and then tweak it by chatting with it. Making game creation conversational is the whole idea.

Milo: Regunow handles regulatory research: you write the regulation question in plain language, it returns AI answers with authoritative sources, and it does document comparison, gap analysis, and structured reports.

Mia: The "authoritative sources" part is what makes that credible rather than just another chatbot — in compliance work, the citation is the product.

Milo: Staffcoder is a browser-based IDE for practicing real engineering, with over a thousand tasks and more than forty learning paths, free to start. So it's structured practice rather than tutorials — you write actual code in the browser.

Mia: And finally Rune, a Mac text editor built around one idea: a single scratch file, shared through iCloud across your Macs, with conflict-free merging — so notes just sync without any merge pain. Nineteen dollars, one time. No subscription.

Milo: One-time pricing feels almost quaint at this point, and honestly, after a whole episode about seats, subscriptions, and who controls the stack, ending on a nineteen-dollar one-time purchase is a nice bit of symmetry.

Mia: It really is. So the picture from today: agents with identities, memories, and their own inboxes; open source pushing back on the SaaS stack; agents getting hands on phones and switchboards to the real world; voice turning into searchable memory; and the hardware and platform layer starting to be rebuilt around them.

Milo: And through it all, the same caveats apply — most of what we covered is described by its makers, not yet independently verified, and the governance questions around autonomous agents are very much open. Thanks for listening, everyone.

Mia: We'll see you next time.