
0918 | Agents, Testing, and a Toolbox of New Launches
Show notes
A rapid-fire tour of new AI-era tools: build and run AI agents, test them before they embarrass you, and a grab-bag of developer, Mac, and creator utilities — plus monetization for AI builders.
Timeline
- 00:00:04 Opening
- 00:00:45 AI that builds and delivers the whole thing
- 00:03:11 Testing AI agents and MCP servers before release
- 00:05:32 Infrastructure: data, cloud dev machines, billing, media models
- 00:08:59 Agents you talk to: iMessage and playful social usage
- 00:10:20 Native mobile from one codebase
- 00:11:40 Distribution: fundraising and competitor monitoring
- 00:13:12 Open-source: community benchmarks and your own knowledge base
- 00:15:01 Mac utilities: local-first, buy-once
- 00:17:53 Closing
Related links
- The Forge by Bob's Workshop
- AskDeck
- QAgent
- MCPJam
- NovaSynth by Noveum
- Axiom
- Bitrise Remote Dev Environments
- CREEM 2.0
- Higgsfield API
- Text Agent Store
- Die With Me
- Modaal for Android
- Pitchfire for Startups
- Figo
- Compute:Arena
- Opyt
- MacSentinel
- Blanc
- TinyKPI
- Zella
This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.
Transcript
Mia: Welcome back to the show, everyone. I'm Mia.
Milo: And I'm Milo. Today's episode is basically a tour of the AI product wave — but not just the shiny demos. We've got tools that build whole products, tools that test those products before anyone ships them, the infrastructure underneath, and even a few things where agents become social.
Mia: Right, and that's the thread I want to pull on. The interesting part isn't any single launch — it's that there's a whole supply chain forming. Someone builds the product, someone else tests it, someone else runs it, bills for it, and markets it. Let's start at the top of that chain.
Milo: So, The Forge. The pitch is pretty bold: you describe an idea in natural language, and an AI team plans it, develops it, deploys it, and then keeps operating it. And not just the code side — it includes CRM, payments, and brand management.
Mia: That's the whole company in a box, basically. And I want to be careful here, because that's a maker description, not a verified result. But if we take it at face value, the claim is that the distance between "I have an idea" and "I have a running product with payments" is basically a conversation.
Milo: And what makes it interesting versus the usual AI coding assistants is the "operate" part. Most tools stop at generating code. The Forge is claiming the whole lifecycle — including the boring business plumbing like CRM and billing.
Mia: The obvious open question is quality. What does the output actually look like when a real product meets real users? Edge cases, maintenance, changes over time — none of that is answered yet.
Milo: Now, AskDeck attacks the same problem — idea to finished artifact — but from the opposite direction. Instead of building a whole product, you contact an AI named Eric. You can literally call, text, or email it. A few minutes later you get back an editable PPTX plus a synced explainer video.
Mia: So it's not "build my business," it's "deliver my idea as something I can present." And the delivery detail matters — it's editable, so you're not stuck with a static render, and the video stays in sync with the deck.
Milo: Pricing is a one-time $29. That's a very different model from a subscription, which tells you they're betting on repeat use without locking people in.
Mia: Put the two side by side: both shrink the gap between an idea and a finished artifact. The Forge goes all the way to a live product; AskDeck stops at the artifact you'd use to sell the idea. Different depth, same instinct.
Milo: The honest caveat for both is the same: we don't know how good the outputs really are. That's the natural segue, though — if agents are building things, who checks the agents?
Mia: That's where it gets interesting. There's a whole crop of QA-for-AI tools. First one: QAgent. It automates QA for AI agents, and the setup claim is a two-minute webhook integration. It evaluates along eight dimensions, and you get 100 free evaluations a month.
Milo: Two minutes to set up is the key claim there, because QA tooling usually dies at the integration step. If it's genuinely a webhook and you're done, that lowers the barrier a lot. And the free tier gives you room to try it before committing.
Mia: Then there's MCPJam, which is a platform specifically for testing and evaluating MCP servers. And it's got a few layers: simulated user swarms, actual user testing, evals, and CI/CD gates.
Milo: The CI/CD gate part is what stands out to me. That's the discipline the rest of software engineering already has — you don't merge without passing checks. Bringing that to MCP servers means agent infrastructure can break loudly before release instead of quietly in production.
Mia: And the swarms idea — simulated users hammering your server — is the agent equivalent of load testing plus behavioral testing. You're testing how the thing behaves with many different users, not just whether it responds.
Milo: Third in this group is NovaSynth, and it's focused on voice agents. Before your voice agent goes live, it simulates real callers — interruptions, background noise, accents. It scores across more than 30 dimensions and automatically fixes the issues it finds.
Mia: Voice is where messy behavior really bites. A chat agent can wait for you to finish typing. A voice agent gets interrupted mid-sentence by someone on a bus. So simulating those conditions before launch is exactly the right test surface.
Milo: And notice the pattern across all three: agents fail in weird, unpredictable ways — so the ecosystem is building QA layers purpose-built for them. This is the same move that happened for web apps a decade ago, replayed for agents.
Mia: Okay, so the agent passed its tests. Now where does it actually run? Infrastructure time.
Milo: Let's take data first. Axiom is a machine data platform, and the headline specs are petabyte-scale schema-less ingestion and fully managed event storage that keeps everything.
Mia: "Schema-less" is doing real work there. If agents and apps are producing events faster than anyone can design schemas for them, you want to be able to ingest first and structure later. And "keep all the data" means you're not forced to throw things away before you know what you'll need.
Milo: Next, compute — specifically, dev machines. Bitrise RDE gives you cloud Mac and Linux development machines that start in seconds. And the detail I like: the hardware and cache are the same as their CI. So what you build on is what you build with.
Mia: That shared-cache detail is not trivial. One of the slowest parts of remote development is cold environments and cold builds. If your dev machine and your CI share hardware and cache, the "works on my machine" gap shrinks.
Milo: And they're pitching it as a way to drive parallel AI agents over MCP — multiple agents working at once on real machines. Pricing is $20 a month.
Mia: Then the money layer. CREEM 2.0 is a merchant-of-record platform aimed specifically at AI developers. It handles payments, tax, revenue share, and usage-based billing across more than 190 countries.
Milo: Usage-based billing is the important word there, because that's how most AI products charge. And merchant-of-record means they take on the tax and compliance burden, which is exactly the part solo developers hate.
Mia: And the twist: it ships with MCP and CLI, so an agent can literally open a store. We'll come back to why that's funny in a minute.
Milo: One more infra piece — media. Higgsfield API exposes more than 50 generative media models — things like Seedance and Kling — through a single asynchronous API. Pay per use, with Python and TypeScript SDKs.
Mia: The value here is the single API. If you want video from Kling today and something else tomorrow, you swap a parameter instead of rewriting an integration. Async matters too, since video generation is long-running.
Milo: Oh, and there's a fun little product that sits inside this infrastructure story: Text Agent Store. It's a marketplace of 137 iMessage AI agents — you add one as a contact and just text it, no account needed.
Mia: Which is a distribution layer riding on top of all this plumbing. And now the picture is complete: Axiom stores the data, Bitrise runs the agents, CREEM handles billing, Higgsfield supplies the media generation, and Text Agent Store is one of the storefronts.
Milo: The takeaway is that the infrastructure around agents is maturing. It's no longer just "a model and a prompt" — there's data, compute, billing, and media available as services.
Mia: Let's stay on that Text Agent Store idea a moment, because it connects to something else playful. Die With Me.
Milo: So Die With Me is an AIM-style buddy list — but instead of showing who's online, it shows your friends' Claude Code and Codex usage remaining. And if anyone drops below 20 percent, you can enter a chat room with them.
Mia: That's such a strange, delightful idea. It's a callback to the old instant-messenger era where your buddy list was ambient awareness of your friends. Here the ambient signal is how close your friend is to running out of AI usage.
Milo: And it's genuinely social in a way most dev tools aren't — agent usage itself becomes something you share and commiserate over. "You're at 15 percent? Get in the chat room, let's panic together."
Mia: Two different ends of the same phenomenon: agents being consumed where people already are. Text Agent Store puts them in iMessage; Die With Me puts their usage in your social life. Neither has a clear monetization story yet — and that's exactly the gap CREEM's agent-openable stores are aiming at.
Milo: Which is a nice loop: the marketplace needs billing, and the billing platform is literally building for agents that can open stores.
Mia: Okay, next subject — and this one is for anyone who's ever groaned at a cross-platform framework. Modaal generates native iOS and Android apps from one project. On the iOS side it's Swift and SwiftUI; on Android it's Kotlin and Compose. The logic is shared, but the UI is native on each platform.
Milo: That's the compromise everyone's always chasing. Web-view approaches are easy but feel wrong on at least one platform. React Native and friends get close but with layers in between. Modaal's pitch is: don't compromise — write once, get genuinely native UI on both sides, share only the logic.
Mia: Now, generated native code raises an obvious reliability question — who checks all that output? And this is where we can point back to what we covered earlier: MCPJam's CI/CD gates and testing discipline are exactly the kind of thing that keeps generated apps trustworthy.
Milo: That's the through-line of the whole episode, honestly. Generation is cheap now; the differentiator is verification.
Mia: Speaking of which — distribution. How do these products find customers? Two tools here.
Milo: First, Pitchfire. You upload your funding deck, and it matches you with VCs that fit and sends intro emails. You get three investor recommendations a day, and it's free.
Mia: Free is notable — fundraising tools are usually expensive. The daily cadence is smart too: three names a day is a pace a founder can actually act on, versus a dump of two hundred leads.
Milo: Second, Figo, which is competitor monitoring on a weekly rhythm. It watches your competitors' rankings, their pages, their ads on Google, Meta, and TikTok, their social presence, and — this is the new one — how often they show up in AI recommendations. Then every Monday you get a briefing.
Mia: That last one is the part I'd flag as genuinely new. "AI recommendations" as a monitored channel means brands now care whether ChatGPT-style tools mention them. That's a competitive surface that didn't exist a couple of years ago.
Milo: And notice how Text Agent Store and CREEM's agent-openable stores fit into the same picture: text itself is becoming the sales channel. Fundraising intros by AI, competitive intel by AI, checkout opened by an agent. Go-to-market is getting automated around AI products, not just the products themselves.
Mia: Now for the open-source corner, where things get more community-driven.
Milo: Compute:Arena is a community-driven benchmarking project for local AI. People submit benchmarks, there's an open-source testing harness, and it covers any hardware, runtime, or quantization. The leaderboard is already live.
Mia: The "any hardware, runtime, or quantization" scope is the point. Local AI is a wildly fragmented world — different chips, different quantizations, different runtimes — so a shared open harness means people can finally compare apples to apples on whatever they own.
Milo: And it's the community providing the benchmarks, not one company deciding what matters. That's a different trust model from vendor benchmarks.
Mia: Then Opyt, which is more personal. It turns your X bookmarks, your Substack subscriptions, blogs, and arXiv papers into a knowledge base. It's MIT-licensed, free, runs locally, and plugs into AI via MCP.
Milo: Local-first is the key phrase. Your reading pile is personal, and it's scattered across five platforms. Opyt pulls it into one place you control, on your own machine, and then your AI tools can query it through MCP.
Mia: And notice the MCP thread again — Opyt connects back to the testing and infrastructure discussions we had earlier. MCP is turning into the connective tissue of this whole ecosystem: servers, tests for the servers, dev machines driving agents over MCP, billing via MCP, knowledge bases exposed over MCP.
Milo: One thread, many products. Okay, last stop — the Mac utilities, and there's a clear pattern here, so let's do them as a group.
Mia: The pattern is local-first and buy-once. Four products, all fitting it.
Milo: MacSentinel first: it monitors your Mac, runs diagnostics locally, and cleans up. It has rules for more than 500 apps, and — a nice detail — deleted junk goes to the trash first so it's recoverable. The Pro version is a one-time $29, no subscription.
Mia: The trash-first deletion is a trust feature. Cleanup tools are scary precisely because they're irreversible. Making it recoverable signals "we're not reckless with your files."
Milo: Blanc is a free, MIT-licensed minimal desktop browser. Instead of a tab bar, you get a floating island, and it blocks ads and trackers at the network layer. Explicitly: no AI, no extensions.
Mia: "No AI" as a feature! That's almost a statement of philosophy at this point. Network-level blocking means every app benefits, not just the browser. And open source plus free means you can verify what it does.
Milo: TinyKPI: it shows your business KPIs — Stripe, PostHog, Google Analytics and more — right in the notch of your MacBook. Runs 100 percent locally, one-time $14.
Mia: Ambient dashboards are underrated. Instead of opening five tabs to check your numbers, the numbers are just... there, in the notch, all day. And local-only means your metrics aren't routed through someone else's server.
Milo: And Zella: it auto-edits screen recordings from your Mac or iPhone — removes silence, adds captions, does zooms — all running locally. Free up to 1080p, and Pro is a one-time $89.
Mia: Screen recordings are the raw material of tutorials and bug reports, and the raw material is always full of dead air. Auto-editing that locally, without uploading your screen to the cloud, is exactly the privacy-conscious version of that job.
Milo: So the wave is clear: privacy-first, local processing, and one-time pricing instead of subscriptions. Four different makers converging on the same model.
Mia: The honest unknown: how long does one-time pricing hold? Subscription revenue is predictable; one-time purchases need a constant stream of new customers. Either these developers find a way to keep it sustainable, or we'll see quietly added subscription tiers. That's the thing to watch.
Milo: And honestly that's a fair note to end the deep dives on — plenty of claims, plenty of interesting directions, and a lot still to be proven.
Mia: Quick recap of the shape of it, though: agents that build whole products, QA layers to test them, infrastructure for data, compute, billing, and media, new consumption and distribution surfaces, native mobile from one codebase, community benchmarks and local knowledge bases, and a wave of buy-once Mac software.
Milo: The thing I'll be watching is verification — every one of these stories eventually runs into "but does the output hold up?" The tools being built to answer that question might be the most important layer of all.
Mia: Agreed. Thanks for listening, everyone — we'll catch you next time.
Milo: Bye, everyone!