1003 | Harnesses, Satellites and Stratego Kings

||Download

Show notes

From SaaS becoming harnesses around models to TPUs in orbit, this episode looks at how AI is reshaping software, hardware, and privacy — plus local inference engines, agent databases, a Stratego-playing AI, exoplanet imaging, Zig's new release, a Gleam blogging pipeline, von Neumann lore, and museum visits.

Timeline

  • 00:00:04 Opening
  • 00:02:53 AI reshapes the software business
  • 00:04:55 Apple locks down the Mac as agents arrive
  • 00:07:01 Local AI hardware: from gadgets to orbit
  • 00:10:31 AI research: reading hidden states
  • 00:13:24 Developer tooling: Zig 0.17 and a Gleam blog pipeline
  • 00:15:37 Culture, museums and von Neumann lore
  • 00:19:23 Closing

Related links

This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.

Transcript

Mia: Welcome back to the show, everyone. I'm Mia.

Milo: And I'm Milo. Today we've got a thread that runs through almost everything we're going to talk about: machines that act. Agents that write code, agents that touch your files, agents in orbit, agents playing board games against world champions.

Mia: Right, and around that, the humans trying to keep up — with new OS controls, new tooling, and, honestly, a few museum trips to decompress. Let's start with the story that got the sharpest arguments going: the idea that every SaaS company is about to become a harness around a model.

Milo: So Shrivu Shankar's argument, as laid out in the piece, is that any SaaS business will end up wrapping a model — the harness being the scaffolding you build around the model to make it actually do the job — and crucially, that harness is the core of your competitiveness, so it should never be outsourced.

Mia: And the comment section immediately split. One camp basically said: yes, finally someone said it out loud. If the model is a commodity, the differentiated part is everything around it — the workflow knowledge, the integrations, the guardrails, the eval loops. Those people argued that founders who treat the model vendor as their entire product are building on rented land.

Milo: But there was real pushback. A number of commenters said the framing is too tidy. They pointed out that harnesses are hard to defend precisely because they're made of prompts, orchestration code, and glue — things models themselves are getting better at generating. So if the harness is just more software, isn't it also the part the next model eats? That was a genuinely uncomfortable question in the thread, and I don't think it got resolved.

Mia: Right — one commenter's counter was that harness work isn't generic glue, it's accumulated domain feedback. You learn, from real customers, where the model fails, and you encode that. That learning loop doesn't ship with the base model. But another commenter countered that labs could capture that same feedback directly if they ever went vertical — which is exactly the fear.

Mia: So you have people saying "the harness is the moat" and people saying "the moat is temporary" in the same thread, and both had reasonable cases.

Milo: The strongest firsthand-ish note was from people who said they'd already restructured their teams around this — model work centralized, harness work treated as the product. Whether that's wisdom or fashion, I guess time will tell.

Mia: And that fear about labs going vertical — the infrastructure stacking on top of infrastructure — leads beautifully into the Supabase story. Supabase acquired Turso, and the stated picture is that Supabase keeps focusing on Postgres while Turso keeps focusing on SQLite, with the combined goal being database infrastructure built for agents.

Milo: The discussion there was less about whether the deal makes sense and more about what "databases for agents" even means. Some commenters read it as: agents need a different shape of database access — tons of small, fast, isolated stores, maybe embedded, hence the SQLite side, versus the shared, heavy relational side Postgres covers. Others were more skeptical, saying a database is a database and "for agents" is partly marketing.

Mia: But the people who took it seriously made a decent case: if you imagine millions of autonomous agents each needing their own scratch state, the per-instance economics of SQLite versus a big shared Postgres become a real architectural question. The unresolved part, which several commenters flagged, is whether consolidation like this helps or hurts that future — one buyer controlling both axes of the answer.

Milo: And notice how this connects back to the harness debate. Supabase is, in a sense, building harness-layer infrastructure. If the harness is the moat, owning the data layer the agents run on looks like a pretty deep part of that moat.

Mia: Which raises the question the plan of the day keeps circling: who owns the agent infrastructure stack? Is it the labs, the database vendors, the OS vendors, or the app companies? Nobody in these threads claimed to know. It's genuinely open.

Milo: Okay, from the money question to the trust question. Apple tightened full disk access on macOS, adding additional controls, and the stated reason is the growing privacy risk from autonomous AI agents.

Mia: The comments here were unusually sympathetic to Apple, which surprised me a little. The mainstream view was: agents do exactly the dangerous thing — they request broad access, they act without a human in the loop per action — so the permission model built for human-driven apps just doesn't map anymore. Full disk access granted to a helpful utility is one thing; granted to a semi-autonomous agent loop is another.

Milo: There was disagreement though. Some commenters worried that more prompts and more controls just train users to click through — permission fatigue. Their argument: every added dialog weakens the signal of the dialogs that remain. The counterargument was that the alternative, no friction, is worse now that software can act at machine speed. Nobody had a clean answer for how you scope permissions to an agent that needs broad access to be useful.

Milo: That's the open question: narrow scopes make agents useless; broad scopes make them dangerous.

Mia: And then the companion piece from Apple: Pass Designer, in beta on macOS 27, for creating and previewing Wallet passes in real time, with semantic tags and validation. Commenters read this as the same story from the tooling side — Apple is building developer surfaces for a world where software generates artifacts on the fly. Real-time creation and built-in validation means less human fiddling in the loop.

Milo: So the synthesis some commenters landed on — carefully, not claiming it's Apple's official line — is that OS vendors are rebuilding both the trust layer and the tooling layer at the same time, because agents are about to touch sensitive data whether we like it or not.

Mia: From trust on the desktop to hardware on the desk — and then in the sky. Meta launched Muse Gadgets: open source hardware, built on ESP32 and Raspberry Pi, for connecting the Muse AI to screens, sensors, and everyday objects.

Milo: The reaction split into two flavors. The enthusiasts said open hardware is the right move — if ambient AI is going to live in your home, it should be inspectable, hackable, and cheap, and ESP32 and Raspberry Pi are exactly the right boring, proven platforms for that. Screens, sensors, everyday objects — that's an ambient computing play, not a smart-speaker play.

Mia: The skeptics asked the obvious question: who is this for? If it's a developer reference design, great. If it's meant to be a consumer category, open hardware historically struggles there. And some wondered aloud whether Meta's incentive is to commoditize the gadget layer so the value accrues to the assistant — which is, funnily enough, the harness argument again, one layer down: the gadget is the harness, the model is the commodity.

Milo: Then DwarfStar 4, from antirez — a C inference engine running DeepSeek V4, GLM 5.x, and Qwen3.8 locally, on Metal, CUDA, and ROCm. Three backends, one small C codebase, models running on your own machine.

Mia: Commenters loved the scope discipline. One view: hand-written C for inference is a statement — that with enough care, you don't need a mountain of framework code to run serious models locally. The other view: this is a delightful hobby project and the real production path is the big runtimes; don't confuse elegance with scale.

Mia: The interesting middle position was that projects like this matter as leverage — they keep the big frameworks honest and they make "local" a real option rather than a slogan.

Milo: And then the scale goes vertical: Google's Project Suncatcher — TPUs in space. The prototype satellite is in orbit, launched via SpaceX's Transporter-18, and contact has been confirmed. That's it, that's the fact — a TPU-carrying test satellite is up there and talking.

Mia: The thread on this had genuine wonder and genuine doubt in equal measure. The doubters: thermal management in space, launch costs, radiation, maintenance — datacenters in space face problems that don't have cheap answers, and a prototype proves none of them solved. The supporters: every one of those problems gets cheaper with a working test article in orbit, and power — raw sunlight — is the one input a datacenter needs that space provides without a grid.

Milo: But look at the arc across these three: Meta puts AI on your desk objects, antirez puts inference on your GPU, Google puts compute above the atmosphere. It's the same question — where does the compute live relative to the human? — answered at three different altitudes.

Mia: Okay, from hardware to the research frontier, and specifically the frontier of hidden information. Ataraxos — a collaboration across CMU, MIT, NYU, and Stanford — beat the best Stratego player 15 to 1, using 16 GPUs.

Milo: And the detail that made the thread light up: a second network whose job is to guess the hidden pieces. Stratego is imperfect information — you can't see your opponent's ranks. So they didn't just build a stronger player; they built a dedicated inference machine about the opponent's hidden state, and then, presumably, played accordingly.

Mia: Commenters drew the obvious line to poker and beyond: perfect-information games were conquered years ago, and imperfect information was the stubborn remainder. The debate was about what this actually proves. One side: a 15–1 score against the best human is a landmark, full stop. The other side: Stratego's hidden information is structured and bounded — a fixed set of piece types — so it's a step toward general hidden-state reasoning, not a proof of it.

Mia: Real-world agents face hidden state that doesn't decompose into a known piece list.

Milo: The unresolved question people kept asking: how much of the win was the second network, versus raw scale from those 16 GPUs? Would the architecture win at a quarter of the compute? Nobody could say from what was published.

Mia: And there's a lovely parallel from an entirely different field: twelve years of telescope images of HR 8799, released together — four exoplanets, each more massive than Jupiter, imaged in orbit at 133 light-years.

Milo: Twelve years. That's the part commenters fixated on — the patience. You can't resolve a planet's orbit from one image; you need a long baseline of observations, years of them, and then the motion pops out of what looked like noise.

Mia: Which is the same epistemology as the Stratego network, in slow motion: hidden state revealed by accumulating evidence over time rather than seeing directly. One does it with a second neural network; the other does it with a decade of telescope time. Both are, in a sense, guessing the hidden pieces and checking against reality.

Milo: Somebody in the comments made exactly that connection — that inference about unobservables is the shared project of modern AI and modern astronomy, just at different timescales. That's the kind of cross-pollination that makes comment sections worth reading.

Mia: Alright, tooling time, and the headline here is Zig 0.17.0 — five months of work, 206 contributors, 925 commits, a reworked build system, and incremental ELF compilation for x86_64-linux.

Milo: The build system rework got the most commentary. The argument in favor: build systems are where language projects quietly lose people, and Zig treating the build as a first-class, language-level concern is ahead of the curve. The skeptics: reworking the build system is disruptive mid-project, and every release that touches builds breaks someone's CI.

Milo: Both sides agreed, though, that incremental compilation is the feature that changes daily feel — edit, compile, keep going, without a full rebuild each time.

Mia: Worth noting the scoping: incremental ELF compilation is for x86_64-linux specifically. Commenters who follow the project closely took that as the pattern — land one target, prove it, then widen. That's a discipline argument, and it seemed to win the room, though nobody claimed the roadmap was confirmed.

Milo: The smaller but charming companion: a developer blogging with Gleam, Org-mode, and Pandoc — writing in Emacs, converting through Pandoc, and a statically typed pipeline using Blogatto and Nix. The blog post's whole thesis is that even a personal site can be a typed, reproducible pipeline.

Mia: Commenters resonated with exactly that: the joy of a blog stack where the type system catches your broken frontmatter before it ships. The dissent was the usual — someone always says "I use markdown and a shell script and it's fine" — and the rebuttal was that for a site you maintain for a decade, reproducibility via Nix pays for itself. It's a micro-version of the Zig story, honestly: invest in tooling rigor for something you'll touch for years.

Milo: Which is a nice segue to the human layer — culture, debate, and a couple of field trips. The LessWrong piece on social reality in China got genuinely heated. The article's material touched on social conformity and status games in China, and the debate spilled into cultural diversity — whether the dynamics described are specific to one society or universal in different costumes.

Mia: And it got heated because both readings had passionate defenders. One side argued the piece illuminates something real about how social pressure structures behavior in ways Western individualist frames miss. The other side pushed back hard that any society's "social reality" gets exoticized when described from outside, and that conformity and status games are human constants.

Mia: There were firsthand perspectives in the thread from people who'd lived in multiple cultures, and even they disagreed about how much was cultural versus universal. That one's unresolved by design — it's a question about interpretation, not data.

Milo: Then the field trips. Three Dutch computing museums: the HomeComputerMuseum in Helmond, Bonami in Zwolle, and what's affectionately called a DEC barn — a collection of Digital Equipment machines.

Mia: The comments on the museum posts had a distinct theme: machines you can touch. HomeComputerMuseum in particular got praise for being a place where the computers are not behind glass. The DEC barn got wistful comments — people who worked on PDPs and VAXen reliving it. Somebody made the point that preserving working machines is different from preserving artifacts; a computer that can't run is a sculpture.

Milo: And the wildcard of the day: the Shimano Bicycle Museum in Kansai, praised specifically for privileging bicycle history over Shimano's own story — no sales floor, no self-promotion.

Mia: That detail sparked a mini-debate about corporate museums generally. The consensus-leaning view — and I say leaning, because there were holdouts — was that most company museums are marketing in a costume, and Shimano's restraint is the rare exception that makes the history more credible, not less. The holdout position was that even restraint is a curated narrative. Fair, but as museum experiences go, it read as one of the good ones.

Milo: And then the reading list: Halmos's "The Legend of von Neumann" from 1973, rediscovered — and commenters piling on with recommendations: "The Man from the Future" and "Turing's Cathedral."

Mia: The von Neumann thread had a tone of reverence. Commenters quoted and paraphrased Halmos's portrait — the speed of thought, the range, the way von Neumann seems to touch every field we've discussed today: computing, games, the very architecture everything runs on. And the Stratego story earlier? Von Neumann basically founded game theory. The thread didn't make that point explicitly, but you can draw the line.

Milo: The two recommended books got characterized as complementary — "The Man from the Future" as the modern biography treatment, "Turing's Cathedral" as the machine-and-institution story of the early computer era. Several commenters said read Halmos first for the voice, then the books for the depth.

Mia: So let's try to tie the bow, because there is one. Today: the harness debate says the value is in what wraps the model. Apple is rebuilding the trust layer for models that act. Meta and antirez and Google are deciding where the compute lives. Ataraxos is teaching machines to reason about what they can't see. Zig and a Gleam pipeline are the humans building careful scaffolding of their own. And von Neumann is lurking under all of it.

Milo: The through-line, if you want one: every layer is asking the same question — who owns the stack around the intelligence? The labs, the OS vendors, the database folks, the hackers with a C compiler and a Raspberry Pi? Every thread today was a different bet on the answer.

Mia: We don't have the answer either. But we'll be watching — and if these threads taught us anything, it's that the bets get settled in comment sections, compiler releases, and occasionally, twelve years of telescope images.

Milo: Thanks for listening, everyone. We'll see you next time.

Mia: Bye.