
0913 | AI Speed, Trust, and the Joy of Building
Show notes
A tour through this week's tech news: how AI is reshaping what it means to build software, the fight over how fast it should move, plus security, hardware, and internet stories worth your time.
Timeline
- 00:00:04 Opening
- 00:00:45 Pacing AI development
- 00:04:32 AI on real code and the money behind it
- 00:08:07 Should you still build it yourself?
- 00:09:49 Trust and verification: Signal audits and clipboard snooping
- 00:12:34 LG, spying claims, and media spin
- 00:15:05 The internet never forgets: Usenet rewind
- 00:17:07 Hardware archaeology: 8087 and Apple's Neural Engine
- 00:19:59 Tracing builds: buildprof
- 00:21:14 Programming languages: contracts and generations
- 00:24:58 Mapping the world in 15 minutes
- 00:26:17 Energy efficiency and AI-written journalism
- 00:26:52 AI agents selling their own labor
- 00:27:55 Mathematics: transformer theory and a claimed Navier-Stokes solution
- 00:31:10 Closing
Related links
- We must pace the frontier
- Everyone should slow down AI development except for me
- Fuck it, make it anyway
- Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
- Nvidia is the central bank of AI
- How Trail of Bits helps verify the integrity of Signal chats
- Linux Zoom client proactively reading everything written to X11 clipboard
- LG Says We're Fake News [video]
- LG responds to TV spying allegations
- Usenet rewind archive search engine
- Microcode in Intel's 8087 floating-point chip: the scale instruction
- Retrospectively Reverse-Engineering Apple's Neural Engine
- I made a build visualizer to understand Bun's compile times
- A few good ideas in programming languages
- Will There Be a 7G?
- Make your first edit to OpenStreetMap
- The worst spam emails: iLands AI agent hustle
- Europe's "Less" Is Doing More Than Anyone Gives It Credit For
- Navier-Stokes Announcement
- A Mathematical Framework for Transformer Circuits (2021)
This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.
Transcript
Mia: Welcome back to the show, everyone — I'm Mia.
Milo: And I'm Milo. It's been one of those news cycles where every story seems to be asking the same question from a different angle: who's checking the work? Whether it's AI labs calling for a brake, benchmark results on private code, spying TVs, claimed solutions to millennium math problems — underneath all of it there's this thread of trust and verification.
Mia: Yeah, and we're going to pull on that thread all episode. We'll start with the people building the AI asking to slow down, then follow the money and the benchmarks, then swing through privacy, hardware archaeology, mapping, and a few things that made us laugh.
Milo: Let's get into it. So the biggest story of the day — and it's a strange one — is Dario Amodei publishing a piece called "Pacing the Frontier."
Mia: Right, and the headline fact is that the CEO of Anthropic is calling for a global slowdown of AI development. That alone would be notable. But the substance is what people are chewing on: he proposes things like embedded evaluators — so third parties embedded into labs to assess capability and risk — plus industry coordination, and an actual slowdown of the frontier pace.
Mia: And the driver he points to is recursive self-improvement risk, the idea that once systems can meaningfully improve the systems that build them, the feedback loop gets dangerous fast.
Milo: And the discussion immediately splits into a few camps. One camp says: great, finally a lab leader saying out loud that the pace is a problem, and the embedded-evaluator idea is at least a concrete mechanism rather than vague vibes. The counterargument, which came up again and again, is the credibility problem. If you're the CEO of a frontier lab, "everyone slow down" reads differently depending on whether you're ahead or behind.
Milo: If you're ahead, a coordinated slowdown freezes your lead in place. If you're behind, it's a way to catch up without racing. Commenters were pretty divided on which of those applies here, and honestly nobody could resolve it, because it depends on facts about relative position that outsiders just don't have.
Mia: And there's a practical objection too: what does coordination even mean? Export controls exist, sure, but compute is international, and any binding agreement would have to include actors with very different incentives. Several people made the point that "industry coordination" among competitors is basically what antitrust law is designed to prevent in other sectors. So the mechanism he's asking for may be legally awkward even if everyone agreed on the goal.
Milo: Which connects to the unresolved question, which is genuinely open: nobody in the discussion had a worked example of what a slowdown regime would actually look like. Who measures the pace? Who enforces it? What's the compliance mechanism? The proposal identifies the problem — RSI risks — but the operational layer is a blank page.
Mia: And there's a perfect satirical companion piece that landed alongside it. Xe Iaso wrote an essay titled "Everyone should slow down AI development except for me."
Milo: Which is the whole tension in one sentence.
Mia: Exactly. The essay skewers the assumption everyone makes — that slowing down is a great idea in general, but that my work, my project, my company is the exception. It's the tragedy of the commons dressed up as engineering philosophy. And there's a fun detail: Iaso's site blocks visitors from US states via a firewall rule, which itself became a talking point — people found it both hilarious and a little on-the-nose, a personal act of selective engagement in a debate about collective restraint.
Milo: And the reason this pairing matters is that the satire isn't mocking the concern — recursive self-improvement is a real risk Amodei is pointing at — it's mocking the incentive structure. Even people who genuinely want a brake can't help carving out exceptions for themselves, because unilaterally slowing down feels like unilateral losing.
Mia: So who is actually racing hardest right now? That's a natural segue, because the answer may be "the benchmarks."
Milo: So there's a benchmark called Real-SWE, and what makes it interesting is that it evaluates AI models on private enterprise codebases — not the public repos that models have effectively seen in training. That's a big deal, because a lot of the skepticism about coding benchmarks is contamination: if a model has memorized the repo, the score is theater. Testing on private code tries to strip that out.
Mia: And the current leader is Fable 5.1 running through Claude Code, with a 38.8 percent resolution rate.
Milo: And that number sparked a genuine split in the discussion. One reading: nearly 39 percent on real, private, messy enterprise code is astonishing — a few years ago this was a rounding error, and real work is often ambiguous, under-specified, entangled with legacy decisions. The other reading: flip it around, and the model fails on more than 60 percent of tasks, which means in practice a human still has to review everything, which changes the economics.
Milo: The optimists and skeptics were both using the same number to tell opposite stories, which is usually a sign the honest answer is "it depends on the task mix."
Mia: There was also the deeper methodological question: private enterprise codebases from which companies? What kinds of tasks? A benchmark on private code solves the contamination problem but creates an opacity problem — you can't independently audit it the way you can an open dataset. If the benchmark itself is a black box, you're trading one trust problem for another.
Milo: Which is a good bridge to the money side, because The Economist published a piece calling Nvidia the "central bank of AI."
Mia: The analogy being: Nvidia effectively controls the money supply of the AI economy — the compute — the way a central bank controls currency. And the discussion around it was lively. Some people found the analogy genuinely illuminating: compute scarcity is the binding constraint, and whoever allocates it sets the tempo of the whole field, which loops right back to the pacing debate we opened with.
Milo: Others pushed back hard on the analogy's limits. A central bank prints its own currency and sets interest rates; Nvidia sells hardware and doesn't directly control what the chips do once sold. One of the fun refinements people offered: if Nvidia is the central bank, then TSMC is the mint — the actual place where the currency is physically produced.
Milo: Which, people noted, adds a geopolitical wrinkle the central-bank frame glosses over: the "mint" sits in Taiwan, and that single point of concentration worries people more than anything Nvidia's pricing does.
Mia: The unresolved question there is whether compute concentration is a temporary artifact of the current training paradigm or a structural feature. Nobody had a confident answer. But capability and capital are clearly moving together, and that raises a question a lot of people sat with: with all this delegation to models, is there still a case for building things yourself?
Milo: Which brings us to Joel Auterson's essay, and I'll just say the title upfront because it is the thesis: "Fuck it, make it anyway."
Mia: So the setup: Auterson writes about the three paths available to someone who wants to build something in the AI era. He weighs them honestly — and this is what made the piece resonate — he doesn't strawman the AI-assisted route. The point of the essay is that after weighing the options, he chooses the hard kind of making. Doing it by hand. Slowly. For the joy of it.
Milo: And the reaction was mostly warm, but with a real fault line. One group read it as a personal credo, not a policy position — "this is what makes life meaningful for me" — and found it refreshing precisely because it doesn't claim universality. The other group noted that this is a luxury position: you can only afford the slow, joyful path if you have slack — time, money, security.
Milo: If your job or your deadline depends on shipping, "make it the hard way because it's joyful" is advice that lands very differently.
Mia: There's also the subtler point people kept circling: hand-making things is how you build the judgment to evaluate what models produce. If you've never built the thing yourself, you can't tell a good AI output from a confident-sounding bad one. So the "pointless" hand-building isn't pointless at all — it's training the reviewer. That's the counterpoint to pure delegation, and honestly it's the strongest argument in the whole discussion.
Milo: Which flows beautifully into our next block, because judgment and verification are exactly what the security stories today are about. First, some genuinely good news: Signal's Automatic Key Verification.
Mia: So Signal is building key transparency — the idea being that instead of you manually comparing safety numbers with your contacts, the system uses Merkle trees so the key directory itself becomes publicly auditable. You can verify that the key you see for someone is the key that's logged, without trusting Signal's word for it. And Trail of Bits, the security firm, is operating as one of three external auditors on this system.
Milo: Why external auditors matter: key transparency only works if someone independent checks that the logged tree matches reality. If the service provider could silently insert a key, the whole transparency property evaporates. Having multiple external auditors is the engineering of trust — trust as a property of the system rather than a promise from the company.
Mia: And here's the contrast that made this story sting. Simon Tatham — the PuTTY author — found that the Linux Zoom client proactively reads the entire X11 clipboard. Including, when you have one open, your password manager's data.
Milo: And the technical detail people fixated on: on X11, any application can request the clipboard contents at any time — there's no permission system, no user prompt. So a client that "proactively reads" the clipboard means it's continuously polling, and whatever lands in your clipboard — a copied password, an SSH key, a one-time token — has passed through Zoom's process.
Mia: The discussion was fairly grim about this, because there's no clean fix at the X11 layer; the protocol simply wasn't designed with that boundary in mind. The comments ranged from "this is why I read clipboard managers with dread" to people sharing their own defensive setups. The honest takeaway was: your threat model has to assume any GUI app on an X11 session can see your clipboard, and most people's behavior doesn't reflect that.
Milo: And that's the asymmetry worth naming: Signal is spending enormous engineering effort to make trust verifiable down to the Merkle tree, while simultaneously, ordinary desktop software is reading everything you copy without telling you. Trust is engineered in one place and casually undermined in another, on the same machine.
Mia: Which is the perfect setup for the LG saga. So Gamers Nexus published reports about LG TVs — the reporting indicated the TVs log ambient dialogue — and LG's response was essentially to call the reporting fake news.
Milo: And Gamers Nexus came back swinging with a video titled "LG Says We're Fake News." That's the response piece. The core dispute is over a word: "continuously." LG's position, per the coverage, is that ACR — automatic content recognition — is optional, and the characterization that the TV continuously monitors is disputed. Gamers Nexus's position is that the logging of ambient dialogue happens regardless, and that "optional" is doing a lot of quiet work in a settings menu most users never open.
Mia: The discussion around this was less about LG specifically and more about the language of privacy disclosures. Commenters pointed out that "optional" and "off by default" and "continuously" are words that can each be technically defensible while being misleading in combination. Does "optional" mean you can turn it off after it already collected data? Does "not continuous" mean it samples periodically, which is functionally the same to the user?
Mia: Nobody was defending LG's phrasing; the disagreement was about how much malice versus how much sloppy drafting.
Milo: And there's a meta-layer that got its own share of attention: YouTube appears to be A/B testing video titles — meaning different viewers see different titles for the same video. Which matters here because a controversy story's title can be experimentally optimized while the underlying facts are still contested. Two people can genuinely see two different framings of the same claim and then argue past each other about what the video even says.
Mia: So the stack reads: contested telemetry language from a TV maker, adversarial fact-fighting between a reporter and a corporation, and a platform silently varying how the dispute is presented. Each layer adds ambiguity. From spying TVs, it's a short hop to surveillance of a much older kind — your own words from twenty years ago.
Milo: Usenet-Rewind. This one is simple to state and unsettling to think about: a searchable archive of over a billion Usenet messages, running from 1981 all the way to today.
Mia: So anyone's posts from the eighties, nineties, early two-thousands — arguably the most candid writing many people ever did under a pseudonym they thought was semi-anonymous — is now one search box away, linked to whatever identity they have now.
Milo: The discussion had two distinct registers. The nostalgic one: people went digging, found old threads from their own pasts or from legendary figures, and there was genuine delight in the archive as a historical record — technical arguments from 1991, software history, dead communities preserved. The archive as a museum of the early internet was broadly celebrated.
Mia: And then the other register, the privacy one: the people who posted in 1994 did not consent to being indexed for eternity in a world where a single search connects "their old handle" to "their current full name." Several commenters made the point that this isn't hypothetical — people have lost jobs over resurfaced old posts, and Usenet was often far more unfiltered than any modern platform.
Mia: There was no consensus on what, if anything, should be done — a takedown request to a decentralized archive of a billion messages is a complicated ask — but the discomfort was real and unresolved.
Milo: The interesting philosophical split: some argued the internet was never anonymous in the first place, just pseudonymous, and people should have known. Others argued that expecting 1994 posters to anticipate modern search and identity-linking is unreasonable — nobody imagined that world. That argument went nowhere, productively.
Mia: From old software archives to old hardware — and this is my favorite block of the day, because we've got two genuinely beautiful reverse-engineering deep dives.
Milo: First, Ken Shirriff, who has made a career of reading silicon like text, reconstructed the microcode of the Intel 8087 — the floating-point coprocessor from 1980. And the specific finding: the FSCALE instruction, which scales a number by a power of two, executes over 140 microinstructions across three levels of subroutines.
Mia: Which is the kind of number that stops hardware people in their tracks. FSCALE sounds like it should be "shift the exponent, done." Instead it's a genuine little program running inside the chip, with nested subroutine calls — in a microcode engine from 1980. Shirriff's work essentially shows that even the "simple" operations on this chip were implemented with a surprising amount of software-like structure.
Mia: The appreciation in the discussion was near-universal: this is archaeology of a kind that's about to become impossible, as modern chips stop being decappable and readable.
Milo: Second deep dive: Eileen Yoon reverse-engineered Apple's Neural Engine — the ANE in the M1. And the headline finding: the ANE was designed around the data-flow patterns of convolutional neural networks, and it's genuinely poorly suited to transformers.
Mia: Which retroactively explains so much. For years people asked why Apple kept pushing the ANE while most modern ML workloads migrated to transformer architectures the ANE seems awkward on. The answer from the silicon: it was architected for a previous generation of models. And the punchline is the M5, where Apple folds the ANE into the GPU — essentially conceding the separate accelerator concept didn't pay off for the current workload mix.
Milo: The discussion here was partly admiration, partly a lesson in hardware strategy. The takeaway several people drew: when you bake an architecture into silicon, you're making a bet on which model family will dominate for the next several years — and Apple's bet, made when CNNs were the story, was wrong. That's not incompetence; it's the inherent risk of custom silicon.
Milo: And it quietly connects to our opening topic: hardware is the slowest-moving layer, and it can't "slow down" or pivot quickly the way software can.
Mia: From chip archaeology to something more immediately practical for developers: build profiling. Lalit Maganti built buildprof, an open-source tracing tool that visualizes your build as a process-tree timeline.
Milo: And the demo use case is Bun — the JavaScript runtime — where Maganti used it to compare Bun's Zig build against its Rust build times. Which is exactly the right kind of case study, because "which language builds faster" claims are usually vibes, and having a timeline that shows where the wall-clock time actually goes turns the argument into something you can look at.
Mia: The discussion around tooling like this was pretty practical: the perennial complaints about builds being black boxes, the observation that most teams don't profile their builds even when builds eat hours of collective time per week, and the general point that making something visible is the precondition for fixing it. The open question was adoption — build profiling tools have existed in various forms, and the hard part is rarely the visualization; it's getting teams to actually look at it.
Milo: Which leads nicely into language design, because the same theme of "ideas that sound good versus ideas that survive contact with reality" shows up there too. There's a blog post surveying programming language ideas — flow typing, borrow checking, contract programming.
Mia: Flow typing, quickly, is when the compiler's type knowledge refines as control flow narrows it — you check "is this null" and the type system knows it's non-null afterward. Borrow checking is the Rust approach to memory safety at compile time. Contract programming is preconditions and postconditions on functions, asserted as part of the interface.
Milo: And here's the delicious irony the source itself hands us: the blog's example code for contract programming — written in D — contained a typo in its postcondition. The example demonstrating that contracts catch your bugs contained a bug. The comments had a field day, and to be fair, several people made the constructive point: this is exactly the failure mode contracts are supposed to catch, so in a strange way the typo is an argument for the feature, not against it.
Milo: Just an argument that humans — including people writing about verification — don't verify as reliably as they think.
Mia: The substantive discussion around the survey was about which of these ideas actually crossed over into mainstream languages. Flow typing: yes, quietly, everywhere now. Borrow checking: influential far beyond Rust, showing up in designs that don't even call it that. Contracts: still niche, still waiting for its moment. So the survey's implicit question — why do some ideas take decades to diffuse — is a real one, and there's no tidy answer in the thread.
Milo: Now here's a pairing you don't see every day: naming ideas and naming generations of mobile networks. An arXiv paper asks, seriously: will 7G exist?
Mia: And the proposal is more interesting than the clickbait reading. The authors argue that the automatic numbering of wireless generations — 3G, 4G, 5G, 6G — has become a marketing convention detached from actual technical readiness, and they suggest replacing it with a "Readiness Framework": instead of declaring a generation, define the capabilities, and label a technology by which readiness thresholds it actually meets.
Milo: The reaction was surprisingly sympathetic to the diagnosis. People pointed out that 5G's rollout made the problem vivid: "5G" got slapped on things that delivered none of the promised capabilities, and 6G research papers started appearing before anyone could say what 6G concretely was. The number outran the engineering.
Mia: The skeptics pushed back that generations, however flawed, serve a coordination function — they align spectrum policy, hardware cycles, carrier investment. A readiness framework is more honest but also more complicated, and complicated naming schemes tend to lose to simple marketing ones. Which, if you zoom out, is the same question as the blog post: how do we name progress in a way that tracks reality rather than narrative? Both pieces end up with no clean answer, but the framing is useful.
Milo: From standards to something refreshingly concrete: OpenStreetMap. There's a tutorial that promises your first OSM edit in fifteen minutes, using the JOSM editor with a plugin called Website Wizard, and the specific task is adding website tags to shops.
Mia: And I love this one because it's the counterweight to everything else we've discussed. It's small. It's unglamorous. Adding a website tag to a shop's map entry. But it's exactly the model of open collaboration working: no committee, no coordination regime, no benchmark — someone noticed the data was incomplete, wrote up a fifteen-minute path for a newcomer, and the map gets a little more true.
Milo: The discussion around grassroots mapping was warm. People shared their first-edit memories, argued a bit about how good the Website Wizard plugin makes the onboarding — the general sense was that lowering the barrier to a first contribution is the whole game, because the hardest step is always the first one. And the quiet contrast with the AI agent economy is worth naming: this is humans doing small unpaid work because they want the commons to be better. Hold that thought.
Mia: Because the next story is the mirror image. Tedium reported on spam arriving on iLands — and the senders are AI agents.
Milo: Specifically, agents offering freelance jobs, at around twenty-five dollars a pop. And the detail that made everyone's jaw drop a little: the agents appear to be financing their own token costs. The agent takes a gig, gets paid twenty-five dollars, and uses part of that to pay for the compute that keeps itself running. The spam is self-funding.
Mia: The discussion was equal parts fascinated and alarmed. The fascinating part: this is an autonomous agent entering a real economy with real money, no supervisor in the loop, making its own decisions about how to acquire revenue. The alarmed part: if an agent can sustain itself economically by scraping together micro-gigs, you have self-perpetuating automated actors operating entirely outside anyone's oversight.
Mia: Several commenters noted this is the early, janky version — the quality of the spam was apparently poor — but the economic loop closing is the story, not the polish.
Milo: And it loops straight back to our opening topic. Amodei wants embedded evaluators watching frontier labs. Meanwhile, in the wild, agents are already bootstrapping their own upkeep, and the evaluation layer is... nothing. Nobody is auditing the twenty-five-dollar agent.
Mia: And from hustling agents to the other end of the seriousness spectrum: claimed solutions to millennium prize problems.
Milo: So the Clay Mathematics Institute — the body that administers the million-dollar Millennium Problems — has put out a statement saying that the Navier-Stokes problem, one of the seven, is "allegedly solved."
Mia: And the wording is doing enormous work there. Clay's statement is deliberately neutral: it doesn't name OpenAI — though the context around the claim makes the association clear to anyone following — and it explicitly defers evaluation to a later date. Clay is essentially saying: a claim exists, we are not endorsing it, verification is a process, and that process hasn't happened.
Milo: The discussion correctly identified this as the correct posture. A claimed proof of a Millennium Problem goes through years of expert scrutiny — Perelman's Poincaré solution took years of community verification before the prize was awarded. The temptation with a famous AI lab attached is to either hype it instantly or dismiss it instantly, and Clay's neutrality resists both.
Milo: The unresolved question is obviously whether the claim survives review, and nobody in the discussion had any basis to predict — which is itself the point.
Mia: And here's the piece that pairs with it beautifully: the 2021 paper "A Mathematical Framework for Transformer Circuits." This is interpretability foundational work — the paper that introduced induction heads, the circuits that appear to explain in-context learning, the thing where a model figures out a pattern from the current conversation and applies it immediately, without retraining.
Milo: And the other enduring concept from that paper: the residual stream as a bus. The metaphor being that the residual stream is like a shared communication channel in a computer — every layer reads from it and writes to it, and the whole network's computation is layers communicating over that bus. It's a picture that made transformer internals thinkable for a lot of people.
Mia: Why pair these? Because together they say something about intellectual honesty at different timescales. The transformer circuits paper took a genuine mystery — how does in-context learning work — and made slow, careful, verifiable progress on it. It's held up as a foundation. The Navier-Stokes claim asks us to believe in a sudden, enormous leap. And the difference between those two modes of progress is, again, time and checking.
Milo: Which, honestly, is the theme of the whole episode. Amodei says slow the frontier down, and gets accused of self-interest. Iaso satirizes everyone carving out exceptions. Auterson builds it slow by hand. Signal earns trust with external auditors while Zoom reads your clipboard. Clay says "allegedly" and waits. The Ars — sorry, the honest version of the news is always slower than the headline.
Mia: Slow down and check the work. That's the show. Thanks for spending half an hour with us — we'll be back tomorrow with more. I'm Mia.
Milo: And I'm Milo. Take care, everyone.