
0903 | Leaving, Splitting, and Trusting Machines
Show notes
From Uber's African retreat to AI benchmarks that don't add up: this episode looks at what big tech is leaving, selling, and claiming — plus dark matter, mushroom-spotting LLMs, and 44 years of the Commodore 64.
Timeline
- 00:00:04 Opening
- 00:00:48 Big Tech reshaping itself: Uber retreats, Google survives
- 00:05:10 Can we trust AI answers and benchmark claims?
- 00:09:55 Same model, wildly different cost — and a fast-moving model race
- 00:13:59 AI in the field: poison mushrooms, aging memory, dark matter
- 00:16:56 Classrooms, credentials, and doing it yourself
- 00:20:56 Fast round: fixes, classics, and internet friction
- 00:27:56 Closing
Related links
- Reasons robotics is hard
- I wanna live an NPC life
- Fable 5.1 World Modeling
- Poisson Disk Sampling
- Qantas Airbus A380 engine failure in 2010 (2023)
- Google avoids a breakup of its ad tech business
- WebLLM: high-performance in-browser LLM inference engine
- Mushroom hunting with LLMs: what can go wrong?
- Can I opt out of my input or output data being used for training?
- A Note from LWN
- Commodore 64 released September 1, 1982
- Muse Spark 1.3
- Check if a file was made with Claude
- GrapheneOS says Pixel 11 has MTE support after all
- Introducing Muse Spark 1.3
- A third of Perplexity's citations don't contain the number they're cited for
- Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
- LLMs: Intelligence vs. Cost
- Exit the Cave
- Saving money on Google Photos with Immich: Your own personal photo storage
- Quasar 438B: Europe's Leading AI Model
- Gemini 3.8 Flash and 3.8 Flash Cyber
- Three sites made 215,128 “best software” pages for AI. Perplexity cites them
- Mamdani Bans AI in NYC Schools
- SteamdDB Joins Nexus Mods
- Aging brains blend memories together instead of just forgetting them
- Six curl CVEs after OpenAI and Anthropic came back with zero
- Biggest dark matter detector spots a single weird particle
This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.
Transcript
Mia: Welcome back to the show, everyone. I'm Mia.
Milo: And I'm Milo. Today's episode is a strange kind of mirror — big companies retreating and defending their empires, and the rest of us trying to figure out what to trust. Uber shrinking its African footprint, Google dodging a breakup, and then a whole block of stories about whether AI's answers, benchmarks, and price tags mean anything at all.
Mia: Right, and then it gets more human from there — classrooms banning AI, tools to prove what a model touched, and people rebuilding their own software instead. Let's start with the retreat. Uber is exiting Nigeria and Uganda immediately.
Milo: Immediately — that's the striking word. No wind-down, no transition period, just gone. And this follows earlier exits from Côte d'Ivoire and Tanzania. So after this move, Uber is down to four countries on the entire continent: Egypt, Ghana, Kenya, and South Africa.
Mia: What I found interesting in the discussion was how people framed this. It's not framed as a collapse — those four remaining markets are the biggest economies, the places where ride-hailing actually has density. The argument was that the exits are pruning: markets where the unit economics never worked, where driver supply or payment infrastructure or regulation made it too expensive to keep pushing.
Milo: And the counterpoint was about what "immediately" means for the people there. Riders who built their lives around the app, drivers whose income depends on it — they get no notice. That's the human cost of a corporate map being redrawn overnight. Some commenters were pretty blunt that "we're optimizing our footprint" reads very differently if you're the footprint.
Mia: It also raises the unresolved question of whether this is retrenchment before a bigger strategic shift, or just the end of the experiment. Nobody in the discussion could say what Uber does next in Africa — whether the four remaining markets get more investment or are next in line eventually. That's genuinely unknown.
Milo: Which makes a nice contrast with our next story, because Google went the opposite direction — not shrinking, but successfully defending itself. Google won its ad tech antitrust case and avoids being broken up.
Mia: Yeah, and this one generated real disagreement. The judge, Brinkema — who, for context, was a Clinton appointee, so not exactly a natural ally for Google — ruled in Google's favor. And the immediate reaction from a lot of commenters was that the ruling felt lenient.
Milo: What did "lenient" mean in that conversation? I want to be careful here, because the sources don't give us the full legal reasoning.
Mia: Fair. What we can say is that people questioned the rigor of the decision — the sense that a company accused of tying its massive ad business together got to walk away without structural remedies. Some argued that after years of litigation, ending with no breakup means the underlying incentives haven't changed. Google still owns the auction, the exchange, and the ad network side in some combination, and the fear is the behavior continues under new compliance language.
Milo: And the other side — and this is worth saying because it wasn't unanimous — would be that remedies are supposed to fix harm to competition, not punish companies, and that forcing a divestiture of ad tech is genuinely hard without breaking products that publishers and advertisers actually use. I don't think the discussion settled that. It's the same argument that's been running for years.
Mia: The thing that connects it to Uber, though, is what you said in the setup: both are companies actively managing their footprint. Uber does it by leaving. Google does it by winning in court. Two totally different strategies for the same problem — the world you built doesn't fit anymore, so you reshape or defend.
Milo: And there was a third, smaller data point in that same space: SteamDB was acquired by the parent company of Nexus Mods. And the immediate worry in the comments was paywalling — that a community resource people rely on for free data might get monetized behind a subscription.
Mia: Right, and there was a second, more philosophical objection: some people questioned whether the data is even that valuable. Which is an interesting debate in itself — is SteamDB valuable because of the data, or because of years of trust and tooling around it? Either way, the community is nervous about what ownership change means for something they use daily.
Milo: OK, so from companies reshaping themselves, let's move to the thing underneath a lot of this: can we actually trust what AI systems tell us? And there's a study that landed hard on this. Haus Research audited Perplexity's citations.
Mia: The numbers are the story here. Out of 1,826 citations, 34.7 percent of the linked pages either wouldn't load or didn't contain the number Perplexity had attributed to them. And when you measure it by claims rather than pages, 14.4 percent of claims failed.
Milo: Let's sit with that for a second, because those two numbers tell different stories. A third of dead or mismatched pages means the browsing and citation layer is shaky. But the 14.4 percent claims failure rate suggests the model sometimes gets to the right answer even when its receipt is broken — or, more worrying, that some of the "right" answers are right by coincidence.
Mia: And the discussion dug into exactly that. One view: this is a fundamental trust problem. If you can't verify the citation, the whole value proposition of a "search engine that shows its work" collapses. You're not getting sourced answers, you're getting confident prose with decorative links.
Milo: The counterargument people made: every search product has some rate of bad results, and the question is whether 14 percent is catastrophically worse or just a visible version of what Google's been quietly doing for years. And there were firsthand accounts from people who said they use these tools daily and mostly catch errors when they matter — the danger being the times they don't catch them.
Mia: The unresolved part is methodology. Whose fault is a dead link — the model, the crawler, the source site that changed? And does "the page doesn't contain the figure" mean the model hallucinated, or that it read a page that's since been edited? The audit doesn't fully disentangle that, and commenters knew it.
Milo: Which is a perfect bridge to the benchmark side of this same distrust. Because Meta shipped Muse Spark 1.3, aimed at agent and coding tasks, with a "max reasoning" mode — and they claim parity with GPT 5.6 Sol and Opus 5 across multiple benchmarks.
Mia: And after a story about a third of citations not checking out, you'll forgive the room for being a little skeptical of vendor benchmark claims. The recurring question was: parity on which benchmarks, measured how, and does it survive real workloads? The honest answer from the discussion was — unknown. That's the thing nobody could resolve: how these claims hold up outside the vendor's own tests.
Milo: Though there was at least one independent data point. Simonw actually ran it, specifically on SVG generation, and found the quality better than version 1.2. That's a narrow test, but it's the kind of ground truth people were asking for — one person, one task, hands on.
Mia: And the pricing gives you a sense of where this sits: a million-token context window, a dollar twenty-five per million input tokens, four twenty-five per million output. So competitively priced, aimed squarely at the coding and agent market.
Milo: Which loops us into OpenTeams' criticism of Artificial Analysis' intelligence-versus-cost charts — because measuring model value is its own minefield. Their complaint was that the charts use a logarithmic axis, which visually flattens huge price differences. On a log scale, something ten times more expensive can look like a modest step up.
Mia: Yes, and I thought their second point was the sharper one: local models shouldn't be priced on API dollars at all. If you run a model on your own hardware, the real cost is electricity and hardware amortization — so putting it on the same dollar axis as API models is comparing fundamentally different things. The chart implies a common currency that doesn't exist.
Milo: So we have three layers of measurement skepticism now: citations that fail, benchmark claims we can't verify, and charts that may visually mislead. It's a coherent theme — the instruments themselves need auditing.
Mia: And that sets up our next story perfectly, because it's about an experiment that showed the instrument matters enormously. FrontierHarness ran the same model — Kimi K3 — through nine different agent harnesses, and the cost per task varied seventeen-fold. From a dollar five to eighteen dollars and thirty-four cents.
Milo: Seventeen times. Same model, same task. That's not a model difference, that's an infrastructure difference. One harness loops efficiently, another burns tokens retrying, over-prompting, re-reading context.
Mia: The takeaway people drew was almost heretical given how we usually talk: the harness matters as much as the model. All the discourse is "which model is best," and this says the scaffolding around the model can swing your bill by an order of magnitude. If you're deploying agents, your engineering effort might be better spent on the harness than on the model swap.
Milo: The open question there is why. Is it retry logic? Context management? Prompt design? The study shows the spread but the discussion was still working through which harness behaviors drive it — because if you know that, you can fix it.
Mia: Meanwhile the model race itself is accelerating in a way that's hard to even track. Google launched Gemini 3.8 Flash — 75 cents per million input tokens, three seventy-five output, with gains in coding and agent tasks — plus a Cyber variant aimed at cybersecurity. And here's the part that made people do a double take: that's the third Flash release in six weeks.
Milo: Six weeks, three releases. Commenters were asking what that cadence even means — is that responsiveness, or churn? If versions ship every fortnight, benchmark comparisons go stale before you finish reading them. There's no stable baseline to evaluate against.
Mia: And on the other end of the market, there's Quasar 438B from Multiverse Computing — the best European model, scoring 43 on the Artificial Analysis index, 500 tokens in 15.3 seconds, English and Spanish. Which is notable as a regional-competitiveness story.
Milo: Though — tying back to OpenTeams — the 43 on that index comes from the very measurement framework they're criticizing. So even the "Europe's best model" headline inherits the log-axis and pricing questions. It all connects.
Mia: And then there's a deeper, almost philosophical wrinkle: McCoy and colleagues found that neural network vectors — including in LLMs — can be approximated by closed-form symbolic structures without changing the model's behavior.
Milo: Yeah, that one had people genuinely excited and unsettled at the same time. The neural-versus-symbolic debate has been running for decades — "are these things just fancy statistics or do they do something else?" — and this result says, at least as an approximation, the boundary is blurrier than the framing suggests. If symbolic structures can stand in for the vectors with no behavior change, what does that tell us about what the network is actually doing?
Mia: The caveat people raised: "without changing behavior" is not the same as "without changing anything." Interpretability, robustness, the ability to inspect and verify — those could differ enormously between the vector version and the symbolic approximation. But as a research direction, it suggests we might be able to translate these systems into something more legible. Nobody knows yet how far that goes.
Milo: From abstract legibility to very concrete stakes. Let's talk about AI out in the real world — and specifically, three cases where the gap between confidence and correctness matters a lot.
Mia: The sharpest one first: researchers used the FungiTastic dataset — 340,000 observations across 2,800 species — to test GPT 5.6 Sol and three other LLMs on identifying poisonous mushrooms. The error rate on dangerous misidentifications was high enough that the clear conclusion was: never rely on these models for that.
Milo: And I want to underscore that, because mushroom ID is exactly the kind of task where a model's confident tone is dangerous. A wrong answer about code wastes an afternoon; a wrong answer about a mushroom can kill you. The discussion was basically unanimous on this one — the failure mode isn't that models are mediocre, it's that they're mediocre with perfect confidence.
Mia: Second case, more subtle: aging and memory. New research found that older brains don't simply forget — the hippocampus over-generalizes, extracting memory too broadly, so memories get confused across categories. You don't lose the memory; you misfile it.
Milo: The important caveat, and commenters were quick to note it: the sample was only 61 people. That's small. Interesting direction, genuinely preliminary. The suggestion that "forgetting" in aging is partly a categorization failure rather than a storage failure is a meaningful reframe — but it needs replication.
Mia: Third: the LZ dark matter detector — the largest ever built — recorded a single anomalous event. One. And the response from the scientists and the discussion alike was admirably sober: far too early to call it a discovery, more data is being collected.
Milo: I love that this is in the same episode as the mushroom study, because it's the same question — what do you do with a signal? — handled at opposite ends of the rigor spectrum. LZ says "one event, we wait." The mushroom result says "high error rate, here's your warning." Both are honest, but only one of them involves possibly eating a death cap.
Mia: And there's an interesting parallel with the memory study too — single observations, small samples, and the discipline of not over-extracting from them. Which is literally what the aging hippocampus fails to do. Over-generalizing from thin data is the failure mode everywhere today.
Milo: OK, next block: control. Who's in charge — schools, tools, or you. And the flashpoint is New York, where Mayor Mamdani has banned AI in elementary and middle school classrooms.
Mia: The supporters' core argument, as it played out: students need to learn to think independently first. The worry is that if you hand a ten-year-old an answer machine before they've ever struggled through a problem on their own, you never build the muscle. It's the calculator argument, but higher stakes because these systems don't just compute — they write, reason, and replace the actual thinking exercise.
Milo: The pushback in the discussion was that bans are blunt. The technology isn't going away, and a student who meets AI for the first time at fourteen meets it without any guidance. Several people argued the real question isn't "ban or allow" but "what skills do you teach alongside it" — and that a flat prohibition avoids that harder conversation. I don't think there was resolution; this one's genuinely contested.
Mia: What's interesting is it's the same distrust theme from earlier, applied to children. We spent the first half of the show on "can we verify what AI says" — and the ban is essentially saying: not yet, not for them, not before they can verify for themselves.
Milo: Now, on the transparency side, Anthropic launched a C2PA content-credential detection tool — you can check whether a file was processed by Claude. And the crucial asterisk: no credential does not mean the file wasn't AI-generated.
Mia: That asterisk is doing a lot of work. It means the tool can give you a positive signal — "yes, Claude touched this" — but absence of evidence is not evidence of absence. Anything that went through a system that doesn't write credentials, or that stripped them, shows up as clean. So it's a partial window, and the discussion appreciated that Anthropic said so plainly rather than overselling it.
Milo: It's the citation problem again, in a different costume. Perplexity's links can be dead or mismatched; Anthropic's credentials can be missing. In both cases, the verification layer is real but incomplete, and the honest position is knowing exactly what it does and doesn't prove.
Mia: Then the DIY corner. WebLLM — a browser-side WebGPU inference engine with a lot of stars — but commenters were saying the project is effectively dead, and pointing people to llama.cpp's WebGPU support or ONNX as the live alternatives.
Milo: It's a small story with a familiar shape: popular open-source project, momentum fades, community forks toward maintained options. The lesson people drew wasn't "don't self-host," it was "check who's still maintaining before you build on something."
Mia: And if you want the positive version of that story: Immich. Self-hosted Google Photos replacement — Docker deployment, automatic phone backup, multi-user sync — and the feedback was genuinely good. People reported the experience holding up to daily use, which for a self-hosted photo solution is the real test.
Milo: And I think Immich resonates in this episode for a reason. We've spent the whole show on trust — trust in citations, benchmarks, price charts, credentials. Self-hosting is the terminus of that line: the only verification layer you fully control is the one you run yourself.
Mia: Let's do the fast round now — a bunch of stories where the discussion itself is the story. First: AISLE found six CVEs in curl. All low severity, all fixed in curl 8.22.0. The wrinkle is that Codex and Mythos — AI agents, doing security reviews — had previously reported zero.
Milo: And that sparked a real debate. One reading: AI review found nothing, human researchers found six, so AI security analysis is still far from sufficient. The other reading: the findings are all low severity, so maybe AI correctly triaged that nothing high-impact existed, and six low-severity bugs is normal background noise for a codebase like curl. Both readings fit the facts, and people argued both.
Mia: It also touches the memory-and-overgeneralization theme, honestly — drawing a big conclusion from six low-severity findings is exactly the kind of inference the episode keeps warning about.
Milo: Next, a lovely one: Bridson's 2007 paper on Poisson disk sampling — it fits on one page — is approaching a thousand citations. And the author himself is still at it, proposing a "parent-point cone culling" optimization that reduces iterations.
Mia: One page, twenty years of relevance. In an episode full of bloated benchmarks and seventeen-fold cost spreads, there's something deeply satisfying about an algorithm so clean it fits on a sheet of paper and still earns improvements from its own author.
Milo: And a birthday: the Commodore 64 launched September 1st, 1982. 64 kilobytes of memory, under 600 dollars, roughly 12.3 million units sold — the best-selling computer in history.
Mia: 64K, and it outsold everything. There was real affection in that thread — and, tying to Immich, a running comparison about how much computing it actually takes to back up your photos versus what a 1982 machine could do with 64 kilobytes.
Milo: Then some internet-economy friction. LWN is raising subscriptions about 20 percent on September 15th, 2026 — the first increase since 2022 — and the stated reason is that subscriber growth has stalled. So the existing base funds the increase.
Mia: Which prompted a genuine discussion about sustainable journalism on the open web: is a 20 percent increase every four years the healthy model, or a symptom of a ceiling? People who pay for LWN mostly seemed inclined to keep paying, but the structural question — can niche technical journalism grow its audience or only deepen its value to the faithful — went unanswered.
Milo: And Mistral did something that upset its users: they turned on training on prompts by default at the Team tier, and in the process removed the central admin toggle that let organizations control it. Docs got corrected after complaints.
Mia: The part that bothered commenters most wasn't the default itself — it was taking away the control. If admins can't centrally enforce a privacy setting, every team member's behavior becomes a policy decision. It's a governance failure dressed as a settings change, and it's exactly the kind of thing enterprises notice.
Milo: Now something more fun: Fable 5.1's agent swarm generated a full 3D world of San Francisco's Union Square — 453 buildings, 129 storefronts, 220 pedestrians — for about 8 million tokens and 33 dollars.
Mia: Thirty-three dollars for a walkable 3D city block. People compared that to the cost of manually modeling even a single one of those buildings. Now, the fair caveat — and it was raised — is that generated isn't the same as game-ready or accurate; it's a world-shaped output, not a finished asset. But the cost curve is the headline.
Milo: Two culture debates to close the round. First: a blogger argued the "NPC lifestyle" — tuning out, not caring about politics or risk — is the ultimate peace of mind. Hacker News pushed back hard.
Mia: The core rebuttal: reality has consequences. Ignoring politics and risk doesn't make them disappear, it just means they hit you unprepared — and calling that enlightenment is, in the commenters' framing, learned helplessness wearing a philosophy costume. The blogger's side would say the constant outrage cycle is its own harm; the reply is that there's a difference between opting out of noise and opting out of agency.
Milo: Second: an essay called "Exit the Cave" argues that failing publicly beats practicing privately — ship the bad work, learn in the open. And the debate was about where that applies. For creative work and skills with feedback loops, lots of people agreed — public failure compounds. But others pointed out boundaries: professions where failure harms other people, or contexts where early public work follows you permanently. "Failing in public" means something different for a comedian and a surgeon.
Mia: Two more quick ones. GrapheneOS notes the Pixel 11 retains minimal hardware MTE support, but degraded performance — the cache acceleration was removed — led to MTE being disabled in firmware. So a security feature people care about is technically present and practically switched off.
Milo: And the ugly mirror of our trust theme: three sites created in late 2023 generated 215,128 pages of "best software" content designed to be cited by AI — and 59.8 percent of Perplexity's citations come from outside the top 100,000 Tranco sites.
Mia: Which reframes the Haus Research audit from earlier. It's not just that the model cites dead pages — there's an entire adversarial industry building pages specifically to be cited. The citation layer isn't just unreliable; it's under attack. Those two stories together say the "show your work" feature of AI search may be its most exploitable surface.
Milo: So let's try to pull the thread through the whole episode. Companies retreat and defend. Models get faster, cheaper, more impressive on paper. And underneath all of it runs one question: what's your verification layer?
Mia: For citations, it's dead links and adversarial content farms. For benchmarks, it's vendor tests and log-scale charts. For security, it's AI reviews that found zero where humans found six. For your photos, apparently, it's a Docker container on your own machine.
Milo: And for a ten-year-old in New York, the answer is: learn to think without it first. Every story today is somebody deciding how much to trust, and building — or refusing to build — the apparatus for checking.
Mia: That's the show. Thanks for listening — we'll be back tomorrow with whatever the internet argues about next.
Milo: See you then.