
0813 | Grok 4.6, DeepSeek V4 Pro, Qwen 3.8, and Breaking the WAL
Show notes
This episode opens with a model-release arms race as xAI's Grok 4.6 lands a 61 on the Intelligence Index, quickly followed by DeepSeek V4 Pro 0813's Opus-class performance at a fraction of the price, and Qwen's massive 2.4-trillion-parameter open-weight release. The hosts then unpack Tailscale's hunt for a 16-year-old SQLite write-ahead log bug, Lovable's $400M Series C raise, a German criminal complaint over Meta's AI glasses, uBlock Origin's retreat from blocking Facebook ads, and mass vulnera
Timeline
- 00:00:00 Opening
- 00:00:46 Grok 4.6 lands with a 61 on the Intelligence Index
- 00:02:21 DeepSeek V4 Pro 0813: Opus-class at 20x less?
- 00:04:23 Qwen3.8-2.4T: a five-terabyte open-weight release
- 00:06:53 Tailscale's 16-year-old SQLite WAL bug, and the open-source fix
- 00:08:28 Lovable raises $400M — but is the AI builder already commoditized?
- 00:10:21 German criminal complaint targets Meta's AI glasses
- 00:12:06 uBlock Origin gives up on Facebook ads
- 00:14:07 Fake ClaudeBot: mass vulnerability scans spoof AI crawlers
- 00:16:18 What maths are LLMs good at? Tim Gowers weighs in
- 00:18:40 Does AI remove software engineering's middle class?
- 00:20:56 Pixel Watch 5's 30-hour battery reignites the charging debate
- 00:22:57 HTML over WebSockets: brilliant or rage bait?
- 00:25:08 License plate readers: general-purpose cameras needing warrants
- 00:27:13 Meta's monetization rewards controversial creators, report finds
Related links
- Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index - Bri Hacker News Campaign Feed
- Grok 4.6 - Bri Hacker News Campaign Feed
- DeepSeek V4 Pro 0813 - Bri Hacker News Campaign Feed
- Qwen3.8-2.4T - Bri Hacker News Campaign Feed
- Tailscale Traces Database Corruption to 16y/o SQLite WAL-Reset Bug - Bri Hacker News Campaign Feed
- Breaking the WAL - Bri Hacker News Campaign Feed
- Lovable raises $400M Series C - Bri Hacker News Campaign Feed
- German advocacy group lodges criminal complaint over Meta AI glasses - Bri Hacker News Campaign Feed
- uBlock Origin Is Giving Up the Fight to Keep Ads Off Facebook - Bri Hacker News Campaign Feed
- Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot - Bri Hacker News Campaign Feed
- What sort of maths are LLMs good at? - Bri Hacker News Campaign Feed
- AI is removing the middle class of software engineering? - Bri Hacker News Campaign Feed
- Pixel Watch 5 - Bri Hacker News Campaign Feed
- HTML over WebSockets: real-time SPAs with barely any JavaScript - Bri Hacker News Campaign Feed
- License plate reader searches should require a warrant - Bri Hacker News Campaign Feed
- Controversial creators are benefiting from monetization programs run by Meta - Bri Hacker News Campaign Feed
This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.
Transcript
Mia: Hey, welcome back to HackerNews Daily. I'm Mia, and as always I'm joined by my co-host Milo. Great to have you with us.
Milo: Thanks, Mia, glad to be here. We've got a packed episode today — big model news, a major funding round, plus some genuinely wild stories from the Hacker News front page.
Mia: Exactly. We'll start with Grok's new top score on a major AI benchmark, then dig into a rumored DeepSeek release, a huge new open-weights model, a brewing privacy fight, and plenty more.
Milo: And we'll round things out with a big gadget launch, some sharp commentary on where software engineering is headed, and a privacy story that hits close to home for a lot of drivers. Let's dive in.
Mia: So let's dig into the new model from xAI. Grok 4.6 just landed with a score of 61 on the Artificial Analysis Intelligence Index, a widely watched benchmark that ties together a lot of individual model tests into a single number. And the company is positioning this as a frontier-level release.
Milo: And the interesting part isn't just the benchmark score, it's the timing and the price talk in the community. One commenter pointed out that this dropped right after DeepSeek released its V4 Pro update, and someone else noted both teams seem to be circling a new release from Qwen. So there's clearly a race going on among these labs right now.
Mia: Right. On the value side, a lot of the discussion is about Cursor, the coding tool. Since Grok 4.5, Cursor has had what one commenter calls an incredible deal for frontier-level models. The argument is that their subscription now goes much further than what OpenAI or Anthropic offer, and even on their lower-tier plans you can burn through a lot of tokens on their first-party models, Grok and Composer, without really running out compared to the competition.
Milo: So the bench score is one headline, but the real story for developers is the cost per token attached to that frontier-level performance. A 61 on that index put Grok 4.6 right in the middle of the top tier conversation, and the community's immediate reaction was about how cheaply you can actually get access to that power through tools like Cursor.
Mia: Now let's talk about DeepSeek V4 Pro 0813, the update that commenters were pointing to as the reason for the timing of Grok's release. It's now listed on OpenRouter as the general availability version of DeepSeek V4 Pro, and the pricing is the headline. Listed at about 43 cents per million input tokens and 87 cents per million output tokens.
Milo: And one commenter put it in stark terms: it's competitive with Anthropic's Opus 4.8, weaker than models called Sol and Fable, but about twenty times cheaper. That's a per-token cost comparison deep enough to change which model you reach for by default.
Mia: The other part people keep comparing is the context window, which is huge here at over a million tokens, with up to 384,000 tokens of output per request. And OpenRouter notes the listed price isn't necessarily what you pay, since caching and discounts often bring the real price well below that. It also tracks throughput, latency, and uptime on the model, so you're not just trusting the card.
Milo: And on the quality side, one developer found Pro clearly better than the updated DeepSeek Flash for code reviews. Flash makes more initial mistakes, has to recheck things, and often produces about five times more output, claiming there's a bug or that code won't compile when those problems don't actually exist, then correcting itself later. So the cheaper, faster model turns out to cost more in developer attention.
Mia: Right. One commenter put the gap at around five percentage points, citing 87 percent versus 82 percent on whatever test they were looking at. So DeepSeek is basically selling a near-frontier result at commodity prices, and the tradeoff is doing the arithmetic on what those follow-up corrections actually cost you in time.
Mia: And that Qwen release everyone was timing their launches around just landed. Qwen3.8-2.4T-A95B is now on Hugging Face with open weights, and Qwen is making a point of this being the first time a Qwen-Max-class model has been opened up. The name tells the shape of it: 2.4 trillion total parameters, with about 95 billion active.
Milo: That's the key open-weight detail, because only a fraction of the weights are active per token, which is what keeps it usable. But the raw file size is the reality check. One commenter bluntly called it roughly a five terabyte model. There's a heavily compressed one-bit version that comes in around 400 gigabytes, and the full lossless version clocks in at nearly five terabytes.
Mia: And the communities are already debating what that means in practice. The one-bit version, at 95 billion active parameters, is being described as putting Opus 4.5-level performance into a machine a normal person could actually buy, with usable tokens per second. One commenter says the card claims the model sits somewhere between Opus 4.8 and a model called Fable 5, and that a machine holding seven terabytes of memory, including context and cache, is within reach of medium-size companies.
Milo: But there are tradeoffs in the open release itself. One commenter noted the vision capability has been removed, and the context is capped at 250,000 tokens, while the official commercial Qwen3.8-Max version built on this same model adds vision input and a larger default context with built-in tools. There's already speculation someone will bolt a vision tower onto the open version to restore images at lower performance. So the open weights get you the core, but at the edges they're a trimmed-down version of the official model.
Mia: And there's honest skepticism too. One commenter said the model card looks almost too good to be true. So the open-weight crowd is weighing genuine frontier-level capability against the hardware you'd need to run five terabytes of model, and how much was stripped out to make the open release possible.
Mia: Finally, there's a deeper story hiding in the infrastructure behind all these releases, and it's about a forgotten bug in a widely used database. Tailscale published about tracing database corruption to a sixteen-year-old SQLite write-ahead log reset bug, and around the same time the testing company Antithesis broke a similar write-ahead log bug in their own product.
Milo: The lesson the community kept coming back to is how a tiny, ancient bug can quietly corrupt data for years. And more interestingly, how companies are choosing to handle it. Simon Willison highlighted that Tailscale funded work on an open-source shim for SQLite's virtual file system, and that shim helped isolate the race condition almost immediately and should help track down similar bugs going forward.
Mia: So a company hit a real-world corruption bug, and instead of just patching their own instance, it paid for a new, very specific debugging tool that's now available to the wider ecosystem. The Antithesis side notes they wanted to publish their findings as soon as possible because people were talking about this bug that same day, with a follow-up on the way.
Milo: I think that's the takeaway to close on. Sixteen years underground in one of the most widely deployed databases on the internet, and it took a specific fundraising effort just to see it clearly. These leaks show that the reliability of the tools everyone builds on is a constant, ongoing commitment, not something you set up once and forget.
Mia: Freelance software tool Lovable just landed a four hundred million dollar Series C round, and it pushes the company's valuation to thirteen point three billion dollars. Menlo Ventures led the round, co-led by a European fund managed by EQT, with a long list of new backers coming in and several existing investors returning. And the company announced it in a blog post in mid-August.
Milo: So what does the growth story look like behind a valuation like that?
Mia: Lovable launched back in November of 2024, and the company says since then people have built more than sixty million projects on it. Their apps get over nine hundred million visits every month, which is a striking number. And business adoption has moved fast too — they say Fortune 500 employee reach went from about half of those companies in the first year to nearly two-thirds less than a year later.
Milo: And the featured capabilities show this is going well beyond a simple app builder — things like payment functionality, tools for search engines and AI search, integrations with the big office suites and payment platforms, plus automatic security scanning and governance controls. There's even a certification and a security and trust center for published apps.
Mia: The user base is fairly serious about this too. Lovable's own survey data says nearly eight in ten builders are creating a business or side project they hope to monetize, and more than a third of those already earn revenue from it.
Milo: The Hacker News reaction was mixed, honestly. Some people are still skeptical. One commenter asked whether anyone even still uses Lovable now that people have moved to tools like Codex and Claude Code, and whether it's still the easiest way to get a site or app going. Another commenter said they use it themselves to work on a side project.
Mia: German digital rights group HateAid has filed a criminal complaint over Meta's AI glasses. According to Reuters, the complaint targets Meta, parts of eyewear maker EssilorLuxottica including Ray-Ban, and several big German retailers that sell the devices. The heart of it is the Ray-Ban Meta Wayfarer specifically, which HateAid says breaks German privacy law because it's designed to film people without them noticing.
Milo: The managing director put it starkly — that there's no place to escape from smart glasses, and you have to expect at any moment to be filmed and then exposed on the internet.
Mia: The legal basis is a federal data protection law that prohibits selling communication devices designed to record people secretly. The complaint was filed with the Frankfurt-based digital crime prosecution unit, which confirmed it received it and said it would do a routine preliminary review to see whether there's grounds for a deeper probe.
Milo: And what's the reaction from the retailers caught up in this?
Mia: MediaMarkt and Saturn's parent group said it's taking the complaint very seriously and noted that suppliers have contractual obligations that all goods be law-compliant. Mister Spex said it hadn't been officially notified but takes privacy protection seriously.
Milo: On Hacker News, commenters got into the legal and practical arguments. One explained that German law requires recording devices to be visibly recognizable as such, and glasses aren't a widely known and accepted camera device — the argument is basically that this model looks too much like normal sunglasses for people to tell it apart. And another commenter said that from a non-lawyer reading of the law, this looks like it could be an easy case to win.
Mia: A post on Hacker News pointing to a writeup from August claims uBlock Origin is giving up the fight to keep ads off Facebook, alongside a similar headline from a tech news site about how hard Facebook ads are becoming to block. The bulk of the discussion, though, spun off into a bigger question — whether AI will eventually end browser advertising altogether.
Milo: That's a big claim. What are people actually arguing?
Mia: One commenter made the case that once AI models are fast enough at processing images in real time, someone like Apple can ship a browser rendering engine that simply masks all ads. It would take a few years, they figure, but the browser ad business is in what they called its terminal phase. Another commenter called it the first and only use case for large language models they're genuinely excited about.
Milo: But there was pushback on that rosy picture too.
Mia: Definitely. One commenter called it naive, arguing the more likely scenario is AI being used to add or replace ads in ways that can't be removed anymore — and that running your own AI inside the browser could be disallowed in the name of a safe browsing experience. And several people doubted Apple in particular would lead that fight. One pointed out that Safari doesn't even come with a built-in ad blocker and that the move could collide with Apple's highly lucrative search placement deal with Google. Another noted that Apple sells ads itself now, and increasingly across its surfaces.
Milo: And there was a darker framing in the thread too — one commenter said it's amazing how people think Apple has a clean reputation because of what they called a thin layer of indirection in its part of the surveillance ecosystem.
Mia: And for what it's worth, at least one commenter said the whole discussion made them even more excited for the open-source Ladybird browser to be released.
Mia: There's a Hacker News story circulating about someone running mass vulnerability scans by spoofing AI bots like ClaudeBot. It points to an analysis tracking bot traffic across more than five thousand websites, which reports that bot traffic now makes up thirty-five percent of all visits — actually down two percent from the previous ninety days.
Milo: So what's actually changed?
Mia: AI-related bot traffic is up eleven percent and now sits at twenty-eight percent of visits. Among the most common visitors are the standard search engine crawlers like Bing and Google, a couple of SEO tool crawlers, and then the AI agents — ChatGPT's user agent, ClaudeBot, and others. By type, AI search crawlers make up about twelve percent of activity and AI data scrapers about eleven percent.
Milo: But the core claim in the post is that some of these visits are fake — that someone's spoofing those AI bot names.
Mia: Right. The author says a lot of these visits use faked user agents, and that they fail IP verification or a check called Web Bot Auth, and they flagged a surge across many websites in the last week. One commenter backed that up, saying many of the listed user agents are commonly faked. Their advice: check the network that actually owns the IP, because blocking most cloud providers makes most faked bots disappear. Though they noted some run from residential and phone IPs using hijacked code, where what you think is a reader is really a multipurpose proxy.
Milo: And were there any ideas on where the scans are coming from?
Mia: That same commenter suggested the surge could be tied to a newly released vulnerability, which would mean looking at exactly which URLs the scans target. And there was useful caution from another commenter about country-level attribution — that fiber leaving a country can be tapped and packets can be injected with arbitrary origin IPs, so an internet provider can't always verify where a connection really comes from.
Mia: Timothy Gowers, writing on his blog just days after OpenAI announced it had solved ten major problems in mathematics and theoretical computer science, has been thinking about what sorts of maths large language models are actually good at. Two of those solved problems stand out to him. One is the first construction ever of a non-sofic group — a construction he calls, based on talks he's attended, one of the most important unsolved problems in group theory. The other is a long-standing problem in Ramsey theory about how fast a certain number grows when you keep adding more colours. He says it was a major open problem he didn't expect to see solved in his lifetime.
Milo: And he's honest about the limits of his analysis. He wrote the post in early August 2026 with the expectation that capabilities will shift quickly, so he expects it may mainly serve as a record. And notably, he doesn't claim a crisp classification of what these models are strong at — instead, he tries to rule out bad answers.
Mia: That's an important distinction. Here's the pattern he noticed: the models aren't just good at finding counterexamples, they also produce proofs of difficult statements. But almost all of the most famous problems they've solved — the two he mentioned, plus the Jacobian conjecture and the unit distance conjecture — are counterexamples. To actually argue that's a reliable pattern, he says, you'd need to decide when solving a problem counts as finding a counterexample, and then explain why these models would be especially suited to that kind of problem.
Milo: There's also a speed argument in there. He points out that if the models were better than all humans at every part of mathematics, their enormous speed advantage would have produced a much bigger flood of results. Instead, the breakthroughs are concentrated in this narrower category.
Mia: And his closing note, quoted in full by a commenter, is a useful way to think about when the picture changes. He says a good sign that the models have reached human level on a much wider class of problems will be when they start proving theorems using methods that are new and surprising — not just the flashy counterexamples.
Mia: A developer going by Florian Herrengt, posting on the eleventh of August, makes the argument that AI is removing the middle class of software engineering — by making projects with a weak engineering culture fail much faster. His scenario is vivid: a senior engineer comes back to seven open pull requests. The first one is a massive change — over twenty-four thousand lines added, almost four thousand removed — with an AI-generated description. And the team changed more on the project over a weekend than during a weeks-long holiday.
Milo: The core behavioural pattern is that someone prompts an agent for a few hours, opens a pull request, and because it looks functional, they repeat it — until nobody on the project actually understands it anymore. He gives the example of one author who couldn't even say where a piece of data came from, and sent a Claude conversation as the design decision.
Mia: And reversing course is genuinely hard, not just annoying. Adding a table to the database might take ten minutes, but removing one later requires a migration plan, keeping the service running for paying users, and carefully handling old data relationships. His point is that accumulating technical debt is fine, as long as it's knowingly a shortcut.
Milo: The commenters pushed back on several fronts. One said this is the opposite of failing faster — people who would never have cleared the early hurdles now go deep, falsely hoping to prompt their way out, into deeper and deeper confusion. Herrengt replied that speed leaves no chance to stop them before the damage compounds. Others questioned whether hours-long AI sessions are even realistic, or, when one commenter called the short-term gains a trap and pointed to an Oracle ban on AI-generated code, others noted that ban was about copyright, not quality.
Mia: Either way, the through-line is the burden it shifts: reviewers end up carrying far more of the load, and a project's failure stops being slow and visible, becoming fast and expensive.
Mia: Google has announced the Pixel Watch 5, and the pitch centres on what they're calling Gemini Intelligence and Google Health — proactive, hands-free help, low-latency AI features, GPS tracking they describe as their most accurate yet even in dense cities or on trails, and a new Health Guardian suite. That includes what they call industry-first breathing emergency detection, which can call for help if the wearer doesn't respond, plus monthly trend summaries for blood pressure, sleep, and metabolic health. Pre-orders opened on the twelfth of August, with general availability starting on the twentieth.
Milo: But the real discussion on Hacker News wasn't about the AI features at all — it was about the battery. One commenter called the reported thirty hours of battery life a deal breaker, pointing to a Garmin that lasts two weeks.
Mia: Others pushed back on that comparison. One said, based on switching from a Garmin to an Apple Watch, shorter battery life matters less than you'd expect — charging phone, watch, and headphones just becomes a nightly checklist. Another noted it depends entirely on how you use it, since sleep data like resting heart rate and heart rate variability are valuable early signals for overtraining. And one commenter made the point that, unlike a phone or headphones, you actually want the watch on at night to monitor your sleep.
Milo: But the skeptics were just as loud. One Pixel Watch 2 owner said their watch barely lasts a day, and described a nightly ritual of charging before bed so the alarm wouldn't die overnight — and said they'd never buy another. Another disliked charging their Apple Watch for the same sleep-tracking reason, wishing battery life were far better.
Mia: So the announcement lands with impressive health and AI features, but the deciding factor for a lot of people seems to be whether a watch that needs charging every night can really earn a permanent place on their wrist.
Mia: There's a good debate running on Hacker News about building real-time single-page apps with barely any JavaScript. The case against the standard approach — a JavaScript framework, a JSON API, and two codebases held together by contracts — is that it's familiar but hardly the only way. The alternative explored here is HTML over WebSockets: the server sends already-built HTML, and the client just places it where it belongs. All the rendering logic stays in the back end in a single language, with no contracts and no separate API.
Milo: There are three flavours of this idea. Over plain HTTP, request by request — that's the style of things like htmx. Over something called SSE — a one-way continuous channel. And over WebSockets, a permanent two-way channel, which is what Phoenix LiveView and Django LiveView use. The article credits Phoenix creator Chris McCord with presenting LiveView at a conference back in 2019, where he built a real-time Twitter clone in fifteen minutes without writing rendering JavaScript or using a front-end framework.
Mia: So JavaScript is still there on the client, but only to open the channel, place received HTML, and handle secondary things like animations and events. After the initial connection, the flow is simple: the client asks for a particular page, the server queries the database, renders the HTML with its template engine, sends the whole thing back, and the client drops it into place. The server can even push changes without the client asking.
Milo: Opinion splits hard on this. One commenter, half sarcastic, asks whether the post is satire or rage bait, saying it's just a normal website, a multi-page app with extra steps. But another defends the approach from real experience with Rails and Turbo, arguing that keeping everything on the server makes this kind of app genuinely simpler to build and maintain, especially when real-time updates are the point.
Mia: Former police analyst Andrew Wheeler argues that searching stored license plate reader data should require a warrant. He made the case in a post drawing on his work as an expert witness in the Virginia case Schmidt v. City of Norfolk, where residents challenged whether police searches of years of cached plate data counted as an illegal search. A judge ruled against them, writing that in Norfolk the answer is, not today.
Milo: And that phrase, not today, is exactly the point Wheeler wants people to focus on. He reads it as a signal that under current case law, the question is not whether automatic plate reader data will eventually need a warrant, but when. He points to earlier rulings like the cell-site location information case, Carpenter versus the United States, alongside decisions on geofence warrants and historical aerial drone footage, as stepping stones in that direction.
Mia: Wheeler is careful not to overstate the case against the cameras themselves. He calls them a good investment mainly because they're cheap enough to pay for themselves, under three thousand dollars per camera, while admitting the evidence that they actually reduce crime is, in his words, pretty meh. He draws a key line between an active flag, like a stolen car alerting police the moment it passes a camera, and historical searches that track where a particular plate has traveled over time.
Milo: On that historical side, he argues a warrant requirement wouldn't seriously slow down real investigations. What he objects to sharply is the current status quo, calling the practice of not retaining the data at all very bad, and describing existing abuse-prevention rules as laughable. His bottom line, states should mandate warrant procedures for these historical searches. And one commenter pushed back with a broader point, noting these aren't just academic traffic cameras, they're general-purpose internet-connected devices that could be reprogrammed for other uses.
Mia: An ABC News Verify investigation says Meta is directly paying several controversial Australian creators through its Facebook monetisation programs. The list includes white nationalist Hugo Lennon, who was moved on by Victoria Police after shouting racist abuse at Indian Prime Minister Narendra Modi in Melbourne, and who's been earning through Facebook's Content Monetization program since September of 2025. Also named is The Noticer, a far-right news site promoting white supremacist and neo-Nazi ideology, which has been earning since November, aside from an unexplained gap in the new year, and whose X account is currently suspended.
Milo: The investigation also flags March for Australia, registered in December, and Monica Smit, founder of the anti-vaccine group Reignite Democracy Australia, who joined the program in September and appears to have earned Facebook ad revenue going back to 2017. Her page carries vaccine misinformation and promotes radiation protection bracelets. None of the creators responded to requests for comment, and the scheme is invitation-only.
Mia: Meta's response was that it has clear policies for its monetisation tools, and that policing offensiveness isn't its role. But ABC News Verify says the content appears, at times, to directly violate Facebook's own monetisation policies. For scale, Facebook says it distributed nearly three billion US dollars to around sixteen million monetised accounts in 2025.
Milo: Now, some readers pushed back on the framing, arguing the headline overstates things because these creators aren't commissioned directly by Meta, they simply qualify for program payouts based on their content's performance. Even so, the finding stands at its core, that the platform's rules are letting figures with openly white supremacist, racist, and anti-vaccine content collect money through a program the company runs, at times even against its own stated policies. And that's the tension that keeps this on the agenda as the program expands.
Milo: And that's a wrap for today. We covered Grok 4.6 topping the charts on the Artificial Analysis Intelligence Index, plus that surprising OpenRouter listing pointing to a new DeepSeek V4 Pro, and a fresh 2.4-trillion-parameter open-weight release from Qwen.
Mia: We also dug into a massive $400 million funding round for Lovable at a $13.3 billion valuation, a criminal complaint filed in Germany against X over handling of radicalized content, and Google's Pixel Watch 5 announcement.
Milo: Plus, a spirited debate over whether AI is hollowing out the middle class of software engineering, and the pushback on uBlock Origin calling it quits in the fight to keep ads off Facebook.
Mia: Thanks for listening. Catch you next time, and stay curious.