0828 | The flash-tier open-weight race: GLM-5.3-Flash vs Qwen3.8-Flash-Next

||Download

Show notes

0828 | The flash-tier open-weight race: GLM-5.3-Flash vs Qwen3.8-Flash-Next

Timeline

  • 00:00:00 Opening
  • 00:00:45 The flash-tier open-weight race: GLM-5.3-Flash vs Qwen3.8-Flash-Next
  • 00:02:17 Speech stacks: Speko's one-API offering and Gemini 3.5 Transcribe
  • 00:03:06 Taming production agents: Traccia's control plane vs IQ Routing's trajectory-aware gateway
  • 00:03:56 Lenz: a fact-checking API with an audit trail instead of a single model's guess
  • 00:05:21 GitNexus: a knowledge graph kernel so coding agents stop guessing
  • 00:07:14 Ojin: interruptible live-voice agents with a face and a voice
  • 00:08:53 Savvy: a meeting copilot grounded in your documents, not the open web
  • 00:10:42 Cobalt: turning Kobo e-readers into an open app platform
  • 00:12:42 Sendra: Figma frames to inbox-ready email, checked against real devices
  • 00:14:05 SpacebarX: a keyboard-first, local-first outliner for notes, tasks, and writing
  • 00:15:53 Yomi: a reading app where kids read aloud and feed a cat
  • 00:17:33 The Million Sad Ducks: name a duck, make it happy, get a certificate

Related links

This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.

Transcript

Mia: Welcome to Product Hunt Daily on Bri Radio. I'm Mia, and alongside me is Milo. Today we're diving into a lineup that spans a natively multimodal AI model, a fact-checking tool built to keep AI-generated claims honest, an open-source knowledge graph kernel for coding agents, and a few others worth catching.

Milo: That's right. We've also got a macOS meeting assistant, an open-source platform letting Kobo e-readers run apps, a Figma plugin that turns email designs into responsive HTML, plus a keyboard-first outliner and a reading app for kids. And we'll finish with a world of one million sad ducks.

Mia: A lot of ground to cover. Let's get into it.

Mia: Z.ai has positioned GLM-5.3-Flash as the first natively multimodal model in its GLM-5 series. The launch describes a 320-billion-parameter model with 18 billion active parameters, and claims it outperforms GLM-5.2 across benchmarks at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic tests.

Milo: That open-weights story resolves an earlier mystery. Community commenters identified the model as the previously anonymous Ox Alpha, and one person said people were already running it on OpenCode and OpenRouter before the name appeared. They described the public version as 320B-A18B with native multimodality, a one-million-token context, and MIT weights, with sparse plus linear attention keeping that long window cheaper to serve.

Mia: One commenter put the API pricing at fifteen cents per million input tokens and fifty cents per million output tokens, with the first two weeks at fifty percent off. Another said they had run it through Kilo Code for a few days, called it fast, and said it held up on long agentic runs while finishing tasks quickly.

Milo: One commenter did ask when GLM models would finally read images sent to them. The launch materials call this model natively multimodal, so that question stays unresolved.

Mia: Speko bills itself as an OpenRouter for voice, meaning a single API covering speech-to-text, language models, and text-to-speech, with public benchmarks sitting beside runtime availability. The founder, Bek, says the motivation came from building voice AI apps before this, where his team had to support more than ten languages every week.

Milo: That pain point also shows up in Gemini 3.5 Transcribe, which its maker is positioning as Google's most precise speech-to-text model yet, built for real-time transcription that handles the way people actually talk. It manages self-corrections and removes filler words while you're speaking naturally rather than forcing you to type every thought.

Milo: Traccia is a vendor-neutral AI agent control plane for teams running autonomous agents in production. It observes agent behavior, evaluates performance, and governs actions through policies and runtime controls, keeping an auditable trail of what happened, built on an open developer-first SDK and OpenTelemetry so it works across models, frameworks, and existing observability stacks.

Mia: That framing lines up with IQ Routing, a drop-in gateway that treats an agent run as a trajectory rather than a stream of independent calls. It classifies each request, serves from cache, and routes every step to the cheapest model that clears the quality bar, so a cheap model handles boilerplate while the strongest model takes on what matters.

Mia: Lenz is an AI fact-checking API, aimed at products and teams that can't afford to ship a wrong AI-generated claim. Its pipeline extracts verifiable claims from text, searches independent sources, runs multi-model debate, and routes evidence through a review panel, then returns a scored verdict with sources, citations, reasoning, confidence, and a full audit trail.

Milo: A Product Hunt exclusive offers one month of Developer access free, with no credit card required, listed at a ninety-nine dollar value. That includes a thousand extracts per day, five thousand assessments a month, five hundred verifications, and a thousand asks per month.

Mia: To support the premise, the team gave a thousand real claims from Lenz to five frontier models, all with identical prompts, web search, and thinking enabled. All five models agreed on only thirty-seven percent of the claims, and on twenty-three percent their verdicts were two or more steps apart on a five-point scale. Confidence proved a weak signal, since three-quarters of answers were self-rated at nine or ten out of ten. The research paper, data, and prompts are published at lenz.io forward slash research, and one community member reported text input worked seamlessly.

Milo: GitNexus, from Akon Labs, is an open-source knowledge graph kernel for coding agents. Maker Subham argues agents spend most of their budget finding context rather than writing code, and that current tooling is a proxy, since a file tree mirrors the disk rather than the system, and embeddings return who-calls-this-function as a ranked guess instead of an exact answer.

Mia: It resolves a codebase into one deterministic graph, with nodes for symbols, files, functions, services, and repos, and edges for calls, imports, implements, and deploys, spanning GitHub, GitLab, Azure DevOps, and self-hosted enterprise repositories. It exposes seven MCP tools and works with Claude Code, Cursor, Codex, OpenCode, Windsurf, and Antigravity, and the company says there's no lock-in, with managed, self-hosted, and air-gapped deployment options. Akon Labs is backed by Y Combinator, and reports more than forty-five thousand GitHub stars and over a million npm downloads.

Milo: The company's benchmark says agent runs come out fifty-one percent cheaper with GitNexus connected. On DeepSWE, a benchmark Akon Labs itself notes is single-issue and can't exercise multi-repo graphs or impact analysis, GitNexus solved sixty-eight point four percent of issues versus thirty-seven percent for a bare model and fifty-four percent for Graphify, at eighty-eight cents per solved task versus a dollar seventy-nine, with ten point six percent fewer steps. One commenter said most context tools are really just fancier grep, and that cross-repo calls are exactly where the graph approach changes the answer.

Mia: The Ojin launch on Product Hunt is an AI agent that comes with a real face and real voice, built specifically for live conversation rather than turn-based messaging, and pitched at fast-moving teams, makers, and enterprises. Founder Mio says setup only needs one still photo, a plain-language persona, and a voice, with no capture session or script involved.

Milo: The hard engineering part there is turn-taking — telling when someone has actually finished a sentence rather than just paused to think. On the face models, the listing advertises two behind one API: Oris Portrait, a fast, scalable model with under two hundred milliseconds of latency, and Oris Presence, described as the most lifelike face model available. Both stream over WebSocket and drop into Pipecat or LiveKit Agents.

Mia: Pricing starts at five cents per minute with ten dollars in free credits, and the end-to-end Human Agents bundle includes speech-to-text, an LLM, voice, and face in a browser-native experience. In the comments, one tester said Sofia handled interruptions and random interjections smoothly, though it was still obviously an AI. Someone else argued endpointing latency, not the two-model API, is the real differentiator, and there was a commenter asking how Ojin differs from LemonSlice.

Milo: Another commenter pushed back on the per-minute pricing itself, saying it's less useful than cost per completed call outcome — because a four-minute call that books a meeting can be cheaper than a sixty-second call that goes nowhere.

Mia: Savvy, new on Product Hunt, is a macOS meeting assistant built from the maker's complaint that existing AI meeting tools join calls and mail summaries afterward — they help the archive, not the conversation. Instead of drawing on the open web, it grounds itself in your own documents: point it at a folder and it builds a versioned brief per client.

Milo: During a call it stays quiet until it speaks up for exactly three reasons — the other side asks a question the brief answers, someone crosses a red line you set, or you press the Advice button. Every card cites its source document. It needs Apple Silicon and macOS 13 or later, and it's open source under MIT.

Mia: On privacy, the launch post says documents, indexes, and transcripts stay on the Mac, with only excerpts and the audio stream leaving it. The README goes further: at recommendation time, it sends selected excerpts, the whole brief, and recent relevant transcript turns. One important caveat — live transcription is not local. Meeting audio streams to Deepgram or AssemblyAI, and the current build has no offline transcription mode. Transcript and audio files are deleted after thirty days.

Milo: Savvy sets Deepgram's mip opt out to true so audio isn't used for their model training there, though users have to check AssemblyAI's terms for an equivalent. Provider credentials live in macOS Keychain, and recommendations run through OpenAI's Codex CLI or Anthropic's Claude Code under your own account. Removing a client deletes derived data but never the source folder.

Mia: Cobalt, according to the maker, is an open-source application platform that lets Kobo e-readers run apps — a launcher, a signed app store, a Rust SDK, and a runtime that keeps every app in its own unprivileged process. Installation happens once over USB; after that, apps install, update, and remove on the device over Wi-Fi, with signatures verified before launch, and a reboot returns to the stock Kobo reader.

Milo: The maker built it partly because long Claude sessions forced laptop use and strained their eyesight — they wanted e-ink for reading and for approving coding-agent requests. Their store includes an RSS reader, Audiobook Studio, which researches, writes, narrates, and plays original audiobooks using deep research and the ElevenLabs API, plus an arXiv browser, a Gutenberg browser, offline games, a Claude Code and Codex controller, and a terminal.

Mia: Sidekick handles approving or denying coding-agent requests away from the keyboard. On the technical side, the SDK models an app as one Rust file, with resources like network, storage, audio, frontlight, and Wi-Fi capability-gated. A commenter said managing Claude Code and Codex from a one-hundred-dollar device with a month-long battery on the go has been a real help.

Milo: Tested models include the Clara BW N365, Clara Colour N367, Elipsa 2E N605, Clara HD N249, Libra 2 N418, and Libra Colour N428 — though the 2025 Clara BW P365 refresh is supported only after matching its specific hardware.

Mia: Sendra is a Figma plugin built by product designer Joe that turns existing Figma email designs into responsive HTML. According to the maker, the exported HTML renders in Gmail, Outlook, Apple Mail, Yahoo, AOL, and every major inbox, without needing a component library or rebuild.

Milo: You select a frame you've already designed, then control stacking, spacing, and sizing on mobile per element, plus dark mode overrides. The maker states every rendering rule is checked against real clients on real devices rather than preview tools, because previews miss problems like emails losing styling on phones. Sendra also hosts images so exports ship with real URLs, and it can send a test email from inside the plugin.

Mia: It's free for the first ten exports, with an unlimited Pro upgrade. The framing is the familiar pain point — designing emails in Figma and then rebuilding them in HTML, with issues often surfacing later in a forwarded screenshot. One community commenter who was converting emails from Figma to HTML at the time said Sendra helped them avoid dealing with code or technical setup.

Milo: Since these are maker claims and individual comments, Sendra's real-world reliability across every named client remains unverified by independent testing.

Mia: SpacebarX is a keyboard-first, local-first outliner covering notes, tasks, writing, Markdown, code, and projects. The maker says he built it because he was constantly splitting his brain between a notes app and a task app, with the context he needed somewhere else. It keeps everything in one nested outline, lets any thought become a task with an inline due date, and requires no project setup or predefined templates.

Milo: It works offline, keeps files on the user's machine, and syncs through your own Google Drive or Dropbox — no SpacebarX server stores, reads, or controls your notes. The maker says the interface stays instant even at a million items. The free tier is free forever; Pro adds visual views, saved searches, encrypted documents, version restore, Calendar, uploads, and bring-your-own-key AI, available monthly, yearly, or as a ninety-nine-dollar option.

Mia: Today view collects what needs attention, including overdue and upcoming tasks, and auto-rolls unfinished tasks into today. The current product is a PWA installable from a browser on desktop or mobile, with native desktop and mobile apps in progress. A Product Hunt commenter who tried Google Drive sync called it seamless, described SpacebarX as a blend of Workflowy and Notion, and said tasks due inside deeply nested lists automatically show up in Today view.

Milo: They also said a native mobile app would improve the current mobile experience, and asked when iCloud Sync will launch. OPML imports are supported from Workflowy, Dynalist, Obsidian, and Logseq.

Mia: There's a new iOS app called Yomi built for kids aged five to nine, and it gives them a reason to practice reading aloud in English and Finnish. A child picks a real short story at one of three reading levels and reads it out loud; speech recognition listens, and every spoken word feeds a little cat named Yomi, who is hungry again the next day, so daily practice comes with a built-in reason to come back.

Milo: And the design is notably gentle about mistakes. Kids can tap any word to hear it spoken, and wrong words are handled without drama or red marks. There's also a simple wizard that lets a child choose a name, a character, a place, and an event, and the app generates their own story for them to read back.

Mia: Designer Elina, launching her second Product Hunt product, says she built Yomi after workshops with first- and second-grade teachers and her own daughter's refusal to read aloud. The apps she tried, she says, made practice stressful, leaning on streaks and fire emojis. Yomi deliberately leaves out streaks, ads, accounts, data collection, leaderboards, coins, and kid-facing levels.

Milo: One honest boundary she states: Yomi will not teach a child to read — teachers and parents do that — so no effect on reading skill is being claimed. In comments, one parent praised the lack of gamification and expected the comprehension checks to be a hit, though the feature list doesn't actually describe those. Another suggested future additions like pictures in books or Yomi sounds.

Mia: Separately, there's a website called The Million Sad Ducks, where people can pan and zoom a world made of one million sad ducks for free. For one dollar, a duck becomes permanently happy, takes whatever name the user gives it, and earns a Certificate of Duck Happiness as a PDF. The ponds range from one dollar to one thousand dollars, depending on how dramatic a duck you want.

Milo: Because the duck is named at checkout, the project works as a gift, with the certificate going to the recipient. The maker, whom a commenter addressed as Bogdan, calls it the Million Dollar Homepage with ducks, explaining that a pixel was ad space while a duck is a gift.

Mia: Every duck's position, colour, pose, and accessory derives from its ID number, so the million-duck world is a pure function, and the database stores only the ducks people made happy. The certificate is a full A4 landscape document with seals, laurels, three fictional signatories, and a Department of Avian Emotional Welfare. The project was built solo, shipped, and takes real payments, and the maker asked whether one dollar feels like the right price.

Milo: One commenter said their duck, Duckingston III, is now a happy quacky. Another said the certificate made them laugh and called the Department of Avian Emotional Welfare exactly the right amount of fake-official, but asked whether a typo or autocorrected name can be fixed after paying — and the discussion included no answer from the maker.

Mia: So to close it out, GLM-5.3-Flash is pitched as the first natively multimodal model in the GLM-5 line, with 320 billion total parameters but only 18 billion active, while outperforming its predecessor at a tenth of the price.

Milo: And the race for open-weight flash tiers, plus taming production agents with control planes and trajectory-aware gateways, framed a lot of that discussion today.

Mia: Thanks for tuning in — take care out there.