smallest.ai's Pulse Model Takes the Diarization Crown

smallest.ai's Pulse Model Takes the Diarization Crown
Speaker diarization — knowing who said what — has quietly been the hardest problem in meeting transcription, and smallest.ai just posted a serious result. Its Pulse model topped the VoiceArena Diarization Bench v1 on October 7, 2026, hitting 24.4% DER across 139 conversations and 280+ speakers, a 1.7x improvement over ElevenLabs Scribe v2 (40.71%), AssemblyAI (44.64%), and Deepgram (67.13%) [4][5][6].
The gains hold up in both messy in-person meetings (24.7% DER) and cleaner online calls (22.4% DER), with streaming multi-speaker support across 21+ languages. That in-person number matters more than it looks — most benchmarks flatter themselves with clean Zoom audio, so a model that holds together in a noisy conference room is the one actually useful for real meeting intelligence.
Community reaction on X was straightforward congratulations mixed with the obvious follow-up question: how fast will this filter down into the consumer-facing meeting tools everyone already uses?
Hybrid Search Is Quietly Becoming the RAG Default
A steady drip of tutorials this cycle converged on the same architecture: pair BM25 keyword search with vector embeddings in a single retrieval pipeline. LangChain's EnsembleRetriever combined with ChromaDB and Nomic embeddings is the reference stack showing up everywhere, typically with 50/50 weighting and Reciprocal Rank Fusion to merge results [7][8][9].
The appeal is practical, not academic — pure semantic search is great at "find me the gist" but bad at "find me the exact phrase someone said." Hybrid search fixes that gap, which matters enormously once you're querying months of meeting transcripts instead of a handful of documents.
Builders on X are treating this less as a research curiosity and more as table stakes for anyone building a serious knowledge base on top of meeting data — a sign the RAG tooling around transcripts is standardizing fast.
Meeting Transcripts Are Becoming Living Knowledge Graphs
The most interesting trend this month isn't a single product — it's a pattern. Multiple projects are now converting raw meeting transcripts into structured, queryable knowledge graphs. Meetily pushes local transcription into Obsidian-based graphs with action items attached [10]. A separate project uses LLM extraction to build self-updating Neo4j graphs linking meetings, people, and tasks via relationships like ATTENDED [11]. And a tl;dv + Cognee integration builds entity graphs explicitly designed to answer questions like "who volunteered for X" or "what's blocking Y" [12].
This is the shift from transcription-as-archive to transcription-as-infrastructure. Nobody wants a pile of searchable text; they want a system that remembers relationships — who committed to what, when, and whether it happened. X threads on this topic increasingly use the phrase "second brain," and it's sticking because it's accurate: these systems are starting to behave less like note apps and more like institutional memory.
What This Means For Your Meetings
Put these four stories together and a clear shape emerges: the meeting intelligence stack is splitting into layers, and each layer is getting a serious upgrade independently. Capture is getting more accurate (Pulse's diarization leap), retrieval is getting smarter (hybrid search becoming default), and the output layer is evolving from flat summaries into structured knowledge graphs that understand relationships between people, decisions, and tasks. The assistant market comparisons are really a proxy war over which company stitches these layers together best.
For Proudfrog, this is validating territory. A Nordic-built tool that already combines accurate speaker ID with a knowledge graph across your entire meeting history isn't chasing a trend — it's built for where the category is heading. The diarization benchmarks show why "who said it" still needs dedicated engineering, not an afterthought bolted onto a transcription API. And the hybrid search pattern confirms that querying a meeting archive well requires more than vector search alone — exact names, project codenames, and specific commitments need keyword precision alongside semantic recall.
The practical upshot for professionals: your meeting history is becoming an asset you can interrogate, not just a transcript you skim once and forget. The tools catching up fastest are the ones treating accuracy, structure, and retrieval as one connected problem rather than three separate features.
Key takeaway: Meeting intelligence is maturing from "record and summarize" into "remember and reason" — and the tools that connect accurate capture to a structured, queryable knowledge graph will define the category.
Sources
- https://www.remogrid.com/blog/ai-tools/best-ai-meeting-assistants-2026
- https://sieva.co/best/ai-meeting-assistants/
- https://www.techrepublic.com/article/news-best-ai-meeting-note-takers-2026/
- https://docs.smallest.ai/models/speech-to-text/benchmarks/performance
- https://x.com/smallest_AI/status/2107680800045183003
- https://docs.smallest.ai/waves/model-cards/speech-to-text/pulse
- https://www.youtube.com/watch?v=Hn0UK2l1Nkw
- https://stackoverflow.com/questions/79477745/bm25retriever-chromadb-hybrid-search-optimization-using-langchain
- https://www.mintlify.com/chroma-core/chroma/guides/hybrid-search
- https://meetily.ai/blog/the-decision-makers-second-brain
- https://pub.towardsai.net/building-a-self-updating-knowledge-graph-from-meeting-notes-with-llm-extraction-and-neo4j-b02d3d62a251
- https://dev.to/insane_odyssey/turning-meeting-transcripts-into-an-ai-knowledge-graph-with-tldv-cognee-57e7
Get the daily briefing
AI, knowledge graphs, and the future of work — in your inbox every morning.
No spam. Unsubscribe anytime.