Skip to main content
New in Meetily Pro 1.8.1

Speaker Diarization: Know Who Said What

Speaker diarization labels every voice in your meeting transcript - answering who spoke when - so a recording reads like a conversation instead of one wall of text. Meetily Pro identifies each speaker on-device, live as you record and on audio you import or re-transcribe. Rename a speaker once and it updates everywhere. No cloud upload.

  • Runs On-Device
  • Live & On Imported Audio
  • Rename Once, Updates Everywhere
  • Works With Any Platform
  • Attributes who said what automatically
  • Diarization runs locally on your device
  • Live, imported, and re-transcribed audio
  • Merge duplicates, rename speakers

Speaker diarization, answered

Last updated: Reviewed by the Meetily team

What is speaker diarization?

Speaker diarization is the process of partitioning an audio recording by speaker - answering “who spoke when” - so a transcript reads as attributed dialogue instead of one undifferentiated block of text. It groups the speech segments that belong to the same voice and labels each one (Speaker 1, Speaker 2, or a real name). It is distinct from speech recognition, which turns audio into words; diarization adds the layer of who said each of those words. In meeting notes it is what turns a raw transcript into a readable record of a conversation.

How does Meetily label who said what?

Meetily Pro identifies each speaker and labels every segment of the transcript with who said it - live as you record, and on audio you import or re-transcribe. It runs on-device, so no audio is uploaded to attribute speakers: the same local pipeline that transcribes with Whisper or Parakeet also handles diarization. Rename a speaker once and the name updates across the whole transcript, and merge duplicates when one person gets split into two labels. Speaker diarization arrived in Meetily Pro 1.8.1.

Speaker diarization vs. speaker identification

Speaker diarization separates a recording into distinct speakers without knowing who they are - it produces anonymous labels like Speaker 1 and Speaker 2. Speaker identification goes a step further and matches a voice to a known identity. In practice, product interfaces use the terms interchangeably: Meetily diarizes each meeting into separate speakers, then lets you name them once so the label carries across the transcript. No enrolled voiceprint database is required - you stay in control of who each speaker is.

Open-source speaker diarization in 2026

Open-source diarization usually means assembling a pipeline yourself - for example OpenAI Whisper for transcription plus pyannote.audio or NVIDIA NeMo for speaker segmentation - which takes setup, tuning, and GPU know-how. Meetily takes a packaged approach: its open-core app pairs on-device Whisper/Parakeet transcription with built-in diarization in the Pro edition, so you get who-said-what labeling without wiring a pipeline together. Everything runs locally on Windows or macOS.

Open Source & Privacy-First

Speaker Attribution Without the Cloud

Most meeting tools diarize by uploading your audio to their servers. Meetily attributes speakers on your own device.

On-Device Attribution

  • Speakers are separated and labeled locally on your machine
  • No audio is uploaded to identify who is speaking
  • The same local pipeline transcribes and diarizes

You Stay in Control

  • Rename any speaker once and it updates everywhere
  • Merge duplicates when one person is split in two
  • No enrolled voiceprint database required

Built for Sensitive Meetings

  • Legal, healthcare, HR, and executive calls stay on your hardware
  • Supports GDPR and HIPAA compliance by design
  • Open-core: inspect how it works

How Does Meetily’s Speaker Diarization Compare?

Most meeting tools offer speaker labels, but they run in the cloud. Whisper on its own does not diarize at all. Here is how the approaches differ.

Speaker diarization: Meetily vs cloud meeting tools vs a raw Whisper setup
CapabilityMeetily ProCloud tools (Otter, Fireflies, Granola, Fathom)Raw Whisper
Speaker diarizationYesYesNo (needs pyannote/NeMo)
Where it runsOn-deviceVendor cloudYour own setup
Live during the callYesVariesNo (batch only)
On imported audioYesVariesYes (if configured)
Rename once, updates everywhereYesVariesNo (manual)
Setup requiredNone (built in)Account + botPipeline assembly
Source codeOpen-core (MIT app)ProprietaryOpen source

Compare Meetily directly with Otter, Fireflies, and Granola.

How Does Meetily’s Speaker Diarization Work?

Automatic attribution you can correct in seconds

1. Record or Import

Record a live call, or import an existing audio file. Meetily transcribes on-device and attributes each segment to a speaker as it goes - or across the whole file for imports and re-transcriptions.

2. Name Your Speakers

Meetily separates the conversation into distinct voices. Rename a speaker once - from Speaker 2 to a real name - and the label updates across the entire transcript automatically.

3. Merge & Refine

If one person was split into two labels, merge them into a single speaker. Your final transcript reads as clean, attributed dialogue - and feeds clearer, speaker-aware summaries.

Who Needs Speaker Diarization?

Anywhere it matters who said what, on the record

Legal Teams

Depositions, intake calls, and multi-party negotiations where every line has to be attributed to the right person - captured locally to protect privilege.

Healthcare & Therapy

Separate clinician from patient in consultation notes and SOAP documentation, with attribution that stays on your own device.

Government & Finance

Hearings, public meetings, and board minutes that need speaker-attributed records for retention and review - kept in your own infrastructure.

Research & Interviews

User interviews, focus groups, and qualitative research where analysis depends on knowing which participant said what.

Sales & Recruiting

Discovery calls and panel interviews where attributed transcripts make follow-up, coaching, and hand-off notes far clearer.

Distributed Teams

Any team on Zoom, Teams, or Google Meet that wants meeting records reading as a real conversation, not an anonymous transcript.

Built on Meetily's On-Device Foundation

The same locally-run engine trusted by a growing open-source community

25K+
GitHub Stars
369K+
Downloads
99+
Languages Supported
On-Device
Speaker Attribution

Speaker Diarization: Common Questions

Speaker diarization is the process of partitioning an audio recording by speaker - answering "who spoke when" - so a transcript reads as attributed dialogue instead of one undifferentiated block of text. It groups speech segments that belong to the same voice and labels each one (Speaker 1, Speaker 2, or a real name). It is distinct from speech recognition, which turns audio into words; diarization adds the layer of who said each of those words.
Meetily Pro identifies each speaker and labels every segment of the transcript with who said it - live as you record and on audio you import or re-transcribe. It runs on-device, so no audio is uploaded to label speakers. Rename a speaker once and the name updates across the entire transcript, and you can merge duplicates if the same person was split into two. Diarization is part of Meetily Pro (since version 1.8.1); the app is open-core with 25K+ GitHub stars.
Speaker diarization separates a recording into distinct speakers without knowing who they are - it produces anonymous labels like Speaker 1 and Speaker 2. Speaker identification (or recognition) goes a step further and matches a voice to a known identity. In practice, product UIs use the terms interchangeably: Meetily diarizes each meeting into separate speakers, then lets you name them once so the label carries across the transcript. No enrolled voiceprint database is required.
In Meetily, diarization runs entirely on-device. Audio never leaves your machine to be split by speaker - the same on-device pipeline that transcribes with Whisper or Parakeet also attributes the speech. This matters for confidential meetings: legal, healthcare, HR, and executive conversations stay on your own hardware. For summaries you separately choose a local model or your own API key (transcript text only, opt-in per meeting).
Yes. Meetily attributes speakers live as you record, and also on audio files you import or on past meetings you re-transcribe. Import a recording from any source, run it through Meetily, and the resulting transcript is labeled by speaker. Re-transcribing an older meeting applies diarization to it as well, so your back catalog can gain speaker labels.
Open-source diarization usually means assembling a pipeline yourself - for example OpenAI Whisper for transcription plus pyannote.audio or NVIDIA NeMo for speaker segmentation - which takes setup, tuning, and GPU know-how. Meetily takes a packaged approach: its open-core app pairs on-device Whisper/Parakeet transcription with built-in diarization in the Pro edition, so you get who-said-what labeling without wiring a pipeline together. Everything runs locally on Windows or macOS.
Speaker diarization is a Meetily Pro feature (introduced in Pro 1.8.1). The free, MIT-licensed Community Edition includes local transcription and AI summaries, but not diarization. Pro is $10/user/month billed annually (early bird) and adds speaker identification, more accurate models, custom summary templates, and priority support. A 14-day trial is available with no card required.
Diarization accuracy depends on audio quality, how many speakers there are, how much they overlap, and microphone setup - clean, low-overlap audio attributes more reliably than a crowded room on one mic. Meetily lets you correct attribution: rename a speaker once and it propagates, and merge duplicates when one person is split across two labels. That keeps the final transcript accurate even when the automatic pass is imperfect.
Meetily captures audio from your own device rather than joining the call as a bot, so speaker diarization works the same across Zoom, Microsoft Teams, Google Meet, and any other platform. There is no per-platform integration to configure. It works on both live calls and imported recordings, on Windows and macOS.

Know Who Said What in Every Meeting

Speaker diarization is in Meetily Pro - $10/user/month billed annually, with a 14-day trial. The free Community Edition covers local transcription and summaries.

On-Device Attribution · Live & Imported Audio · Windows & macOS · GDPR & HIPAA Compliant by Design

Why Choose Meetily Pro?

Privacy-First Architecture

Transcription runs 100% locally; recordings and audio never leave your device. For summaries, you choose your provider: local AI or your own API key. Compliant by design with GDPR and HIPAA workflows.

Advanced AI Features

Real-time transcription, speaker identification, action item extraction, meeting summaries, and custom AI workflows.

Priority Support

Dedicated support team, direct access to engineers, custom integration assistance, and priority bug fixes.

Flexible Deployment

On-Device deployment. Local file access for custom integrations and workflows.

Buy Now

Have questions? or join our GitHub community