What is speaker diarization?
Speaker diarization is the process of partitioning an audio recording by speaker - answering “who spoke when” - so a transcript reads as attributed dialogue instead of one undifferentiated block of text. It groups the speech segments that belong to the same voice and labels each one (Speaker 1, Speaker 2, or a real name). It is distinct from speech recognition, which turns audio into words; diarization adds the layer of who said each of those words. In meeting notes it is what turns a raw transcript into a readable record of a conversation.