Why We Chose Experience Over Features And How Hard That Was
Speaker diarization was our most requested feature, and we said, "Not yet." Here's what we built instead, why the discipline was harder than the engineering, and what we finally shipped in July 2026.

TL;DR
- We passed on the easy shortcut: When everyone asked for speaker diarization, we delayed it. We wanted to build proper, audio-based speaker attribution instead of taking the easy route of matching names from calendar data.
- We fixed the unglamorous stuff: Instead of chasing the next big feature, we spent months working on things people rarely talk about: making onboarding smoother, adding simple reminders to hit record, and letting people save files wherever they wanted.
- Growth came from getting the basics right: Adoption went up just by fixing these small frustrations. It showed us that making the existing experience work well can matter more than adding another shiny feature.
- Patience paid off: When real, on-device diarization finally launched in July 2026, it landed in an app that was already much better to use. We took the longer route instead of rushing to ship a feature just because people were asking for it.
The day after Meetily went public, the feature requests started coming in.
At the top of almost every list: speaker diarization. Who said what in the meeting. Not just a transcript and an attributed one. It made sense. It was the obvious next step. And we said: not yet.
This is the story of why, and what we did instead.
There's a shortcut available here. You can attribute speakers from account or calendar data - you know who was invited to the meeting, so you label the transcript with their names. It isn't voice recognition. It's database matching. We could have done the same thing quickly. It would have looked like the feature customers wanted. We chose not to, because it wasn't actually the feature customers wanted, it was a shortcut dressed up as one.
Real speaker diarization means the system listens, distinguishes voices, and attributes speech accurately without knowing who the speakers are in advance. That's a meaningfully harder problem. We decided that if we were going to build it, we would build it properly with genuine R&D, proper validation, and the kind of robustness that a privacy-first, local application demands. We believed we could be one of the first self-hosted applications to do this well. But doing it well would take time.
So we looked away from the roadmap and looked at the application.
What we saw was an experience with rough edges. Not broken but not seamless. Users were getting lost in the setup. First impressions were inconsistent. Small frustrations were accumulating in ways that weren't visible until you looked at the data and the feedback carefully.
We made a decision about what Meetily should feel like. Simple. Purposeful. We wanted the UI to feel like a Sony recorder from the late 1970s - a red button for recording, a clean black and grey theme. The best recording device you ever used had one job and did it without confusion. You pressed record - it recorded! That was the experience we were after.
What followed was unglamorous work. We combed through the onboarding from the website to the first time a user opened the application. We looked at every point where someone might give up or feel uncertain. We fixed things that didn't make headlines.
Two pieces of feedback kept appearing. First: users were forgetting to start recording when their meetings began, and forgetting to stop when they ended. They'd sit down for a call, get absorbed in the conversation, and realize twenty minutes in that nothing had been captured. We built a prompt, a gentle nudge when a meeting starts and stops. Second: users wanted to export their summaries to a folder of their own choosing, not wherever we defaulted to. We built that too. Neither of these features generated excitement. Both of them made Meetily significantly more useful in practice.
Adoption increased. Not because we added the feature everyone had asked for, but because the experience of using what already existed became reliably good.
The internal discipline this required was harder than any technical problem we solved. The engineering team wanted to build speaker diarization. The community was asking for it by name. Choosing to spend months improving onboarding and a meeting prompt rather than shipping the headline feature needed constant justification to the team, to ourselves. The temptation to build for recognition rather than for quality is real, and it doesn't go away just because you're aware of it.
Our internal principle became: do not release a feature for name's sake. Research it. Implement it robustly. Make sure what you already have works seamlessly before you layer something new on top of it.
Speaker diarization shipped in July 2026. Genuine diarization from the audio, locally processed, no account data required. We haven't seen another self-hosted meeting application do it this way. The reception was better than we expected, not just because the feature worked, but because it arrived into an application that was already working well. It had somewhere solid to land.
Features built in a hurry for name's sake add weight to a product without adding value. Features built with patience, after the foundation is right, compound. One makes your users louder. The other makes them stay.
We chose patience and discipline. Most days, it was the harder choice.
Frequently Asked Questions
From the audio. Meetily listens to the recording and separates it into different speakers based on their voices. It labels them Speaker 1, Speaker 2, and so on. You can then name each speaker yourself, and those names carry over to the transcript and summary and are remembered for future meetings.
There is no voiceprint database, and Meetily doesn't match speakers against calendar invitations or account information. That's also why it works the same way with an imported recording as it does with a live call.
No. Speaker diarization and speaker identification are Meetily Pro features, introduced in Pro 1.8.2 on 22 July 2026. They are also included in Enterprise.
The free, MIT-licensed Community Edition gives you local transcription and AI summaries, but it doesn't separate the recording into different speakers or label who said what. You can try Pro with a 14-day trial, no card required.
Yes. Diarization runs on-device, as part of the same local pipeline that handles transcription with Whisper or Parakeet. Your audio doesn't need to leave your machine to figure out who spoke when.
Summaries are a separate choice. You can run a local model, in which case nothing leaves your machine, or use your own API key, in which case only the transcript text is sent to the provider you choose.
Yes. Meetily Pro can detect when a meeting starts and ends and prompt you, so you don't get twenty minutes into a call before realizing you forgot to record.
Pro can also start and stop recording automatically, with a countdown that you can cancel. This is turned off by default and can be enabled in Settings. The Community Edition uses manual start and stop, either from the interface or with a keyboard shortcut.
Ready to try Meetily?
Join 475,000+ users who use Meetily for private meeting transcription. No bots, privacy first. Community Edition free.
Star on GitHub (29K+) · Open source & self-hostable
Get Started with Meetily
Meetily Pro
Advanced features for individuals and teams.


