Call analytics, meeting assistants, voice agents and media indexing all need to know who spoke when. That is speaker diarization, and training or evaluating it requires carefully segmented reference audio.
What a Diarization Label Contains
- Start and end timestamps for every speech segment
- A consistent speaker ID across the whole recording
- Overlapping speech marked explicitly when two people talk at once
- Non-speech regions such as silence, music, hold tones and background noise
How Diarization Is Measured
The standard metric is Diarization Error Rate (DER): the sum of missed speech, false alarms (non-speech labeled as speech) and speaker confusion, divided by total reference speech time. Evaluations usually ignore a 0.25-second collar around each boundary to allow for small human timing differences.
Because the reference labels are the yardstick, errors in the ground truth directly distort the score. A model can look worse, or better, than it really is.
Where Diarization Breaks
- Overlapping speech during interruptions and crosstalk
- Short turns such as “yes” or “okay” that are easy to miss or misattribute
- Similar voices within the same recording
- Noisy channels such as phone lines, body cameras and far-field microphones
Correct Machine Output or Label From Scratch?
Pre-segmenting audio with a model and having humans correct boundaries and speaker IDs is usually faster and more consistent than labeling from silence. The trade-off is pre-annotation bias: reviewers may accept a wrong boundary because it looks plausible. Spot-checking a sample from scratch keeps that bias visible.
Guidelines to Agree Up Front
Define the minimum segment length, how to treat backchannels and laughter, how to label unknown or new speakers, and whether overlaps are split or double-labeled. These choices affect DER more than most teams expect.
About Smart Annotahub
Smart Annotahub is a managed data annotation company based in Ha Noi, Viet Nam. Our in-house annotators, QA leads and project managers turn raw image, video, 3D, geospatial, text and audio data into training-ready datasets for teams building computer vision, robotics and language AI. Every project starts with a free pilot, runs on your guidelines and tools, and ships with multi-stage quality checks.
Our Services
Frequently Asked Questions
How does the free pilot work?
Send us a sample of your data and your guidelines. We annotate it at no cost, report accuracy and turnaround, and return a precise quote, usually within a few working days.
How do you ensure annotation quality?
Every batch passes annotator self-checks, peer review and a dedicated QA lead. We agree accuracy targets up front and share QA reports with each delivery.
Can you work in our annotation tool?
Yes. Our team works in your platform or ours and delivers in the formats your pipeline expects, such as COCO, YOLO, Pascal VOC or custom JSON.
How is my data kept secure?
All work is done by our in-house team under NDA, with role-based access and no data leaving approved environments. See our Data Privacy Notice for details.
How is pricing calculated?
Per object, per hour or per project, depending on the task. See Pricing for reference rates, or request a pilot for an exact quote.
Want to see the difference on your own data?
Request Free Pilot





















