Is there an algorithm for Speaker Error Rate for speech-to-text diarization?

Some speech-to-text services, like Google Speech-to-Text, offer speaker differentiation via diarization which attempts to identify and separate multiple speakers on a single audio recording. This is often needed when multiple speakers are in a meeting room sharing a single microphone.

Is there an algorithm and implementation to calculate the correctness of speaker separation?

This would be used in conjunction with Word Error Rate which is often used to test correctness of baseline transcription.

Solution

The commonly used approach for this appears to be the Diarization Error Rate (DER) defined by NIST in the NIST-RT projects.

A newer evaluation metric is the Jaccard Error Rate (JER) introduced in DIHARD II: The Second DIHARD Speech Diarization Challenge.

Two projects for measuring these include:

https://github.com/nryant/dscore
https://github.com/wq2012/SimpleDER

DER is referenced in these papers:

A Comparison of Neural Network Feature Transforms for Speaker Diarization
The ICSI RT-09 Speaker Diarization System

Math.Sin() gives incorrect value
How to run my python script when the sunOS is start booting
Express-session: not resetting cookie expiration on each request
Getting a stack overflow exception when normalizing a vector
Edit default summary function in R gives error for multiple variables
What was a For loop? Why isn't it needed in R?
How to use download button in shiny and save results in various formats (csv, texte, pdf, spss...)?
Why are there two assignment operators, `<-` and `->` in R?
lm()$assign: what is it?
How to get the value of list(...) in R and S functions
Design matrix for MLM from library(lme4) with fixed and random effects
how to generate elements not included in my sample
Create a matrix with gradually changing values without a for loop
Emacs ESS and S-plus ( S+ ) 8.1 compatability
How to lag date-index in a time-series in R?
Nonlinear regression in R / S
Calling R from S-Plus?