Work sample: annotator calibration for a TTS rating study
Paste the rating matrix a study produces. This returns inter-rater agreement computed three ways, which rater is drifting and in which direction, the clips the rubric is actually failing on, and a written QA assessment ready to send. It reports three coefficients because they disagree, and the disagreement is the finding.
A working prototype by Edward Tay, built for the Cartesia Language Lead application. Mine, not Cartesia code. Method: cartesia.edwardtay.com
CSV or tab separated. First row is rater names, first column is the clip id. Leave a cell blank where a rater did not score that clip. Nothing is uploaded: the arithmetic runs in this browser.