Turn-taking evals
Hand-labelled turn events on unscripted conversations, held out for evaluation
Sample audio
Data description
A held-out evaluation set for turn-taking: unscripted two-speaker English conversations, hand-labelled with the events a voice agent has to get right. Every end of turn, interruption, backchannel and mid-turn pause is marked with its onset on the timeline. Each event is labelled by hand, checked by a second annotator, and kept only where both agree. The people recording are friends and family at home, not actors in a studio, so the timing and the acoustics match what an agent meets in production.
Each speaker’s microphone is captured locally as uncompressed 16-bit PCM and delivered as a separate stem, so a model can be fed one side of the conversation and scored on when it would have spoken against what the other person actually did. Labels ship as a time-aligned events.jsonl per conversation, alongside word-timed per-speaker transcripts and each speaker’s age range, gender, native language and derived accent. None of the audio has been published, so it cannot have leaked into a training set.
Built for evaluating end-of-turn detection, interruption handling and backchannel timing in full-duplex and voice-agent systems, with splits by accent and first language that show where timing breaks down for non-native speakers.
| Detail | Value |
|---|---|
| Speakers per recording | 2 |
| Channels | Dual (one file per speaker) |
| Audio format | .flac / .wav |
| Sample rate | 48 kHz |
| Bit depth | 16-bit PCM |
| Languages | English (en) |
| Turn events | End of turn, interruption, backchannel, mid-turn pausehand-labelled |
| Metadata format | .json / .jsonlUTF-8 |
| Transcription type | ASR, word-level with timings(not human-verified) |
| Transcription engine | Deepgram Nova-3 |
