Multi-party conversations
Unscripted conversations between three to five people, channel-separated
Sample audio
Data description
Unscripted conversations between three to five people who already know each other, in English. Everyone joins from their own location, so each voice is captured by its own microphone with no bleed from the others. There are no prompts or topics, and the recordings keep what a group adds to a two-person call: several people bidding for the floor at once, side conversations that split off and rejoin, and the short vocal cues that pass the turn around a group.
Each speaker’s microphone is captured locally as uncompressed 16-bit PCM and delivered as a separate stem, aligned to the others, with word-timed per-speaker transcripts and a JSON record of each speaker’s age range, gender and native language. With every voice on its own clean track, who spoke when is exact for every overlap, and no hand-labelling is needed. Mixed down, or convolved with a room response, the stems become far-field audio with per-speaker ground truth that a same-room recording cannot give.
Built for speaker diarization, speech separation and multi-talker ASR, and for full-duplex models that have to hold their own in a group: predicting who speaks next, handling overlapping speech, and knowing when a remark is meant for them.
| Detail | Value |
|---|---|
| Speakers per recording | 3to5 |
| Channels | Multi-track (one file per speaker) |
| Audio format | .flac / .wav |
| Sample rate | 48 kHz |
| Bit depth | 16-bit PCM |
| Languages | English (en) |
| Metadata format | .jsonUTF-8 |
| Transcription type | ASR, word-level with timings(not human-verified) |
| Transcription engine | Deepgram Nova-3 |
