Accented English
Two-speaker English conversations across many accents and first languages
Sample audio
Data description
Unscripted, two-speaker English conversations, with the two people in separate locations and each microphone recorded on its own track. There are no prompts or topics, so the recordings carry the real turn-taking of natural talk — the pauses, interruptions, and moments where both voices overlap.
Accent is provenance-based, derived from each speaker’s self-reported demographics: the country they grew up in (origin_country) and whether English is their first language (english_native).
Each speaker’s microphone is captured locally as uncompressed 16-bit PCM and delivered as a separate stem, alongside time-aligned transcripts and a JSON record of each speaker’s age range, gender, native language, and derived accent. It’s built for accent-fair, full-duplex speech AI: training and evaluation sets that expose where ASR fails on non-native and under-represented English, and where turn-taking models break down.
| Detail | Value |
|---|---|
| Speakers per recording | 2 |
| Channels | Dual (one file per speaker) |
| Audio format | .flac / .wav |
| Sample rate | 48 kHz |
| Bit depth | 16-bit PCM |
| Languages | English (en) |
| First languages represented | 38 |
| Metadata format | .jsonUTF-8 |
| Transcription type | ASR, word-level with timings(not human-verified) |
| Transcription engine | Deepgram Nova-3 |
