The Agentic Data CompanyAccented English
Request access

Accented English

Two-speaker English conversations across many accents and first languages

Not published. Available to buyers under a commercial license.

Sample audio

16:57
Speaker A
Speaker B

Data description

Unscripted, two-speaker English conversations, with the two people in separate locations and each microphone recorded on its own track. There are no prompts or topics, so the recordings carry the real turn-taking of natural talk — the pauses, interruptions, and moments where both voices overlap.

Accent is provenance-based, derived from each speaker’s self-reported demographics: the country they grew up in (origin_country) and whether English is their first language (english_native).

Each speaker’s microphone is captured locally as uncompressed 16-bit PCM and delivered as a separate stem, alongside time-aligned transcripts and a JSON record of each speaker’s age range, gender, native language, and derived accent. It’s built for accent-fair, full-duplex speech AI: training and evaluation sets that expose where ASR fails on non-native and under-represented English, and where turn-taking models break down.

DetailValue
Speakers per recording
2
Channels
Dual (one file per speaker)
Audio format
.flac / .wav
Sample rate
48 kHz
Bit depth
16-bit PCM
Languages
English (en)
First languages represented
38
Metadata format
.jsonUTF-8
Transcription type
ASR, word-level with timings(not human-verified)
Transcription engine
Deepgram Nova-3