The Agentic Data CompanyMultilingual Conversations
Request access

Multilingual Conversations

Unscripted two-speaker conversations in the pair’s own shared language

Not published. Available to buyers under a commercial license.

Sample audio

18:46
Speaker A
Speaker B

Data description

Multilingual Conversations is a dataset of unscripted, two-speaker conversations, each held in the pair’s shared native or fluent language. There are no prompts or topics — people talk as they naturally would, so the audio keeps the prosody, disfluencies, and overlap of real conversation. It spans European and Scandinavian languages alongside a long tail of lower-resource languages.

Each speaker’s microphone is captured locally as uncompressed 16-bit PCM and delivered as a separate stem, so every recording arrives with its native dynamics intact and a BCP-47 language code attached. It works as a low-resource-language training set, a cross-lingual prosody corpus, and a multilingual evaluation benchmark for ASR and conversational TTS.

DetailValue
Speakers per recording
2
Channels
Dual (one file per speaker)
Audio format
.flac / .wav
Sample rate
48 kHz
Bit depth
16-bit PCM
Languages
14English excluded
Metadata format
.jsonUTF-8
Transcription type
ASR, word-level with timings(not human-verified)
Transcription engine
Deepgram Nova-3