Open Yap 1K
Channel-separated English natural two-speaker conversations
Sample audio
Samples stream here for listening. All rights reserved: no downloading or redistribution. Terms of use.
Data description
The Agentic Data Company advances conversational AI with Open Yap 1K, the largest free collection of natural conversation ever released: 1,000 hours of two-speaker English speech, open to labs and research teams worldwide.
With this open release, we aim to close a gap in the literature: conversation recorded as it happens in real life. Fisher and Switchboard assigned partners and topics to maximize variety, but the phone network capped their audio near 4 kHz. Newer corpora record at full bandwidth but still pair strangers. Open Yap 1K assigns nothing: each speaker invites someone they already know and talks freely among friends and family.
That choice shows in the speech. People who know each other interrupt more, backchannel more, and leave shorter gaps between turns: the behaviour a full-duplex model has to learn. We compute these dynamics for every conversation (overlap, turn-taking gaps, speaking rate).
These 1,000 hours are one part of a larger licensed corpus we build with frontier labs and research teams. We release them because progress in conversational AI is slower than it needs to be, and open data is the fastest way to change that for everyone.
| Detail | Value |
|---|---|
| Speakers per recording | 2 |
| Channels | Dual (one file per speaker) |
| Audio format | .flac / .wav |
| Sample rate | 48 kHz |
| Bit depth | 16-bit PCM |
| Languages | English (en) |
| Metadata format | .jsonUTF-8 |
| Transcription type | ASR, word-level with timings(not human-verified) |
| Transcription engine | Deepgram Nova-3 |
Intended use
Speech-to-speech and full-duplex
Both sides as independent signals on one timeline. Overlap and turn-taking intact.
Expressive TTS
Spontaneous prosody on clean, isolated 48 kHz tracks. Laughter, fillers, and hesitation.
Audio understanding
Unprompted speech with word-level transcripts and speaker metadata.
Comparable datasets
Open Yap 1K is the largest publicly available dataset of natural two-speaker English conversation, licensed for commercial use. For this comparison, non-commercial releases are left out, as well as scripted corpora and corpora assembled by diarizing in-the-wild audio.
Delivered hours
Speaker demographics
All fields are self-reported by speakers.
Gender
n=239Age
n=239Education
n=239Native language
n=239Childhood country
n=239Speaker contribution
n=239Conversation statistics
Relationship between speakers
n=1,602Spoken language
Conversation length
From the longer of its two tracks. It ships as duration_seconds.
Words per conversation
Both speakers’ transcript entries, added together.
Fillers, laughter, coughs and noise markers count alongside words, so this runs higher than a plain word count. Each entry carries its type in the transcript file.
Audio metrics
Signal
DNSMOS background noise
Effective bandwidth
The highest frequency at which a track still carries real sound. A file header can say 48 kHz while the audio inside stops at 8 kHz.
We read the frequency spectrum of a track’s speech in one pass, up to 262,144 samples, or 5.5 seconds at 48 kHz. We smooth it over 100 Hz either side, then take the highest frequency within 60 dB of the loudest.
Telephone audio cliffs near 4 kHz. A Bluetooth headset in hands-free mode caps near 8 kHz. It ships as effective_bandwidth_hz.
Conversational dynamics
Overlap
The share of talking time where both people speak at once. Almost none usually means echo cancellation gated one microphone, which erases the very thing a full-duplex model needs to learn. Below 3 percent we flag the conversation.
We cut both tracks into 20 millisecond frames. A frame counts as speech when it passes three times that track’s own noise floor, or a fixed floor of 0.004, whichever is higher. The figure is the frames where both speak, divided by the frames where either does.
Turn-taking gap
The silence between one person finishing and the other starting. Get this wrong in a voice model and it either interrupts the user or leaves them hanging.
We merge the voiced stretches of both tracks onto one timeline. Every change of speaker gives one gap, and the figure is the median of them. Where the next speaker started early the two are talking at once, so that goes under Overlap instead.
CANDOR, a large public corpus of natural conversation, reports 380 milliseconds. It ships as turn_taking_gap_ms.
Speech dominance
Speaker A’s share of the talking, where 0.5 is an even split.
We count the 20 millisecond frames in which A is voiced, then divide by the frames in which either speaker is. Overlapping speech counts for both people, so the two shares add up to more than 1 when they talk at once.
The metadata field speech_dominance is a different number. That one is A’s share of the transcript words.
Speaking rate
How fast a person speaks, in words per minute of actual speaking time.
The denominator is the choice that matters, so we tested three across 28 channels of real conversation. Dividing by the whole recording gave 74 words per minute at the median. That measures how much of the call each person held, not how fast they talk. Voice activity detection gave 192, but it counts breaths, laughter and a partner’s bleed as talking, and it swung from 82 to 281 on the same audio.
So we sum each word’s own duration from the transcript and divide the word count by that. Only entries typed as words count, on both sides of the division. It ships as avg_wpm.
Conformance
Provenance
All audio was recorded on our own platform.
- Speakers register, give explicit consent before their first recording, and are paid for their time.
- Demographics are self-reported at registration, before any recording, and are never inferred from audio.
- Speaker identifiers are pseudonymous and stable within the release. Names, contact details and account identifiers are excluded.
Quality assurance
Two checks run before a conversation is eligible for delivery.
Human linguistic QA
A human reviews each conversation for language proficiency, accent classification, and ratings for naturalness and expressivity.
Transcript screening
Every transcript is passed through an LLM screen for personally identifying information and for content-policy violations. Flagged conversations are excluded.
Metadata
manifest.json
meta.json
speaker_a_meta.json
speaker_b_meta.json
speaker_a_transcript.json
speaker_b_transcript.json
speaker_a_dnsmos.json
speaker_b_dnsmos.json
Citation
Published work using this dataset should cite it as below, including the version.
The Agentic Data Company. Open Yap 1K: channel-separated English natural two-speaker conversations. 2026. Version 1.0. https://theagenticdatacompany.com/open-yap-1k
@misc{openyap1k,
title = {Open Yap 1K: Channel-Separated English Natural Two-Speaker Conversations},
author = {The Agentic Data Company},
year = {2026},
version = {1.0},
publisher = {The Agentic Data Company},
url = {https://theagenticdatacompany.com/open-yap-1k},
}Access
This is a public release: it is not licensed exclusively, anyone may request it, and it is free for both commercial and research use. Access is granted per recipient under a data use agreement.
Request
Tell us who you are, what you are building, and how the audio will be used.
Review
We read every request. Expect a reply within a few hours, either way.
Delivery
Approved recipients get a portal account and either a direct download or delivery into their own S3 bucket.
Use under the agreement
- Commercial use, including in products you sell
- Research use, published or internal
- Training, fine-tuning and evaluating models, and deploying what you train
- Internal copies, and access for staff and contractors under the same terms
- Retaining the delivered dataset, and anything derived from it, after a speaker withdraws
- Redistributing, resharing, sublicensing or reselling the dataset or any part of it
- Attempting to identify a speaker, or link a recording to any external record
- Creating voice clones, replicas or generative reproductions identifiable as a speaker in the corpus
- Holding the dataset under weaker controls than your own confidential material
- Retaining any copy of the dataset after a breach of these terms, or after a written request from The Agentic Data Company
