The Agentic Data CompanyOpen Yap 1K
Request access

Open Yap 1K

Channel-separated English natural two-speaker conversations

Sample audio

Speaker A
Speaker B

Samples stream here for listening. All rights reserved: no downloading or redistribution. Terms of use.

Hours of audio1,000
Conversations1,602
Unique speakers239
Avg. duration37.5 min

Data description

The Agentic Data Company advances conversational AI with Open Yap 1K, the largest free collection of natural conversation ever released: 1,000 hours of two-speaker English speech, open to labs and research teams worldwide.

With this open release, we aim to close a gap in the literature: conversation recorded as it happens in real life. Fisher and Switchboard assigned partners and topics to maximize variety, but the phone network capped their audio near 4 kHz. Newer corpora record at full bandwidth but still pair strangers. Open Yap 1K assigns nothing: each speaker invites someone they already know and talks freely among friends and family.

That choice shows in the speech. People who know each other interrupt more, backchannel more, and leave shorter gaps between turns: the behaviour a full-duplex model has to learn. We compute these dynamics for every conversation (overlap, turn-taking gaps, speaking rate).

These 1,000 hours are one part of a larger licensed corpus we build with frontier labs and research teams. We release them because progress in conversational AI is slower than it needs to be, and open data is the fastest way to change that for everyone.

DetailValue
Speakers per recording
2
Channels
Dual (one file per speaker)
Audio format
.flac / .wav
Sample rate
48 kHz
Bit depth
16-bit PCM
Languages
English (en)
Metadata format
.jsonUTF-8
Transcription type
ASR, word-level with timings(not human-verified)
Transcription engine
Deepgram Nova-3

Intended use

Speech-to-speech and full-duplex

Both sides as independent signals on one timeline. Overlap and turn-taking intact.

Expressive TTS

Spontaneous prosody on clean, isolated 48 kHz tracks. Laughter, fillers, and hesitation.

Audio understanding

Unprompted speech with word-level transcripts and speaker metadata.

Comparable datasets

Open Yap 1K is the largest publicly available dataset of natural two-speaker English conversation, licensed for commercial use. For this comparison, non-commercial releases are left out, as well as scripted corpora and corpora assembled by diarizing in-the-wild audio.

Delivered hours

WidebandTelephone band
Fisher English
2004, paid
1,959h
Open Yap 1K
2026, free
1,000h
Switchboard-2
1998, paid
898h
otoSpeech full-duplex
2026, free
280h
Switchboard-1
1993, paid
260h
AMI Meeting Corpus
2006, free
100h
CALLHOME English
1996, paid
56h
CALLFRIEND English
1996, paid
52h
Fisher English
8 kHz, telephone band
Open Yap 1K
48 kHz, wideband

Speaker demographics

All fields are self-reported by speakers.

Gender

n=239
Male12050%
Female11950%

Age

n=239
Under 255523.0%
25–347431.0%
35–445322.2%
45–544117.2%
55+166.7%

Education

n=239
Primary10.4%
Secondary7933.1%
Vocational3514.6%
Bachelor10543.9%
Master156.3%
PhD41.7%

Native language

n=239
English21288.7%
Arabic72.9%
Tagalog52.1%
German31.3%
Nepali20.8%
Other104.2%

Childhood country

n=239
United States12652.7%
South Africa2510.5%
Canada187.5%
United Kingdom104.2%
New Zealand93.8%
Philippines52.1%
Other4619.2%

Speaker contribution

n=239
Top contributor
1.8%
Top 10 contributors
17.9%

Conversation statistics

Relationship between speakers

n=1,602
Friends1,13670.9%
Colleagues17310.8%
Romantic partners1579.8%
Family1368.5%

Spoken language

English1,602100%

Conversation length

1.2176.3
p5
2.5
p50
30.1
p95
100.6
min

Words per conversation

16529.0k
p5
353
p50
4.5k
p95
17.3k
words
Unique words51,250
Total words9.8M

Audio metrics

Signal

DNSMOS background noise

SampleBAKShare
4.1
13.4%
4.0
30.2%
3.9
20.6%
3.8
12.5%
3.7
7.9%
3.6
5.5%
3.5
4.6%
3.4
3.0%
3.3
2.2%
p5
3.49
p50
3.97
p95
4.14

Effective bandwidth

524
p5
6.9
p50
11.6
p95
22.3
kHz

Conversational dynamics

Overlap

1.727.7
p5
2.8
p50
8.3
p95
20.9
% of voiced time

Turn-taking gap

2201440
p5
280
p50
580
p95
1120
ms

Speech dominance

0.160.87
p5
0.27
p50
0.56
p95
0.81
share

Speaking rate

142234
p5
156
p50
186
p95
218
words/min

Conformance

CheckPass
Track pairs sharing a common timeline anchor100%
Conversations with word-level transcripts100%
Tracks whose background noise is not intrusive100%
Tracks with zero clipped samples98.38%
Tracks above telephone bandwidth96.44%

Provenance

All audio was recorded on our own platform.

  • Speakers register, give explicit consent before their first recording, and are paid for their time.
  • Demographics are self-reported at registration, before any recording, and are never inferred from audio.
  • Speaker identifiers are pseudonymous and stable within the release. Names, contact details and account identifiers are excluded.

Quality assurance

Two checks run before a conversation is eligible for delivery.

01

Human linguistic QA

A human reviews each conversation for language proficiency, accent classification, and ratings for naturalness and expressivity.

02

Transcript screening

Every transcript is passed through an LLM screen for personally identifying information and for content-policy violations. Flagged conversations are excluded.

Metadata

Schema
conversations/
manifest.json
dataset_version: schema version
generated_at: ISO 8601
archive_contents: "full" | "incremental"
audio_format: "wav" | "flac"
total_conversations
total_duration_seconds
languages: BCP-47
conversation_ids: stable pseudonyms
transcript_schema_version
conv_a1b2c3d4e5f6/
speaker_a.flac
speaker_b.flac
meta.json
conversation_id: stable pseudonym
language: BCP-47
relationship: self-reported
duration_seconds
recorded_at: YYYY-MM-DD
speaker_a_id: stable pseudonym
speaker_b_id: stable pseudonym
conversation_summary: LLM, 2-3 sentences
topics[]
speech_dominance: speaker A's share of spoken words, 0-1 (0.5 = equal)
turn_taking_gap_ms: median gap between turns
speaker_a_meta.json
speaker_id: stable pseudonym
age_range: bucketed, e.g. "25-34"
gender: self-reported
country: ISO-3166 alpha-2, self-reported
education_level: self-reported
native_language: BCP-47, self-reported
accent: derived from self-reported demographics, not measured from the audio
origin_country: ISO-3166, childhood country
english_native
recording
sample_rate: Hz, as delivered in this archive
duration_seconds
bit_depth: 16
channels: 1
integrated_lufs: measured LUFS; audio ships un-normalized
true_peak_dbtp: measured dBTP
transcript: filename
avg_wpm: words per minute
echo_cancellation: applied at capture
headphones: exact on mobile, inferred on web
device: capture device or audio route
audio_metrics: measured from this track, at the rate this archive ships
noise_floor_dbfs: lower = quieter background. Null = no room tone to measure, not unmeasured
silence_profile: "room_tone", "noise_gated" (digital silence between turns), "dropout" (mic cut out mid-recording), "no_signal"
silent_while_partner_spoke_pct: % of the partner's speech this mic went silent for. 0 = overlap survived, unless either track's silence_profile is no_signal
effective_bandwidth_hz: highest frequency holding real energy: the mic's limit, not the container's
dnsmos_bak_median: ITU-T P.835 background noise, 1-5 (5 = not noticeable)
dnsmos_sig_median: the voice itself, 1-5
dnsmos_ovr_median: both together, 1-5
dnsmos_time_series: filename; the same three scores over time
speaker_b_meta.json
speaker_id: stable pseudonym
age_range: bucketed, e.g. "25-34"
gender: self-reported
country: ISO-3166 alpha-2, self-reported
education_level: self-reported
native_language: BCP-47, self-reported
accent: derived from self-reported demographics, not measured from the audio
origin_country: ISO-3166, childhood country
english_native
recording
sample_rate: Hz, as delivered in this archive
duration_seconds
bit_depth: 16
channels: 1
integrated_lufs: measured LUFS; audio ships un-normalized
true_peak_dbtp: measured dBTP
transcript: filename
avg_wpm: words per minute
echo_cancellation: applied at capture
headphones: exact on mobile, inferred on web
device: capture device or audio route
audio_metrics: measured from this track, at the rate this archive ships
noise_floor_dbfs: lower = quieter background. Null = no room tone to measure, not unmeasured
silence_profile: "room_tone", "noise_gated" (digital silence between turns), "dropout" (mic cut out mid-recording), "no_signal"
silent_while_partner_spoke_pct: % of the partner's speech this mic went silent for. 0 = overlap survived, unless either track's silence_profile is no_signal
effective_bandwidth_hz: highest frequency holding real energy: the mic's limit, not the container's
dnsmos_bak_median: ITU-T P.835 background noise, 1-5 (5 = not noticeable)
dnsmos_sig_median: the voice itself, 1-5
dnsmos_ovr_median: both together, 1-5
dnsmos_time_series: filename; the same three scores over time
speaker_a_transcript.json
conversation_id: stable pseudonym
speaker_index: "a" | "b"
language: BCP-47
text: full transcript
words[]
word
start: seconds; null where a reviewer typed the word in
end
type: "word" | "filler" | "laugh" | "cough" | "noise"
corrections_applied: a reviewer edited the ASR output
speaker_b_transcript.json
conversation_id: stable pseudonym
speaker_index: "a" | "b"
language: BCP-47
text: full transcript
words[]
word
start: seconds; null where a reviewer typed the word in
end
type: "word" | "filler" | "laugh" | "cough" | "noise"
corrections_applied: a reviewer edited the ASR output
speaker_a_dnsmos.json
conversation_id: stable pseudonym
speaker_index: "a" | "b"
metric: always "dnsmos_p835" — ITU-T P.835 scores predicted by Microsoft DNSMOS
model: model file and licence the scores came from
measurement_seconds: seconds of audio each score describes (fixed by the model)
step_seconds: seconds between consecutive measurements; measurements overlap
measurement_count: number of entries in each of the arrays below
bak[]: background-noise score per measurement, 1-5 (5 = not noticeable). Entry i covers [i*step_seconds, i*step_seconds+measurement_seconds) on the delivered timeline; null where the audio held too little speech to score
sig[]: speech-signal score per measurement, 1-5; null on the same entries as bak
ovr[]: overall score per measurement, 1-5; null on the same entries as bak
speaker_b_dnsmos.json
conversation_id: stable pseudonym
speaker_index: "a" | "b"
metric: always "dnsmos_p835" — ITU-T P.835 scores predicted by Microsoft DNSMOS
model: model file and licence the scores came from
measurement_seconds: seconds of audio each score describes (fixed by the model)
step_seconds: seconds between consecutive measurements; measurements overlap
measurement_count: number of entries in each of the arrays below
bak[]: background-noise score per measurement, 1-5 (5 = not noticeable). Entry i covers [i*step_seconds, i*step_seconds+measurement_seconds) on the delivered timeline; null where the audio held too little speech to score
sig[]: speech-signal score per measurement, 1-5; null on the same entries as bak
ovr[]: overall score per measurement, 1-5; null on the same entries as bak
conv_7f3e9d014a6c/
...

Citation

Published work using this dataset should cite it as below, including the version.

Reference

The Agentic Data Company. Open Yap 1K: channel-separated English natural two-speaker conversations. 2026. Version 1.0. https://theagenticdatacompany.com/open-yap-1k

BibTeX
@misc{openyap1k,
  title     = {Open Yap 1K: Channel-Separated English Natural Two-Speaker Conversations},
  author    = {The Agentic Data Company},
  year      = {2026},
  version   = {1.0},
  publisher = {The Agentic Data Company},
  url       = {https://theagenticdatacompany.com/open-yap-1k},
}

Access

This is a public release: it is not licensed exclusively, anyone may request it, and it is free for both commercial and research use. Access is granted per recipient under a data use agreement.

01

Request

Tell us who you are, what you are building, and how the audio will be used.

02

Review

We read every request. Expect a reply within a few hours, either way.

03

Delivery

Approved recipients get a portal account and either a direct download or delivery into their own S3 bucket.

Use under the agreement

Permitted
  • Commercial use, including in products you sell
  • Research use, published or internal
  • Training, fine-tuning and evaluating models, and deploying what you train
  • Internal copies, and access for staff and contractors under the same terms
  • Retaining the delivered dataset, and anything derived from it, after a speaker withdraws
Not permitted
  • Redistributing, resharing, sublicensing or reselling the dataset or any part of it
  • Attempting to identify a speaker, or link a recording to any external record
  • Creating voice clones, replicas or generative reproductions identifiable as a speaker in the corpus
  • Holding the dataset under weaker controls than your own confidential material
  • Retaining any copy of the dataset after a breach of these terms, or after a written request from The Agentic Data Company