Tool calling
Scripted voice-agent calls, with mock tools fired at realistic latencies
Sample audio
Data description
Scripted voice-agent calls: one human performs as agent while a second human performs as the caller. Agent fires mock tools during the call with realistic simulated latencies.
Each speaker’s microphone is captured locally as uncompressed 16-bit PCM WAV - never re-encoded from the call audio - and delivered as 24 kHz mono per-speaker stems (FLAC or WAV). Every take ships with the full scene script including that take’s rolled variable values, word-timed per-speaker transcripts keyed agent / caller, and a time-aligned events.jsonl logging every tool_call and tool_result against the delivered audio.
Each script is performed by multiple distinct speaker pairs, never the same speaker twice. Built for training and evaluating tool-calling voice agents: latency-bridging speech, turn-taking around function calls, read-backs and confirmations, and agent-persona TTS.
| Detail | Value |
|---|---|
| Speakers per recording | 2 |
| Channels | Dual (one file per speaker) |
| Audio format | .flac / .wav |
| Sample rate | 24 kHz |
| Bit depth | 16-bit PCM |
| Languages | English (en) |
| Tool calls per conversation | 1to6 |
| Metadata format | .json / .jsonlUTF-8 |
| Transcription type | ASR, word-level with timings(not human-verified) |
| Transcription engine | Deepgram Nova-3 |
