A Japanese speech dataset of natural two-party dialogues collected in the everyday-life domain. Native speakers recorded free-form conversations in a quiet environment so that the natural characteristics of real dialogue — intonation, speech rate, and turn-taking timing — are preserved as they occur rather than reproduced from a script.
| ID | Audio File | Speaker 1 | Speaker 2 | Description | Duration |
|---|---|---|---|---|---|
| 1 | Audio_0002_ja_Life_Multi-turn_01 | Female | Female | 日本語で行われた自然な2人対話。幼いころの将来の夢についての日常会話。 | 00:05:55 |
| 2 | Audio_0002_ja_Life_Multi-turn_02 | Female | Female | 日本語で行われた自然な2人対話。幼いころに好きだった遊びについての日常会話。 | 00:05:40 |
Natural two-party conversations arising in everyday life are organized, and sessions with an average conversation length of approximately 5 to 7 minutes are selected. The final recording list is fixed after filtering out non-standard expressions such as profanity and unintelligible speech.
One pair of native speakers records the free-form conversation in a quiet environment at 44.1 kHz / stereo. Recording proceeds with minimal intervention so that turn-taking timing remains natural.
Speech and text are aligned at the turn level and reviewed by native-speaker reviewers. Records with alignment errors, speaker misattribution, or pronunciation mismatches are rejected and reworked, and only high-precision records are included in the final dataset.