Pre-training DataAudio

Japanese Natural Dialogue Speech Dataset

Type
Multi-turn Speech
Domain
Lifestyle
Language
Japanese

Overview

A Japanese speech dataset of natural two-party dialogues collected in the everyday-life domain. Native speakers recorded free-form conversations in a quiet environment so that the natural characteristics of real dialogue — intonation, speech rate, and turn-taking timing — are preserved as they occur rather than reproduced from a script.

Applicable Areas

  • Training Automatic Speech Recognition (ASR) and Speaker Diarization models
  • Training voice interfaces for conversational AI and chatbots
  • Learning natural speaking styles for Text-to-Speech (TTS)
  • Research on multi-turn dialogue flow and intonation

Dataset Preview

IDAudio FileSpeaker 1Speaker 2DescriptionDuration
1Audio_0002_ja_Life_Multi-turn_01FemaleFemale日本語で行われた自然な2人対話。幼いころの将来の夢についての日常会話。00:05:55
2Audio_0002_ja_Life_Multi-turn_02FemaleFemale日本語で行われた自然な2人対話。幼いころに好きだった遊びについての日常会話。00:05:40
  • Example conversation topics: childhood career aspirations, favorite childhood games, and other everyday subjects
  • Session length: approximately 5 min 40 sec to 5 min 55 sec in the sample (recorded as long-form, session-level conversations)
  • Please review the data specifications and actual samples in advance.
  • The complete data shown in this preview is available in the sample download.

How to create

Collection

Natural two-party conversations arising in everyday life are organized, and sessions with an average conversation length of approximately 5 to 7 minutes are selected. The final recording list is fixed after filtering out non-standard expressions such as profanity and unintelligible speech.
One pair of native speakers records the free-form conversation in a quiet environment at 44.1 kHz / stereo. Recording proceeds with minimal intervention so that turn-taking timing remains natural.

Validation

Speech and text are aligned at the turn level and reviewed by native-speaker reviewers. Records with alignment errors, speaker misattribution, or pronunciation mismatches are rejected and reworked, and only high-precision records are included in the final dataset.

Flitto Curation Data Flitto Curation Data