A parallel speech dataset composed of English and Korean utterances aligned at the sentence level. Real-world sentences from the advertising and marketing domain were recorded by native speakers, and every recording was mapped to its transcript on a sentence-by-sentence basis to produce a high-quality aligned corpus.
| ID | Source Text (EN) | Source Audio | Target Text (KO) | Target Audio |
|---|---|---|---|---|
| 1 | I would like to share the interim results of the ongoing marketing campaign. | en_1.wav | 현재 진행 중인 마케팅 캠페인의 중간 결과를 공유해 드리겠습니다. | ko_1.wav |
| 2 | We are seeing positive signs, with website visitors increasing by about 20% compared to last week. | en_2.wav | 지난주 대비 웹사이트 방문자 수가 약 20% 증가하며 긍정적인 신호를 보이고 있습니다. | ko_2.wav |
| 3 | However, the final purchase conversion rate for new sign-ups is slightly lower than expected. | en_3.wav | 다만, 신규 가입자의 최종 구매 전환율이 예상보다 다소 낮게 나타나고 있습니다. | ko_3.wav |
A corpus of real-world advertising and marketing text is collected, and sentences whose average spoken length is approximately 10 seconds are selected. The final recording script is fixed after deduplication and filtering of non-standard expressions. Each utterance is recorded at least twice, and the best take is adopted.
Speech and text are aligned at the sentence level using a forced alignment algorithm, then cross-reviewed in two passes: first by native-speaker reviewers, then by expert-level annotators. Records with alignment errors, typos, or pronunciation mismatches are rejected and reworked. Only records that reach 99% alignment accuracy are included in the final dataset.