A parallel speech dataset composed of English and Indonesian utterances aligned at the sentence level. Real-world sentences from the advertising and marketing domain were recorded by native speakers, and every recording was mapped to its transcript on a sentence-by-sentence basis to produce a high-quality aligned corpus.
| ID | Source Text (EN) | Source Audio | Target Text (ID) | Target Audio |
|---|---|---|---|---|
| 1 | I would like to share the interim results of the ongoing marketing campaign. | en_1.wav | Saya ingin membagikan hasil sementara dari kampanye pemasaran yang sedang berjalan. | id_1.wav |
| 2 | We are seeing positive signs, with website visitors increasing by about 20% compared to last week. | en_2.wav | Kami melihat tanda-tanda positif, dengan jumlah pengunjung situs web yang meningkat sekitar 20% dibandingkan minggu lalu. | id_2.wav |
| 3 | However, the final purchase conversion rate for new sign-ups is slightly lower than expected. | en_3.wav | Namun, tingkat konversi pembelian akhir dari pendaftar baru sedikit lebih rendah dari yang diharapkan. | id_3.wav |
A corpus of real-world advertising and marketing text is collected, and sentences whose average spoken length is approximately 10 seconds are selected. The final recording script is fixed after deduplication and filtering of non-standard expressions. Each utterance is recorded at least twice, and the best take is adopted.
Speech and text are aligned at the sentence level using a forced alignment algorithm, then cross-reviewed in two passes: first by native-speaker reviewers, then by expert-level annotators. Records with alignment errors, typos, or pronunciation mismatches are rejected and reworked. Only records that reach 99% alignment accuracy are included in the final dataset.