A parallel speech dataset composed of English and Vietnamese utterances aligned at the sentence level. Real-world sentences from the advertising and marketing domain were recorded by native speakers, and every recording was mapped to its transcript on a sentence-by-sentence basis to produce a high-quality aligned corpus.
| ID | Source Text (EN) | Source Audio | Target Text (VI) | Target Audio |
|---|---|---|---|---|
| 1 | I would like to share the interim results of the ongoing marketing campaign. | en_1.wav | Tôi muốn chia sẻ kết quả sơ bộ của chiến dịch tiếp thị đang diễn ra. | vi_1.wav |
| 2 | We are seeing positive signs, with website visitors increasing by about 20% compared to last week. | en_2.wav | Chúng tôi đang thấy những tín hiệu tích cực, với lượng khách truy cập trang web tăng khoảng 20% so với tuần trước. | vi_2.wav |
| 3 | However, the final purchase conversion rate for new sign-ups is slightly lower than expected. | en_3.wav | Tuy nhiên, tỷ lệ chuyển đổi mua hàng thành công của những người đăng ký mới thấp hơn một chút so với dự kiến. | vi_3.wav |
A corpus of real-world advertising and marketing text is collected, and sentences whose average spoken length is approximately 10 seconds are selected. The final recording script is fixed after deduplication and filtering of non-standard expressions. Each utterance is recorded at least twice, and the best take is adopted.
Speech and text are aligned at the sentence level using a forced alignment algorithm, then cross-reviewed in two passes: first by native-speaker reviewers, then by expert-level annotators. Records with alignment errors, typos, or pronunciation mismatches are rejected and reworked. Only records that reach 99% alignment accuracy are included in the final dataset.