Pre-training DataAudio

English-Vietnamese Parallel Speech Dataset

Type
Single-turn Speech
Domain
Ad and Marketing
Language
EnglishVietnamese

Overview

A parallel speech dataset composed of English and Vietnamese utterances aligned at the sentence level. Real-world sentences from the advertising and marketing domain were recorded by native speakers, and every recording was mapped to its transcript on a sentence-by-sentence basis to produce a high-quality aligned corpus.

Applicable Areas

  • Training and fine-tuning Speech-to-Speech (S2S) translation models
  • Training multilingual Automatic Speech Recognition (ASR) models
  • Voice style learning for Text-to-Speech (TTS)
  • Speech–text alignment research and benchmark evaluation

Dataset Preview

IDSource Text (EN)Source AudioTarget Text (VI)Target Audio
1I would like to share the interim results of the ongoing marketing campaign.en_1.wavTôi muốn chia sẻ kết quả sơ bộ của chiến dịch tiếp thị đang diễn ra.vi_1.wav
2We are seeing positive signs, with website visitors increasing by about 20% compared to last week.en_2.wavChúng tôi đang thấy những tín hiệu tích cực, với lượng khách truy cập trang web tăng khoảng 20% so với tuần trước.vi_2.wav
3However, the final purchase conversion rate for new sign-ups is slightly lower than expected.en_3.wavTuy nhiên, tỷ lệ chuyển đổi mua hàng thành công của những người đăng ký mới thấp hơn một chút so với dự kiến.vi_3.wav
  • Audio specification: WAV, 16-bit, 48,000Hz
  • Please review the data specifications and actual samples in advance.
  • The complete data shown in this preview is available in the sample download.

How to create

Collection

A corpus of real-world advertising and marketing text is collected, and sentences whose average spoken length is approximately 10 seconds are selected. The final recording script is fixed after deduplication and filtering of non-standard expressions. Each utterance is recorded at least twice, and the best take is adopted.

Validation

Speech and text are aligned at the sentence level using a forced alignment algorithm, then cross-reviewed in two passes: first by native-speaker reviewers, then by expert-level annotators. Records with alignment errors, typos, or pronunciation mismatches are rejected and reworked. Only records that reach 99% alignment accuracy are included in the final dataset.

Flitto Curation Data Flitto Curation Data