This dataset is a Korean raw-text corpus collected from game media, including publisher press releases, news updates, event coverage, and developer interviews.
Potential Use Cases
- Domain Pre-Training for Korean Gaming Text: Supplies the title, studio, genre, and live-service vocabulary that general Korean corpora underrepresent, adapting models to the gaming domain.
- Announcement vs. Reporting Discrimination: The presence of both promotional and editorial text on the same subjects supports training classifiers that distinguish announcement framing from independent assessment.