Pre-training DataText

IT Articles Raw Dataset (Feature · Interview · Reporting · Review)

Type
Document Dataset
Domain
IT and Tech
Language
Korean

Overview

This dataset is a Korean raw-text corpus collected from IT trade media, including feature articles, interviews, field reporting, and product reviews written in professional editorial style.

Potential Use Cases

  • Domain Pre-Training for Korean Technology Text: Supplies dense, correctly used Korean IT terminology — product, standard, and company names in natural context — for adapting general models to technology writing.
  • Named Entity Recognition & Relation Extraction: Provides material for extracting products, companies, specifications, and release events, and the relations among them, from Korean technology reporting.
Flitto Curation Data Flitto Curation Data