Korean Real-World OCR Image Dataset
A Korean real-world OCR image dataset built with OCR annotations for Korean text detection, recognition, and Document AI model development.
Use verified, licensed data with confidence. You can download right away or check the data through inquiry.
A Korean real-world OCR image dataset built with OCR annotations for Korean text detection, recognition, and Document AI model development.
A labeled image dataset capturing surface defects on automotive body panels, including scratches, dents, and paint irregularities. Built to support vision inspection model training in automotive manufacturing processes.
A red-teaming image dataset designed for AI safety evaluation. Includes adversarial inputs, harmful content bypass attempts, and policy-violating scenarios to support model safety assessment and AI governance enhancement.
An English multi-turn conversational speech dataset covering end-to-end business negotiations across diverse B2B domains, including SOW scope alignment, cost and timeline risk discussions, and contract revision workflows between stakeholders.
A multi-turn conversational speech dataset in the Gulf Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.
A multi-turn conversational speech dataset in the Egyptian Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.
A multi-turn conversational speech dataset in the Levantine Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.
A multi-turn conversational speech dataset in the Maghrebi Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.
A multilingual OCR image dataset for menu images in seven languages, including Korean, English, Chinese, and Japanese, comprising original images, translated-text images, and inpainted images with professionally reviewed annotations.
An English handwritten document OCR image dataset with bounding-box annotations for multilingual document understanding, OCR model training, and layout-aware text recognition.