Data Marketplace

Use verified, licensed data with confidence. You can download right away or check the data through inquiry.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 323 datasets
  • Pre-training DataImage

    Korean Real-World OCR Image Dataset

    A Korean real-world OCR image dataset built with OCR annotations for Korean text detection, recognition, and Document AI model development.

  • Pre-training DataImage

    Automotive Body Defect Detection Image Dataset

    A labeled image dataset capturing surface defects on automotive body panels, including scratches, dents, and paint irregularities. Built to support vision inspection model training in automotive manufacturing processes.

  • Pre-training DataImage

    Jailbreak Image Dataset (Red-teaming)

    A red-teaming image dataset designed for AI safety evaluation. Includes adversarial inputs, harmful content bypass attempts, and policy-violating scenarios to support model safety assessment and AI governance enhancement.

  • Pre-training DataAudio

    Large-Scale English Business Multi-Turn Conversational Speech Dataset

    An English multi-turn conversational speech dataset covering end-to-end business negotiations across diverse B2B domains, including SOW scope alignment, cost and timeline risk discussions, and contract revision workflows between stakeholders.

  • Pre-training DataAudio

    Large-Scale Gulf Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Gulf Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

  • Pre-training DataAudio

    Large-Scale Egyptian Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Egyptian Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

  • Pre-training DataAudio

    Large-Scale Levantine Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Levantine Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

  • Pre-training DataAudio

    Large-Scale Maghrebi Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Maghrebi Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

  • Pre-training DataImage

    Multilingual Menu OCR Image Dataset

    A multilingual OCR image dataset for menu images in seven languages, including Korean, English, Chinese, and Japanese, comprising original images, translated-text images, and inpainted images with professionally reviewed annotations.

  • Pre-training DataImage

    English Handwritten Document OCR Image Dataset

    An English handwritten document OCR image dataset with bounding-box annotations for multilingual document understanding, OCR model training, and layout-aware text recognition.