Data Marketplace

Flitto Data Marketplace offers license-verified AI training and evaluation datasets in text, speech, image, and video across 100+ languages and 23+ domains. Try a free sample, then contact us to purchase.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

  • Pre-training DataImage

    Korean Real-World OCR Image Dataset

    A Korean real-world OCR image dataset built with OCR annotations for Korean text detection, recognition, and Document AI model development.

    DomainHumanities and Social
    LanguageKorean
  • Pre-training DataImage

    Automotive Body Defect Detection Image Dataset

    A labeled image dataset capturing surface defects on automotive body panels, including scratches, dents, and paint irregularities. Built to support vision inspection model training in automotive manufacturing processes.

    DomainScience and Engineering
    Language-
  • Pre-training DataImage

    Jailbreak Image Dataset (Red-teaming)

    A red-teaming image dataset designed for AI safety evaluation. Includes adversarial inputs, harmful content bypass attempts, and policy-violating scenarios to support model safety assessment and AI governance enhancement.

    DomainIT and Tech
    LanguageKorean|English
  • Pre-training DataAudio

    Large-Scale English Business Multi-Turn Conversational Speech Dataset

    An English multi-turn conversational speech dataset covering end-to-end business negotiations across diverse B2B domains, including SOW scope alignment, cost and timeline risk discussions, and contract revision workflows between stakeholders.

    DomainManagement, Economic and Finance|IT and Tech
    LanguageEnglish
  • Pre-training DataAudio

    Large-Scale Gulf Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Gulf Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

    DomainHumanities and Social|Lifestyle
    LanguageArabic (Saudi Arabia)|Arabic|Arabic (Kuwait)|Arabic (Qatar)|Arabic (Bahrain)|Arabic (Oman)
  • Pre-training DataAudio

    Large-Scale Egyptian Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Egyptian Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

    DomainHumanities and Social|Lifestyle
    LanguageArabic (Egypt)
  • Pre-training DataAudio

    Large-Scale Levantine Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Levantine Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

    DomainHumanities and Social|Lifestyle
    LanguageArabic (Syria)|Arabic (Lebanon)|Arabic (Jordan)|Arabic (Palestine)
  • Pre-training DataAudio

    Large-Scale Maghrebi Arabic Multi-Turn Conversational Speech Dataset for Daily Communication and Opinion Exchange

    A multi-turn conversational speech dataset in the Maghrebi Arabic dialect, covering a wide range of real-life topics such as everyday conversations, hobbies, shopping experiences, and opinion exchanges related to social media.

    DomainHumanities and Social|Lifestyle
    LanguageArabic (Morocco)|Arabic (Algeria)|Arabic (Tunisia)|Arabic (Libya)|Arabic (Mauritania)
  • Pre-training DataImage

    Multilingual Menu OCR Image Dataset

    A multilingual OCR image dataset for menu images in seven languages, including Korean, English, Chinese, and Japanese, comprising original images, translated-text images, and inpainted images with professionally reviewed annotations.

    DomainAd and Marketing|Lifestyle
    LanguageChinese (Simplified)|Vietnamese|Russian|Korean|Japanese|English|Arabic
  • Pre-training DataImage

    English Handwritten Document OCR Image Dataset

    An English handwritten document OCR image dataset with bounding-box annotations for multilingual document understanding, OCR model training, and layout-aware text recognition.

    DomainHumanities and Social
    LanguageEnglish