General Preference Evaluation Chat Arena Multi-Turn Dataset
A Korean-based Chat Arena preference evaluation multi turn dataset built by collecting human preferences for AI model responses.
Flitto Data Marketplace offers license-verified AI training and evaluation datasets in text, speech, image, and video across 100+ languages and 23+ domains. Try a free sample, then contact us to purchase.
A Korean-based Chat Arena preference evaluation multi turn dataset built by collecting human preferences for AI model responses.
A Korean-based long context benchmark text dataset built to evaluate AI models' ability to process and reason over extended contexts.
A high-difficulty Korean reasoning benchmark dataset built according to global HLE standards, including complex expert-level reasoning problems.
A high-difficulty Korean HLE reasoning and benchmark dataset built with Korea-specific expert-level reasoning problems.
A benchmark text dataset developed for Arena evaluation based on Gulf Arabic.
A benchmark text dataset developed for Arena evaluation based on Egyptian Arabic.
A benchmark text dataset developed for Arena evaluation based on Bengali.
A benchmark text dataset developed for Arena evaluation based on Hindi.
A benchmark text dataset developed for Arena evaluation based on Indonesian.
A benchmark text dataset developed for Arena evaluation based on Japanese.