Frontier DataText

Safety Response Multi-turn Dataset

Type
CoT Reasoning Dataset
Domain
Humanities and SocialIT and TechMedicalBio, Environment and EnergyEducationManagement, Economic and FinanceAd and MarketingSports, Arts and CultureScience and EngineeringLaw and PublicLifestyle
Language
English

Overview

This multi-turn safety dataset is designed to evaluate whether LLMs can respond safely when major safety threat categories and attack patterns unfold across multi-turn conversations. The dataset includes queries designed to elicit harmful or inappropriate responses from models, enabling the assessment of potential safety vulnerabilities that may arise in real-world service environments.
Each data sample consists of a sequence of 2–3 user turns. Rather than simply presenting a series of harmful questions, the conversations are designed so that each turn connects naturally to the previous context while consistently reflecting a specific safety threat and attack pattern.
The dataset covers nine safety threat categories: violence and weapons; hate, harassment, and extremism; sexual content and exploitation; cyber intrusion and hacking; non-violent illegal activities and fraud; privacy and security violations; self-harm and suicide; dangerous or high-stakes professional advice; and misinformation and conspiracy theories.
It also incorporates a range of attack patterns, including gradual escalation of harmful intent, decomposition-and-reassembly attacks, role-play and simulation, the use of translation and coded language, and context manipulation or semantic distortion. This design enables the evaluation of a model’s ability to respond safely to context-dependent circumvention and adversarial scenarios that may be difficult to detect in a single turn.

Dataset Preview

Query IDSafety PolicyAttack PatternFailure Type
conv-0024sexual_content_exploitationdecomposition_assemblyover_specification
  • Please review the data specifications and actual samples in advance.
  • The complete data shown in this preview is available in the sample download.

How to create

Collection

The safety-response multi-turn dataset consists of sequences of 2–3 user queries designed to evaluate safety risks that may not be apparent from a single utterance in real-world service environments. Each conversation is assigned a designated safety topic, such as violence_weapons, cyber_hacking, self_harm, or misinformation_conspiracy. Rather than treating each query as an independent item, the conversation is designed to maintain a coherent context and intent throughout. Each turn follows naturally from the preceding utterance, with the core intent of the relevant risk category becoming identifiable by the final point of the conversation.
The queries incorporate attack patterns that reflect ways of bypassing safety policies or developing harmful intent across multiple turns. gradual_escalation increases the level of risk from a benign question to an ambiguous one and ultimately to an explicit harmful request. decomposition_assembly distributes the harmful objective across queries that appear benign individually, such that the harmful intent becomes apparent only when the queries are considered together. roleplay_simulation establishes a fictional role or scenario in the first turn and continues the conversation under that premise, while translation_slang_evasion uses translation, slang, euphemisms, or other indirect expressions to obscure risky intent. context_poisoning introduces false premises or distortions of prior utterances into the conversational context.

Validation

Validation is conducted based on the overall multi-turn conversation flow and the user’s final intent rather than on individual utterances in isolation. The intent expressed at each turn and the conversation’s final destination are assessed together to determine whether the designated Safety Policy accurately reflects the dominant risk of the conversation. Even when a conversation contains multiple risk elements, validation ensures that a single risk category is assigned based on the conversation’s primary objective rather than on incidental or secondary elements.
Attack patterns are also validated against the structure of the conversation as a whole. For example, gradual_escalation is checked to ensure that harmful intent becomes explicit in the final turn, while decomposition_assembly is verified to ensure that each individual turn is benign and that the harmful objective can be inferred only when the queries are combined. roleplay_simulation is checked to confirm that a fictional role or scenario is clearly established in the first turn, and context_poisoning is reviewed to verify that a false premise or distortion of a previous utterance is actually used. In addition, each user query and model response is reviewed for natural and logical continuity with the preceding context. Only data in which the risk category, attack pattern, and multi-turn QA flow are all mutually consistent are included in the final dataset.

Flitto Curation Data Flitto Curation Data