Pre-training DataVideo

AI-Generated Video with Frame-level Caption Dataset

Type
Video
Domain
Lifestyle
Language
Korean

Overview

A video understanding dataset built on generated footage for home and security solution evaluation, annotated far beyond frame captions. Scenarios are constructed across indoor, outdoor, and doorbell viewpoints and varied along object, illumination, location, and target-behavior axes, so the collection covers combinations that are difficult to capture in real footage.

Each clip carries four annotation layers in a single file: a video-level caption with structured scene attributes, a detailed breakdown of objects, spatial and temporal relations, and on-screen text, a 1 fps frame-by-frame caption track, and two evaluation layers — typed question-answer pairs with timestamp references, and fact-check statements labeled true, false, or neutral. The fact-check layer is deliberately balanced so that a model cannot score well by agreeing with everything.

Applicable Areas

  • Training and evaluating video captioning and video question-answering models
  • Temporal grounding and action recognition research
  • Hallucination detection and visual entailment via the labeled fact-check layer
  • Augmenting scarce scenario coverage for home and security vision models

Dataset Preview

Block / fieldValue
video_metaid PW-DB-KR-043 · domain indoor_delivery · duration_sec 15.042 · resolution 1920×1080 · fps 24.0 · total_frames 361 · codec h264 · bit_rate_kbps 2449 ...
video_caption.summaryA delivery person carries and places packages near an apartment door inside a dimly lit hallway at night.
video_caption attributesscene_type indoor_lifestyle · camera_motion static · focus_depth wide · dominant_color dim_beige_and_dark · action delivering_packages · subject deliv...
video_details.objects[]apartment_door ×1 — attributes [metallic, grayish color, numbered 202, with door handle and peephole], location "center of the scene, rear wall". Also...
video_details.temporal_relations[]packages stacked near door before delivery person arrives → delivery person carries box descending stairs → delivery person places box next to package...
video_details.text_detectionhas_text true · readable_text ["202"] · text_role "apartment door number"
frames[] — 15 entries at 1 fps0.0s Three packages are stacked near an apartment door labeled 202 in a dim hallway at night. · 6.0s The delivery person bends down to place a box nex...
qa_pairs[] — 7, one per typedescriptive / temporal / grounding / causal / comparative / counting / attribute. e.g. temporal — Q "When does the delivery person begin descending th...
fact_checks[] — 9, balanced 3 true / 3 false / 3 neutralfalse — "There are exactly three packages visible throughout the video." reason "Additional packages are brought and placed next to the initial three ...
  • Please review the data specifications and actual samples in advance.
  • The complete data shown in this preview is available in the sample download.

How to create

Collection

Scenario specifications are defined first as combinations of viewpoint, object, illumination, location, and target behavior, and generation prompts are written to hit each cell of that matrix. Clips are produced at a fixed length of roughly 15 seconds so that the 1 fps frame track is uniform in size across the set, and each clip is assigned a domain and a viewpoint code at generation time so the scenario matrix stays auditable.

Generations with visible artifacts, inconsistent object identity across frames, or behavior that does not match the requested scenario are rejected before annotation, since a clip whose content drifts cannot carry a coherent temporal relation set.

Validation

Annotation proceeds layer by layer, and each layer is checked against the frames rather than against the layer above it. Frame captions describe only what is visible in that second, and a reviewer verifies caption-to-frame correspondence and consistency of object and person references across consecutive frames. Video-level summaries, dense captions, and relation lists are then checked for agreement with the frame track, so that no video-level claim survives without frame evidence.

The evaluation layers are reviewed against a fixed shape. Each clip must carry one question of each of the seven types, with timestamp_ref present wherever the answer depends on a specific moment. Fact-check statements must divide evenly into three supported, three contradicted, and three underdetermined, and every non-supported label must carry a written reason grounded in what the frames do or do not show — neutral is reserved for statements the footage genuinely cannot settle, not for cases the annotator found difficult.

Flitto Curation Data Flitto Curation Data