Harindham/Video_Detections
Each row in the dataset corresponds to one detected object instance in a single frame, along with spatial and temporal metadata. Detection & Temporal Indexing The video was sampled at a constant rate (5 seconds per frame). Each sampled frame is assigned: A frame index (sequential order) A timestamp in seconds indicating its position in the video. YOLO was applied to each frame to produce: Bounding box coordinates Component class labels Confidence scores This temporal… See the full description on the dataset page: https://huggingface.co/datasets/Harindham/Video_Detections.
04
1 2Each row in the dataset corresponds to one detected object instance in a single frame, along with spatial and temporal metadata.3 4# Detection & Temporal Indexing5 6The video was sampled at a constant rate (5 seconds per frame).7 8Each sampled frame is assigned:9 10 - A frame index (sequential order)11 12 - A timestamp in seconds indicating its position in the video.13 14YOLO was applied to each frame to produce:15 16 - Bounding box coordinates17 18 - Component class labels19 20 - Confidence scores21 22This temporal indexing allows detections to be grouped into continuous time intervals, making it possible to determine when specific vehicle components appear in the video and retrieve all occurrences over time.23 24### Schema25 26Each detection record includes:27 28- video_id – identifier of the source video29 30- frame_index – index of the sampled frame31 32- timestamp – time position in seconds33 34- class_label – detected vehicle component35 36- x_min, y_min, x_max, y_max – bounding box coordinates37 38- confidence_score – detection confidence from the model39 40Multiple rows may exist for a single frame if multiple objects are detected.