CoolFace
Datasetpublic

u94fmn391j/SAVANT-CODALM-large

SAVANT CODALM Large Dataset This dataset is part of the SAVANT framework described in the SAVANT paper, currently under peer review. This repository is provided for peer-review purposes only. After the review process, the dataset will be made publicly available through the authors' main account. Dataset Description CODALM large is the complete 9,640-image dataset, fully labeled using the SAVANT framework with the best-performing VLM. This represents the largest… See the full description on the dataset page: https://huggingface.co/datasets/u94fmn391j/SAVANT-CODALM-large.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes22downloads
Dataset Card

SAVANT CODALM Large Dataset

This dataset is part of the SAVANT framework described in the SAVANT paper, currently under peer review.

This repository is provided for peer-review purposes only. After the review process, the dataset will be made publicly available through the authors' main account.

Dataset Description

CODALM large is the complete 9,640-image dataset, fully labeled using the SAVANT framework with the best-performing VLM. This represents the largest semantically-annotated dataset for anomaly detection in autonomous driving, and demonstrates the framework's capability as a scalable data annotation engine.

Dataset Structure

SplitSamples
Train8,676
Test964
  • —Features:
  • —image: Front-camera driving scene image
  • —description: Textual scene description
  • —classification: Binary anomaly label
  • —image_name: Image filename

Limitations

  • —Derived from the CODA dataset; domain-specific to driving scenes
  • —Annotations generated automatically by the SAVANT framework