datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Griffin
Griffin: Aerial-Ground Cooperative Detection and Tracking Dataset and Benchmark
📄 arXiv |
🐙 GitHub |
💾 Baidu Netdisk |
🤗 Hugging Face
Griffin is the pioneering publicly available dataset for aerial-ground cooperative 3D perception. Built using CARLA-AirSim co-simulation, it features over 200 dynamic scenes—totaling more than 30,000 frames and 270,000 images. With instance-aware occlusion quantification, variable UAV altitudes (20–60 meters), and realistic… See the full description on the dataset page: https://huggingface.co/datasets/wjh-svm/Griffin.skin-svm
Skin Lesion SVM Dataset
This dataset contains RGB skin lesion images, binary lesion masks, and class labels for a three-class biomedical image classification task. It is designed for evaluating classical image-processing features and machine-learning models, especially the SVM baseline implemented in the companion code repository.
Code repository: https://github.com/RuiqiYang77/skin-svm
Dataset Structure
The dataset contains only two splits: train and test.
data/… See the full description on the dataset page: https://huggingface.co/datasets/RuiqiYang77/skin-svm.DataTestSVMA-dataset
SVMA Dataset
SVMA is a comprehensive benchmark of 1,009 short videos designed to evaluate the content safety of modern MLLMs. Unlike prior datasets focused on isolated modality attacks, static image-text pairs, or static audio-text pairs, SVMA introduces coordinated tri-modal adversarial prompts targeting the model's visual, auditory, and perception (cross-modal and general content reasoning) reasoning systems.
Repository: ChimeraBreak
Paper: arxiv, openaccess - coming soon
Point… See the full description on the dataset page: https://huggingface.co/datasets/smji/SVMA-dataset.SVMMTOtickets_svm_artifactssvmdlofschallenge-svmodStanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable.svm_mix_amsvm-wikipediasvm_mal_ama1svm-wikipedia-1svm_mal_bing01svm-microsoftmy-distiset-91003ec9
Dataset Card for my-distiset-91003ec9
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/svmguru/my-distiset-91003ec9/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/svmguru/my-distiset-91003ec9.finalchallenge-svmodStanford Question Answering Dataset (SQuAD) is a reading comprehension dataset, consisting of questions posed by crowdworkers on a set of Wikipedia articles, where the answer to every question is a segment of text, or span, from the corresponding reading passage, or the question might be unanswerable.svm_yahoo_1svm_amazon_01svm-json01svm-youtube-1svm_mal_app1svm_mal_apple01svm_mal_microsoft01svm-bingsvm-googlesvm_mal_applesvm_mix_apsvm_mal_ama01svm-microsoft-1svm-bing-01
