datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ttm-validation-datasetIn-the-wild_validation-dataset
DFE-Val: In-the-Wild Audio Deepfake Proxy Validation Set
DFE-Val is a small curated collection of 102 audio clips (51 real, 51 fake) gathered from public social media platforms to approximate the distributional characteristics of the :contentReference[oaicite:0]{index=0} benchmark.
This dataset was created as part of the :contentReference[oaicite:1]{index=1} research project and is released as an open-source contribution for the audio deepfake detection community.
Why… See the full description on the dataset page: https://huggingface.co/datasets/ALLA1N/In-the-wild_validation-dataset.OmniEdit-validation-datasetAcr_Prophage_External_Validation_Dataset
Prophage ACR Screening Pipelines
This directory contains two analysis pipelines for screening Anti-CRISPR (Acr) proteins from prophage genomes.
Directory Structure
Folder
Description
double_diff/
Double-divergence screening pipeline: identifies novel Acr candidates that are significantly different from known Acr proteins at both sequence and structure levels
double_similar/
High-confidence Acr identification pipeline: identifies high-confidence Acr candidates… See the full description on the dataset page: https://huggingface.co/datasets/Jumbol/Acr_Prophage_External_Validation_Dataset.ccpc-dataset-v2-validation-testmusic-validation-datasetArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationMNLP_M3_validation_datasetroco2-question-dataset-validationArabic-Dataset-for-Commonsense-Validationion
Dataset Card for "Arabic-Dataset-for-Commonsense-Validationion"
Paper:
Tawalbeh, Saja, and Mohammad Al-Smadi. "Is this sentence valid? an arabic dataset for commonsense validation." arXiv preprint arXiv:2008.10873 (2020).
d1_d2_d4_validation_dataset_3_episodesvalidation_countdown_sft_deepseek_qwen_distilled_32b_dataset_v2standard_dataset_nonsynthetic_sorted_validation_set
Dataset Card for "standard_dataset_nonsynthetic_sorted"
More Information needed
ttm_validation_dataset_45secvalidation-datasetophthalmology_validation_datasetvalidation_countdown_sft_deepseek_qwen_distilled_32b_datasetmedical-chatbot-validation-datasetDataset-500-validationdataset-train-validationears-reverb-dataset-validation
EARS-Reverb_v2 Dataset Card
Overview
EARS-Reverb_v2 is a large-scale dataset designed for speech enhancement and dereverberation research. It contains reverberant speech data generated as the output of the code from the ears_benchmark repository. The dataset is intended for training and validation purposes and does not include a test set.
Dataset Structure
validation/: Contains the validation data.
validation.csv: Metadata for the validation set.
There is no… See the full description on the dataset page: https://huggingface.co/datasets/Amayas/ears-reverb-dataset-validation.squad_contrasting_validation_dataset
Dataset Card for "squad_contrasting_validation_dataset"
More Information needed
Canny-image-validation-datasetsquad_adversarial_validation_dataset
Dataset Card for "squad_adversarial_validation_dataset"
More Information needed
mini-validation-datasetCOCO-validation_dataset_balanceadottm_validation_dataset_10secM3_DPO_dataset_validation_onlyNightPrompt-V1-Alpha-Dataset-validationQualified_Syntax_Reentrancy_Dataset_Validation
