datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olmo-2-pretrain-validationsae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playVQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
Korea-AIHub-middlesenior-dialect-speech-validation-part2VQAv2_minival_validation_vprevious
Dataset Card for "VQA_minival_validation"
More Information needed
VQAv2_validation_no_image
Dataset Card for "VQAv2_validation_no_image"
More Information needed
ccpc-dataset-v2-validation-testVQAv2_minival_validation
Dataset Card for "VQAv2_minival_validation_v2"
More Information needed
sada-validation-wav2vec2-xls-r-300m-ar-preprocessedentity_extraction_ade_v2_with_validation
Dataset: Entity Extraction Adverse Drug Events with Validation Split
This dataset is a modified version of the harpreetmann/entity_extraction_ade_v2 dataset that includes a validation split.
Dataset Structure
The dataset contains:
Training set: 3458 examples
Validation set: 385 examples
Test set: 428 examples
Features
text: A string containing medical text with adverse drug events
relations: A list of dictionaries containing drug-ADE relationships… See the full description on the dataset page: https://huggingface.co/datasets/mihirhirave/entity_extraction_ade_v2_with_validation.SNOMED-CT-NER-V.2-k-fold-validationTopiOCQA_validation_top_250_only_w_correct-v2
TopiOCQAHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
TopiOCQA (Human-in-the-loop Attributable Generative Retrieval for Information-seeking Dataset) is information-seeking conversational dataset with challenging topic switching phenomena. It consists of conversation histories along with manually labelled relevant/gold passage. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TopiOCQA_validation_top_250_only_w_correct-v2.VQAv2_sample_validation_google_flan_t5_xl_mode_A_D_PNP_GENERIC_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_A_D_PNP_GENERIC_C_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
D-SFTv1_C-cd3arg-Qwen2.5-1.5B-MockSearchV2-7_24_25-sft_test_with_validation_tracking-sft-dataVQAv2_sample_validation_google_flan_t5_xxl_mode_T_D_PNP_GENERIC_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_T_D_PNP_GENERIC_C_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xl_mode_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xl_mode_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_C_Q_rices_ns_1000"
More Information needed
c4_validation
Dataset Card for "c4_validation"
More Information needed
roco2-question-dataset-validationLRS2-Validation
Usage
import cv2
import torch
import datasets
from torchcodec.decoders import AudioDecoder
from torchcodec.decoders import VideoDecoder
def load_audio(source:str|bytes, start_time:int=0, end_time:int|None=None):
audio_decoder = AudioDecoder(source)
if end_time is None:
end_time = audio_decoder.metadata.duration_seconds_from_header
waveform = audio_decoder.get_samples_played_in_range(start_time, end_time).data
return waveform.transpose(1, 0) # T x 1… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodAnaam/LRS2-Validation.VQAv2_sample_validation_google_flan_t5_xxl_mode_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xxl_mode_A_T_CM_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_A_T_CM_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xl_mode_D_PNP_GENERIC_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_D_PNP_GENERIC_C_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xl_mode_T_D_PNP_GENERIC_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_T_D_PNP_GENERIC_C_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xxl_mode_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_C_Q_rices_ns_1000"
More Information needed
whylab-gemini-2-5-docker-validation
🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced.
DOI
This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.VQAv2_sample_validation_google_flan_t5_xxl_mode_A_C_D_PNP_GENERIC_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_A_C_D_PNP_GENERIC_Q_rices_ns_1000"
More Information needed
VQAv2_minival_validation_google_flan_t5_xxl_mode_Q_rices_ns_1000
Dataset Card for "VQAv2_minival_validation_google_flan_t5_xxl_mode_Q_rices_ns_1000"
More Information needed
