datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
olmo-2-pretrain-validationsae-skeskinen-TinyStories-hf-validation-tokenizer-gpt2_playVQAv2_validation
Dataset Card for "VQAv2_validation"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
prolong-smollm2-validation
ProLong SmolLM2 validation
Two validation sets for evaluating DCLM-trained language models, retokenized from the code, books, and textbooks subsets of princeton-nlp/prolong-data-64K.
Folder
Context length
Sequences
Usable tokens
Stored tokens
Trailing EOS filler (unused)
4k/
4,096
24,414
99,999,744
100,000,000
256
32k/
32,768
3,051
99,975,168
100,000,000
24,832
Each folder contains prolong_val_100m.bin, per-sequence source metadata in prolong_val_100m.json, and… See the full description on the dataset page: https://huggingface.co/datasets/sycmucmu/prolong-smollm2-validation.Korea-AIHub-middlesenior-dialect-speech-validation-part2sroiv2_strawberry_picking_lab_validationThis dataset was created using LeRobot.
SROI v2 — Strawberry Picking (Lab) — Validation Set
Held-out validation set for the SROI v2 strawberry-picking data (project page, Zhejiang University): 100 human strawberry-picking demonstrations recorded with the SROI V2 handheld data-acquisition device — a UMI-style gripper with an integrated Intel RealSense D405 stereo camera — on live plants in a laboratory setup. No robot arm is involved during collection: the 7-DoF… See the full description on the dataset page: https://huggingface.co/datasets/zfff/sroiv2_strawberry_picking_lab_validation.lrs2_train_validation_testVQAv2_minival_validation_vprevious
Dataset Card for "VQA_minival_validation"
More Information needed
VQAv2_validation_no_image
Dataset Card for "VQAv2_validation_no_image"
More Information needed
ccpc-dataset-v2-validation-testVQAv2_minival_validation
Dataset Card for "VQAv2_minival_validation_v2"
More Information needed
sada-validation-wav2vec2-xls-r-300m-ar-preprocessedboltz2_quotient_validation_20260918_060803
Boltz 2 quotient-validation study
Inference-only experiments with group-indexed pair features and tiled attention, compared with the supplied Anthropic optimization kit's actual big mode on H100 GPUs.
Memory savings were measured; improved or preserved folding accuracy was not established. This is not a newly trained model. The checkpoint is the unchanged upstream Boltz 2 structure checkpoint, SHA-256 090e82ac8c92f5e943fa1b39e7410a44027bea7243c0bbb3caa67a77fc1428e1 (2,286,561… See the full description on the dataset page: https://huggingface.co/datasets/raftbioworks/boltz2_quotient_validation_20260918_060803.real-marker-d2-c00-teleop-validation
real-marker-d2-c00-teleop-validation
Part of Mulligan. Browse the release at Policy Arena.
Property
Value
Task
real-marker-d2
Role
validation-view
Release grouping
mainline
Variant
validation
Model rounds
R0
Episodes
50
Frames
9483
Recording FPS
15
Cameras
observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2
Release tag
release-2026-09-22
Parent
mulligan/real-marker-d2-c00-teleop-mixed… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-marker-d2-c00-teleop-validation.entity_extraction_ade_v2_with_validation
Dataset: Entity Extraction Adverse Drug Events with Validation Split
This dataset is a modified version of the harpreetmann/entity_extraction_ade_v2 dataset that includes a validation split.
Dataset Structure
The dataset contains:
Training set: 3458 examples
Validation set: 385 examples
Test set: 428 examples
Features
text: A string containing medical text with adverse drug events
relations: A list of dictionaries containing drug-ADE relationships… See the full description on the dataset page: https://huggingface.co/datasets/mihirhirave/entity_extraction_ade_v2_with_validation.real-square-d2-c00-teleop-validation
real-square-d2-c00-teleop-validation
Part of Mulligan. Browse the release at Policy Arena.
Property
Value
Task
real-square-d2
Role
validation-view
Release grouping
mainline
Variant
validation
Model rounds
R0
Episodes
50
Frames
7915
Recording FPS
15
Cameras
observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2
Release tag
release-2026-09-22
Parent
mulligan/real-square-d2-c00-teleop-mixed… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-square-d2-c00-teleop-validation.SNOMED-CT-NER-V.2-k-fold-validationTopiOCQA_validation_top_250_only_w_correct-v2
TopiOCQAHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
TopiOCQA (Human-in-the-loop Attributable Generative Retrieval for Information-seeking Dataset) is information-seeking conversational dataset with challenging topic switching phenomena. It consists of conversation histories along with manually labelled relevant/gold passage. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.… See the full description on the dataset page: https://huggingface.co/datasets/mteb/TopiOCQA_validation_top_250_only_w_correct-v2.real-routing-d2-c00-teleop-validation
real-routing-d2-c00-teleop-validation
Part of Mulligan. Browse the release at Policy Arena.
Property
Value
Task
real-routing-d2
Role
validation-view
Release grouping
mainline
Variant
validation
Model rounds
R0
Episodes
50
Frames
13768
Recording FPS
15
Cameras
observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2
Release tag
release-2026-09-22
Parent… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-routing-d2-c00-teleop-validation.VQAv2_sample_validation_google_flan_t5_xl_mode_A_D_PNP_GENERIC_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_A_D_PNP_GENERIC_C_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation
Dataset Card for "VQAv2_sample_validation"
More Information needed
D-SFTv1_C-cd3arg-Qwen2.5-1.5B-MockSearchV2-7_24_25-sft_test_with_validation_tracking-sft-dataVQAv2_sample_validation_google_flan_t5_xxl_mode_T_D_PNP_GENERIC_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_T_D_PNP_GENERIC_C_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xl_mode_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_Q_rices_ns_1000"
More Information needed
VQAv2_sample_validation_google_flan_t5_xl_mode_C_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xl_mode_C_Q_rices_ns_1000"
More Information needed
c4_validation
Dataset Card for "c4_validation"
More Information needed
roco2-question-dataset-validationLRS2-Validation
Usage
import cv2
import torch
import datasets
from torchcodec.decoders import AudioDecoder
from torchcodec.decoders import VideoDecoder
def load_audio(source:str|bytes, start_time:int=0, end_time:int|None=None):
audio_decoder = AudioDecoder(source)
if end_time is None:
end_time = audio_decoder.metadata.duration_seconds_from_header
waveform = audio_decoder.get_samples_played_in_range(start_time, end_time).data
return waveform.transpose(1, 0) # T x 1… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodAnaam/LRS2-Validation.VQAv2_sample_validation_google_flan_t5_xxl_mode_Q_rices_ns_1000
Dataset Card for "VQAv2_sample_validation_google_flan_t5_xxl_mode_Q_rices_ns_1000"
More Information needed
