datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.ttm-validation-dataseteskulap_validation_datasetsccpc-dataset-v2-validation-testOmniEdit-validation-datasetdrkernel-validation-data
DR.Kernel Validation Dataset (KernelBench Level 2)
Paper | GitHub
This directory documents the format of hkust-nlp/drkernel-validation-data.
This validation set is built from KernelBench Level 2 tasks and is used for DR.Kernel evaluation/grading.
Overview
Purpose: validation/evaluation set for kernel generation models.
Task source: KernelBench Level 2.
Current local Parquet (validation_data_thinking.parquet) contains 100 tasks.
Dataset Structure
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/drkernel-validation-data.Acr_Prophage_External_Validation_Dataset
Prophage ACR Screening Pipelines
This directory contains two analysis pipelines for screening Anti-CRISPR (Acr) proteins from prophage genomes.
Directory Structure
Folder
Description
double_diff/
Double-divergence screening pipeline: identifies novel Acr candidates that are significantly different from known Acr proteins at both sequence and structure levels
double_similar/
High-confidence Acr identification pipeline: identifies high-confidence Acr candidates… See the full description on the dataset page: https://huggingface.co/datasets/Jumbol/Acr_Prophage_External_Validation_Dataset.apriltag-validation-data
Official README
See the README.txt
Dataset Sources
A non-sharepoint hosting of the "ICRA 2020 - Determining and Improving the Localization Accuracy of AprilTag Detection" dataset
All rights belong to the original authors.
Repository: https://rzunibw-my.sharepoint.com/personal/thorsten_luettel_rzunibw_onmicrosoft_com/_layouts/15/onedrive.aspx?id=%2Fpersonal%2Fthorsten%5Fluettel%5Frzunibw%5Fonmicrosoft%5Fcom%2FDocuments%2Fdatasets%2Ficra2020%2Dapriltag%2Ddataset&ga=1… See the full description on the dataset page: https://huggingface.co/datasets/NoeFontana/apriltag-validation-data.graphql-schema-resolver-execution-validation-cmskdigf
GraphQL Schema Resolver Execution Validation
Each item is a graphql schema resolver execution validation example providing Schema (SDL), Resolver map, Sample query, Expected response (JSON). Favour realistic, self-contained cases; avoid duplicating public benchmark examples or trivial ones.
About
This dataset was produced by the DataBounty community and published here as part of an open, karma-only program.
Accepted items: 1000
Language: JavaScript
Framework:… See the full description on the dataset page: https://huggingface.co/datasets/databounty-io/graphql-schema-resolver-execution-validation-cmskdigf.MMMU-Reasoning-Distill-ValidationD-SFTv1_C-cd3arg-Qwen2.5-1.5B-MockSearchV2-7_24_25-sft_test_with_validation_tracking-sft-databangla-ocr-validation_data_printed
Bangla OCR Validation Dataset (Printed + Scanned)
📌 Description
This dataset is a Bangla OCR validation dataset containing a mix of printed document images and their corresponding text annotations. It is designed to evaluate OCR and vision-language models on both clean digital text and scanned document images.
📊 Dataset Composition
1507 line-level images with text annotations
50 full-page document images with text
Data includes:
Printed/typed Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/arobin79/bangla-ocr-validation_data_printed.roco2-question-dataset-validationmusic-validation-datasetBuild_Bench_Validation_DataThis repository contains the validation set of BuildBench paper. It contains 70 data samples.
MNLP_M3_validation_datasetArabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationonly-code-validation-datavalidation-datastandard_dataset_nonsynthetic_sorted_validation_set
Dataset Card for "standard_dataset_nonsynthetic_sorted"
More Information needed
nepali-validation_datavalidation_countdown_sft_deepseek_qwen_distilled_32b_dataset_v2zelda_validation_data
Zelda Validation Data
Small validation bundle for the FastVideo Matrix-Game 2.0 Zelda world-model
training scenarios.
Layout
validation_zelda.json
images/
actions/
validation_zelda.json uses paths relative to this directory. Download the
dataset into the FastVideo repository root as:
python scripts/huggingface/download_hf.py \
--repo_id mignonjia/zelda_validation_data \
--local_dir data/zelda_validation_data \
--repo_type dataset
ttm_validation_dataset_45secvalidation_countdown_sft_deepseek_qwen_distilled_32b_datasetdbpedia-hindi-validation-data
DBpedia Hindi — Validation Data (Relational Triple Extraction)
3,634 real Hindi Wikipedia sentences, held out during training, used to evaluate the fine-tuned Gemma 3 4B model for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Format
Same chat-format JSONL as the training dataset — messages (system/user/assistant), plus score, source, trace_type fields.
Composition
Real Hindi Wikipedia sentences only (not synthetic), each scored ≥9/10 by an… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-validation-data.text-and-code-validation-dataoct-validation-datasets-en
OCT Validation Datasets (EN)
Validation and reproducibility datasets for the Technology of Expressions (TE) and Ordinative Category Theory (OCT) framework.
Release
Version: 5.3.0
Date: 2026-05-01
Source commit: 9c13db7 (anckhalion/te-oct-framework-en)
Included
Cycle inputs and processed outputs
Raw source references used in reproducibility runs
Scripts and manifests for reconstruction
Related resources
Framework repo:… See the full description on the dataset page: https://huggingface.co/datasets/anckhalion/oct-validation-datasets-en.Dataset-500-validationophthalmology_validation_dataset
