datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drkernel-validation-data
DR.Kernel Validation Dataset (KernelBench Level 2)
Paper | GitHub
This directory documents the format of hkust-nlp/drkernel-validation-data.
This validation set is built from KernelBench Level 2 tasks and is used for DR.Kernel evaluation/grading.
Overview
Purpose: validation/evaluation set for kernel generation models.
Task source: KernelBench Level 2.
Current local Parquet (validation_data_thinking.parquet) contains 100 tasks.
Dataset Structure
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/drkernel-validation-data.dbpedia-hindi-validation-data
DBpedia Hindi — Validation Data (Relational Triple Extraction)
3,634 real Hindi Wikipedia sentences, held out during training, used to evaluate the fine-tuned Gemma 3 4B model for the DBpedia Hindi Chapter (Google Summer of Code 2026).
Format
Same chat-format JSONL as the training dataset — messages (system/user/assistant), plus score, source, trace_type fields.
Composition
Real Hindi Wikipedia sentences only (not synthetic), each scored ≥9/10 by an… See the full description on the dataset page: https://huggingface.co/datasets/Nitin1211/dbpedia-hindi-validation-data.oct-validation-datasets-en
OCT Validation Datasets (EN)
Validation and reproducibility datasets for the Technology of Expressions (TE) and Ordinative Category Theory (OCT) framework.
Release
Version: 5.3.0
Date: 2026-05-01
Source commit: 9c13db7 (anckhalion/te-oct-framework-en)
Included
Cycle inputs and processed outputs
Raw source references used in reproducibility runs
Scripts and manifests for reconstruction
Related resources
Framework repo:… See the full description on the dataset page: https://huggingface.co/datasets/anckhalion/oct-validation-datasets-en.ophthalmology_validation_datasetcombined_ophthalmology_validation_dataset
