datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tadabur-align-references
tadabur-align-references
Precomputed reference embeddings powering tadabur-align — word-level timestamp extraction for Quranic recitation via DTW alignment transfer (no ASR).
What this is
For 5,481 of the Quran's 6,236 ayahs, this dataset holds frame-level tadabur-embedding features for up to 8 reference reciters, plus each reference's word-level timestamps and internal-pause intervals. No audio is included — only model outputs and timing data. tadabur-align… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur-align-references.agents-last-exam-reference
Agents Last Exam — Reference (Ground-Truth) Data
⚠️ Gated dataset. This repo contains the ground-truth / reference outputs
used to score the Agents Last Exam (ALE) benchmark. Access requires login,
agreement to the terms on the access-request form, and manual approval.
Note (06/16/26): This repository was accidentally deleted and has been recreated. The
previous list of approved requesters could not be restored, so even if you
were granted access before, you will need to… See the full description on the dataset page: https://huggingface.co/datasets/agents-last-exam/agents-last-exam-reference.multi_reference_image_editing
Multi-Reference Instruction-Based Image Editing Dataset
Overview
This dataset contains 20,000 high-resolution image pairs and multi-modal instructions designed for training advanced image-to-image editing models. It combines two complementary example types: 10,000 reference-grounded edits, where structural or stylistic changes are driven by up to three provided visual reference images, and 10,000 occlusion-based inpainting/outpainting edits, where the model must… See the full description on the dataset page: https://huggingface.co/datasets/molbal/multi_reference_image_editing.watercolour-reference-pool
Watercolour reference pool
The reference paintings that define the reward in the watercolour RL environment: an
agent writes a p5.brush sketch, the sketch is
rendered, and a vision judge compares the render against paintings sampled from this pool.
What the pool contains is the reward function. Replace it and you have changed what
the environment rewards, without touching a line of code.
178 paintings in two tiers, each with the JavaScript source that produced it.
tier… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-reference-pool.references
GEM References
What is it?
This repository contains all the reference datasets that are used for running evaluation on the GEM benchmark. Some of these datasets were originally hosted as a GitHub release on the GEM-metrics repository, but have been migrated to the Hugging Face Hub.
Converting datasets to JSON
We provide a convert_dataset_to_json.py conversion script that converts the datasets in the GEM organisation to the JSON format expected by the… See the full description on the dataset page: https://huggingface.co/datasets/GEM/references.xet-spec-reference-filesThe files in this repository are intended to provide a reference to content processed using the xet protocol, relative to the original file: Electric_Vehicle_Population_Data_20250917.csv
The original file was exported from https://data.wa.gov/Transportation/Electric-Vehicle-Population-Data/f6w7-q2d2/about_data on September 16, 2025.
The contents are as described:
Electric_Vehicle_Population_Data_20250917.csv - the original file
Electric_Vehicle_Population_Data_20250917.csv.chunks - a… See the full description on the dataset page: https://huggingface.co/datasets/xet-team/xet-spec-reference-files.asset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.moss-character-reference-voices
MOSS character reference voices (1336 voices)
1336 distinct synthetic character voices, each mined from a cluster of generated MOSS-VA-v2 character
audio and auto-annotated by Gemini-3-Flash. For every cluster the model was shown the 3 cluster samples
their automatic voice scores, chose the single most representative sample, and wrote a full
casting-style profile.
Contents
dataset.jsonl — one row per voice: cid, name, tagline, description, age, gender, register… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-character-reference-voices.ssim-reference-videosMedical-Data-ReferenceFrontierChallenge-reference
FrontierChallenge reference data
FrontierChallenge reference data provides 97 authenticated, encrypted
verifier archives.
Path
Contents
tasks/<task-id>/verifier.fcref
encrypted tests/: grader, rubric, fixtures, validation code, and reference outputs
manifest.jsonl
archive paths, sizes, and SHA-256 commitments
source_registry.json
release binding shared with GitHub and the solve dataset
tools/
integrity checker and standalone unsealer
The archive password is… See the full description on the dataset page: https://huggingface.co/datasets/apodex/FrontierChallenge-reference.kimi-k3-full-mxfp4-kld-reference-32x2048
Kimi K3 full-MXFP4 KLD reference logits
This dataset contains the canonical full-vocabulary reference logits for
quantization comparisons of Kimi K3. The source is the original full MXFP4
checkpoint served as W4A16 on TP16 with vLLM dev/gg-k3, SparkInfer, and
InstantTensor.
Contents
32 independent 2048-token windows
65,504 scored next-token positions (32 * 2047)
vocabulary size 163,840
one [2047, 163840] F32 safetensors tensor per window
tensor key: logits
total… See the full description on the dataset page: https://huggingface.co/datasets/festr2/kimi-k3-full-mxfp4-kld-reference-32x2048.asr-reference-set-eval-temp
Temporary ASR evaluation audio
Temporary public audio files used for hosted ASR evaluation.
low-high-reference
Reference Directory
git clone https://github.com/PKU-YuanGroup/Helios.git
git clone https://github.com/NVlabs/LongLive.git
git clone git@github.com:bingreeky/MemGen.git
这个目录用于存放项目设计、实现和训练过程中会反复参考的外部资料。它不是运行时必须的源码目录,而是研究与开发参考层。
目录定位
reference/ 主要存放以下几类内容:
论文 PDF
论文配套笔记
方案草稿
外部开源项目的结构化阅读记录
和当前项目直接相关的训练/规划/critic/memory 参考材料
它的作用不是“被 import”,而是帮助回答这些问题:
当前系统应该如何拆成 planner / edit / critic / memory
哪些训练阶段适合先做监督、后做偏好、再做 RL / GRPO
哪些工作可以作为 skill / memory / reflection… See the full description on the dataset page: https://huggingface.co/datasets/Ouzhang/low-high-reference.moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.gdpval-office-round-trip-reference-filespeg_rand_05_01_cam_reference_cam0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_cam_reference_cam0.entity-references
Entity References Database
A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation.
Dataset Summary
This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.captini-scoring-referencesCeltic_Stems_Reference_Sessions_Preview
Harmonic Frontier Audio – Celtic Constellation Reference Sessions (Preview, v0.9)
A high-fidelity music-production dataset designed to connect isolated source performances, production processing, arrangement context, and finished musical outcomes.
Celtic Constellation Reference Sessions (Preview), created by Harmonic Frontier Audio, introduces the Reference Sessions product vertical through a compact proof-of-concept built around purpose-recorded Celtic ensemble material.… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Celtic_Stems_Reference_Sessions_Preview.tldr-with-sft-referenceGreek_Legal_Reference_Texts
NOMOS_Greek_Legislation
NOMOS_Greek_Legislation is a Greek-language legal text corpus derived from the National Printing House (Ethniko Typografeio - ET.gr).
The dataset focuses on Greek national legislation (Laws, Presidential Decrees, Ministerial Decisions) and was developed in the context of the +NOMOS project.
It provides full-text legal documents enriched with metadata and thematic classification tags, suitable for Legal NLP tasks such as Text Classification and Language… See the full description on the dataset page: https://huggingface.co/datasets/syn-nomos/Greek_Legal_Reference_Texts.peg_rand_05_01_tip_reference_cam0_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_tip_reference_cam0_1.RFT-reference-trajectoryspain-reference-personas-frontier
Spain Reference Personas Frontier
Spain Reference Personas Frontier is an open synthetic reference population and benchmark substrate for evaluating and designing socially grounded AI systems for Spain.
It is not observed microdata, not a survey, not a prediction of real citizens, and not a substitute for fieldwork, administrative data, or domain-specific validation.
The package is designed for simulation, evaluation, prompt conditioning, subgroup analysis, service design… See the full description on the dataset page: https://huggingface.co/datasets/apol/spain-reference-personas-frontier.peg_rand_05_01_tip_reference_cam0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_tip_reference_cam0.scancestry-reference-data
scAncestry Reference Panel
Reference data for scancestry, a tool for inferring genetic ancestry from single-cell genomics data.
Reference genome build: GRCh38 / hg38.
Contents
This dataset bundles imputation, phasing, and population-reference files used by the scAncestry pipeline:
gnomad.genomes.v3.1.2.hgdp_tgp.miss0.01.maf0.01.vcf.gz (+ .tbi) — gnomAD v3.1.2 HGDP+1000G common variants, used for PCA reference… See the full description on the dataset page: https://huggingface.co/datasets/powellgenomicslab/scancestry-reference-data.Reference-Update-Benchmark
COVER-Fish Reference-Update Benchmark
This benchmark studies how reference-corpus updates change a frozen recognition
system, instantiated on fine-grained fish identification. It publishes the
versioned control plane behind COVER-Fish: row-level manifests, gallery states,
taxonomy, tensor bindings, transition evidence, protocols and dependency locks.
The benchmark does not duplicate the 83 GB Full Payload Archive. Large source
archives and frozen tensors remain in the immutable… See the full description on the dataset page: https://huggingface.co/datasets/COVER-Fish/Reference-Update-Benchmark.test_referencespeg_rand_05_01_tip_reference_cam1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 85,
"total_frames": 45695,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:85"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_rand_05_01_tip_reference_cam1.
