datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Uno-Curriculum
Uno-Curriculum
Training corpus for a hierarchical-delegation router: a small language
model that decomposes a task into subtasks and routes each subtask to a
(worker model, skill) pair.
Every row comes from a real public HuggingFace dataset — the
question and gold_answer are sampled verbatim from the dataset
identified by the source field. Every row then goes through the
same three-stage pipeline (router probe → teacher trajectory →
noise removal) to obtain the multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/tinaxie/Uno-Curriculum.thinklite_curriculum_wrongonly_l2hgraph-reasoning-foundational-curriculum
Foundational Structural Curriculum — v1.1
TEMPORARY MIRROR. This repository is a convenience copy for browsing and
exploration, published so a reader can look at the data without S3/DVC
credentials. It is not the canonical dataset and is not guaranteed to
stay in sync. The source of truth is the DVC-tracked corpus in the
graph-reasoning-llm repo (pin by the git SHA of dvc.lock, never "latest").
The usual byte-for-byte verification against dvc.lock was skipped for
this upload.… See the full description on the dataset page: https://huggingface.co/datasets/akumch/graph-reasoning-foundational-curriculum.adaptive-curriculum-tool-calling-pool
Adaptive Curriculum Tool-Calling Pool — v2.1
A gated snapshot of the tool-calling training-data pool produced by the
Adaptive Curriculum for Tool Calling sub-experiment. This is the additive v2.1
revision: it keeps the entire v1 + v2 payload and adds the six per-campaign
partition manifests under manifests/partitions/. Nothing from v1 or v2 was
re-encoded, recompressed, moved, or rewritten.
Access is gated. The repository uses manual gating. You must be granted access
by the… See the full description on the dataset page: https://huggingface.co/datasets/kesava89/adaptive-curriculum-tool-calling-pool.square-curriculum-paired-fixed-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1260,
"total_frames": 345576,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1260"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-curriculum-paired-fixed-v1.africa-synth-education-curriculum-metadata-nigeria
Nigeria Education - Curriculum Metadata | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Education… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-curriculum-metadata-nigeria.RIDE-RL-Curriculum-v1
RIDE-RL Curriculum v1
Dataset Summary
RIDE-RL Curriculum v1 is a versioned training-target pool for reinforcement-learning post-training of RNA inverse-folding models. It contains 744 unique RNA structural targets selected from the RIDE training data and the RNA3DB train hierarchy.
The release was designed for the modular RIDE-RL training stack. A pretrained RIDE policy generates candidate RNA sequences conditioned on a target backbone, RhoFold+ predicts their… See the full description on the dataset page: https://huggingface.co/datasets/GuoJicz518/RIDE-RL-Curriculum-v1.digitalisierungsmanager-curriculum-azav-2026
Digitalisierungsmanager für Prozessautomatisierung und Künstliche Intelligenz: Curriculum und AZAV-Zulassung
Änderungsvermerk (19.09.2026): berichtigte Fassung
Diese Fassung ersetzt die Fassung vom 25.05.2026. Berichtigt wurden:
Module und Unterrichtseinheiten: Modultitel und UE je Modul stehen jetzt im Wortlaut der AZAV-Zulassung (13 Module, zusammen 720 UE). Die Vorfassung enthielt Titel und eine UE-Verteilung, die es in der Zulassung nicht gibt, sowie eine… See the full description on the dataset page: https://huggingface.co/datasets/SkillSprinters/digitalisierungsmanager-curriculum-azav-2026.square-curriculum-paired-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1092,
"total_frames": 239489,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1092"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-curriculum-paired-v1.square-curriculum-paired-fixed-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1368,
"total_frames": 369564,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1368"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-curriculum-paired-fixed-v2.square-curriculum-paired-v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1092,
"total_frames": 209820,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1092"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-curriculum-paired-v4.square-curriculum-paired-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1092,
"total_frames": 220620,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1092"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-curriculum-paired-v2.square-curriculum-paired-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1020,
"total_frames": 213876,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 20,
"splits": {
"train": "0:1020"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-curriculum-paired-v3.multidomain-planner-mixture-v6-curriculum-2000-glm5-trim-4-reversedsvara-indic-curriculum-tokenized
svara-indic-curriculum-tokenized
Malayalam + Hindi TTS dataset with text normalization (TN+TTS format), structured for curriculum learning.
Stats
Total: 3,430 records
Malayalam: 1,754 samples
Hindi: 1,676 samples
Curriculum Difficulty
Tier
Count
Categories
1 — Easy
1,025
Simple cardinals, clean prose
2 — Medium
1,445
Currency, units, ordinals, time
3 — Hard
960
Dates, phone numbers, mixed, complex
Format
Each record contains… See the full description on the dataset page: https://huggingface.co/datasets/sreerag/svara-indic-curriculum-tokenized.thinklite_curriculum_wrongonly_l2h_newr2t_24yoram-culture-curriculumcurriculum_embeddings
German Curriculum Concept Embeddings
Sentence embeddings for three focus concepts from the German school curriculum analysis project.
Model: paraphrase-multilingual-mpnet-base-v2Generated: 2026-05-10T09:50:17
Structure
concept/
embeddings.npy # float32 (N, 768) L2-normalised
metadata.parquet # one row per excerpt, all CSV columns
metadata_preview.json # schema + first 5 rows
Concepts
mensch — 4,852 excerpts · dim=768… See the full description on the dataset page: https://huggingface.co/datasets/deirdosh/curriculum_embeddings.nanochat-depo-l0-depth1-curriculum-20260715
Nanochat Depo-L0: symbolic
This is a diagnostic, separately versioned Depo source. Each row contains one
16-node cycle and eight queries under the depth1_only schedule. Only the eight
single-letter answers and terminal token are supervised. It is designed for a
one-document-per-sequence training protocol and must not be treated as public
Depo v3 data.
multidomain-planner-mixture-v6-curriculum-2000-glm5-trimUrdu-Munch-Curriculum-hard
Urdu Munch Curriculum Hard
Long-form and complex utterances for robust TTS training.
Split
train: 269981 rows
Curriculum Rule
audio_token_len >= 223
Notes
Filtered from zuhri025/Urdu-Munch-Processed-s-merged
for long-form / hard TTS training.
multidomain-planner-mixture-v6-curriculum-2000-glm5Urdu-Munch-Curriculum-easy
Urdu Munch Curriculum Easy
Filtered Urdu TTS training subset for easier, shorter audio sequences.
Split
train: 242549 rows
Curriculum Rule
audio_token_len < 120
Fields
id
transcript
voice
text
timestamp
duration
audio_content_token_indices
audio_global_embedding
audio_token_len
Notes
This dataset was created by filtering zuhri025/Urdu-Munch-Processed-s-merged for curriculum learning.
24yoram-department-curriculumMSPI_Features_Curriculumgpe-curriculum-v1Urdu-Munch-Curriculum-next
Urdu Munch Curriculum Next
Filtered Urdu TTS training subset for the next curriculum stage.
Split
train: 552470 rows
Curriculum Rule
120 <= audio_token_len < 223
Fields
id
transcript
voice
text
timestamp
duration
audio_content_token_indices
audio_global_embedding
audio_token_len
Notes
This dataset was created by filtering zuhri025/Urdu-Munch-Processed-s-merged for curriculum learning.
eleusis-hf-hard-curriculumIEMO_Features_Curriculumbioreason-kegg-length-curriculum-task1
