datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
markuplm-toy-datasetNatural_Questions_HTML_Toysoft-toy-wbcd-khlptoy_struc_datasettoy-models-of-sft-data
Toy Models of SFT Data
This is a public-clean candidate data package for the Toy Models of SFT project.
It is built for researcher inspection first.
The package answers two questions:
What were the models trained on?
How did the models actually behave under evaluation?
The package includes training data, eval inputs, model rollouts, judge scores,
parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and
provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.LongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.website_metadata_c4_toyA smaller version (100 samples) of https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4
retinoblastomaRetinoblastoma Dataset
This dataset contains information related to retinoblastoma from ClinvarTuring https://github.com/ToyokoLabs/ClinvarTuring
Licensing Information
License: cc-by-4.0
Authors
Morgan Lyu, Sebastian Bassi and Virginia Gonzalez
LongLive2.0-Toy-Dataset
LongLive2.0 Toy Dataset
This dataset is a toy format-checking dataset for the LongLive2.0 release
code. It is intended to help users verify AR diffusion training, DMD
distillation, and prompt formatting before preparing a larger dataset.
Dataset placeholder:
https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset
Expected Layout
The released toy dataset will contain two separate training folders:
ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Perflow-Shuai/LongLive2.0-Toy-Dataset.toy_data#toy dataset
This is a small portion of the full dataset, used for testing and formatting purposes.
rag-observatory-toy-traces
RAG Observatory Toy Traces
Three small, synthetic traces for testing RAG diagnostics and report interfaces.
Each example isolates a different outcome:
a supported answer with one irrelevant retrieved document;
a retrieval miss that sends the wrong evidence to the generator;
an answer that contradicts relevant selected context.
The records mirror the examples used by
GioiaZheng/rag-observatory
and its interactive Space.
Intended use
This dataset is suitable for:… See the full description on the dataset page: https://huggingface.co/datasets/GioiaZheng/rag-observatory-toy-traces.toy_wikitrain split:
20k documents from Wikipedia (The Pile)
valid split:
5k documents from Wikipedia (The Pile)
MORPHEUS_Datasets
MORPHEUS
MORPHEUS: Modeling Role from Personalized Dialogue History by Exploring and Utilizing Latent Space(EMNLP 2024)
Paper
EN: ConvAI2
ZH: Baidu PersonaChat
intel_orca_dpo_toyboxmy-toy-dbraft_toy_big_dataset_v2raft_toy_90_jinagpt-5.4-mini-math-depth2-toyRED6k-toyeval3_TOY_video_images
Eval3 TOY Video Images
Image-question-answer grounding dataset extracted from the Eval 3 TOY celebrity permutation videos.
Each episode contributes three frames from the first 6 seconds:
start frame
middle frame
end frame
The labels use the user-provided episode block ground truth:
episodes 0-29: Taylor Swift
episodes 30-59: Barack Obama
episodes 60-89: Yann LeCun
within each 30-episode block: first 10 left, next 10 middle, last 10 right
Files:
all.jsonl: all 270 examples… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval3_TOY_video_images.RED6k-toy-anothergoodwiki_long_toyluvorae-support-toytoy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k
ToyMaze2D Hard All-Step Future-Rollout COT
This dataset is generated from the local VisGym ToyMaze2D maze_2d/hard environment.
Rows:
train/: 500,000 gzip-compressed JSONL rows.
test/: 100 gzip-compressed JSONL rows.
Each row is a full trajectory conversation. Every user turn stores a prompt and
one JPEG image item with image_prev, image, and image_next; image_prev == image
is validated for every step, and the final step has image_next == image.
Two non-stop move steps per… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/toy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k.qwen3-8b-math-depth2-toyraft_toy_big_datasetThis is a toy dataset for raft pipleine. It contains 133 examples. The columns are : category, question, context and answer.
ruozhiba-llama3-ttmultiturn_toysetting_stage2raft_toy_130_nomicraft_toy_130_dragon
