CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nielsr /markuplm-toy-datasettextn<1K2 likes260 downloads4y agoHugging Face02SaulLu /Natural_Questions_HTML_Toytextn<1K0 likes254 downloads5y agoHugging Face03kjswaroopNU /soft-toy-wbcd-khlptabularn<1K0 likes146 downloads3mo agoHugging Face04SaulLu /toy_struc_datasettabularn<1K0 likes141 downloads5y agoHugging Face05matonski /toy-models-of-sft-data Toy Models of SFT Data This is a public-clean candidate data package for the Toy Models of SFT project. It is built for researcher inspection first. The package answers two questions: What were the models trained on? How did the models actually behave under evaluation? The package includes training data, eval inputs, model rollouts, judge scores, parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.tabulartext-generation10K<n<100K0 likes141 downloads2mo agoHugging Face06Efficient-Large-Model /LongLive2.0-Toy-Dataset LongLive2.0 Toy Dataset This dataset is a toy format-checking dataset for the LongLive2.0 release code. It is intended to help users verify AR diffusion training, DMD distillation, and prompt formatting before preparing a larger dataset. Dataset placeholder: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset Expected Layout The released toy dataset will contain two separate training folders: ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2.0-Toy-Dataset.texttext-to-videon<1K0 likes113 downloads4mo agoHugging Face07shanya /website_metadata_c4_toyA smaller version (100 samples) of https://huggingface.co/datasets/bs-modeling-metadata/website_metadata_c4 textn<1K1 likes41 downloads5y agoHugging Face08Toyokolabs /retinoblastomaRetinoblastoma Dataset This dataset contains information related to retinoblastoma from ClinvarTuring https://github.com/ToyokoLabs/ClinvarTuring Licensing Information License: cc-by-4.0 Authors Morgan Lyu, Sebastian Bassi and Virginia Gonzalez textquestion-answering10K<n<100K0 likes41 downloads3y agoHugging Face09Perflow-Shuai /LongLive2.0-Toy-Dataset LongLive2.0 Toy Dataset This dataset is a toy format-checking dataset for the LongLive2.0 release code. It is intended to help users verify AR diffusion training, DMD distillation, and prompt formatting before preparing a larger dataset. Dataset placeholder: https://huggingface.co/datasets/Efficient-Large-Model/LongLive2-Toy-Dataset Expected Layout The released toy dataset will contain two separate training folders: ar_training/: paired video/caption data for AR… See the full description on the dataset page: https://huggingface.co/datasets/Perflow-Shuai/LongLive2.0-Toy-Dataset.texttext-to-videon<1K0 likes40 downloads4mo agoHugging Face10multiIR /toy_data#toy dataset This is a small portion of the full dataset, used for testing and formatting purposes. image10K<n<100K0 likes33 downloads5y agoHugging Face11GioiaZheng /rag-observatory-toy-traces RAG Observatory Toy Traces Three small, synthetic traces for testing RAG diagnostics and report interfaces. Each example isolates a different outcome: a supported answer with one irrelevant retrieved document; a retrieval miss that sends the wrong evidence to the generator; an answer that contradicts relevant selected context. The records mirror the examples used by GioiaZheng/rag-observatory and its interactive Space. Intended use This dataset is suitable for:… See the full description on the dataset page: https://huggingface.co/datasets/GioiaZheng/rag-observatory-toy-traces.textquestion-answeringn<1K0 likes25 downloads2mo agoHugging Face12yurakuratov /toy_wikitrain split: 20k documents from Wikipedia (The Pile) valid split: 5k documents from Wikipedia (The Pile) text10K<n<100K0 likes21 downloads3y agoHugging Face13Toyhom /MORPHEUS_Datasets MORPHEUS MORPHEUS: Modeling Role from Personalized Dialogue History by Exploring and Utilizing Latent Space(EMNLP 2024) Paper EN: ConvAI2 ZH: Baidu PersonaChat texttext-generation100K<n<1M0 likes21 downloads2y agoHugging Face14mayflowergmbh /intel_orca_dpo_toyboxtext10K<n<100K1 likes17 downloads3y agoHugging Face15rezashamji /my-toy-dbtextn<1K0 likes14 downloads10mo agoHugging Face16srinjoyMukherjee /raft_toy_big_dataset_v2textn<1K0 likes10 downloads2y agoHugging Face17nishitneema /raft_toy_90_jinatabularn<1K0 likes10 downloads2y agoHugging Face18teetone /gpt-5.4-mini-math-depth2-toytextn<1K0 likes10 downloads3mo agoHugging Face19JoyboyBrian /RED6k-toytextn<1K0 likes9 downloads1y agoHugging Face20robot-learning-group47 /eval3_TOY_video_images Eval3 TOY Video Images Image-question-answer grounding dataset extracted from the Eval 3 TOY celebrity permutation videos. Each episode contributes three frames from the first 6 seconds: start frame middle frame end frame The labels use the user-provided episode block ground truth: episodes 0-29: Taylor Swift episodes 30-59: Barack Obama episodes 60-89: Yann LeCun within each 30-episode block: first 10 left, next 10 middle, last 10 right Files: all.jsonl: all 270 examples… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval3_TOY_video_images.imagen<1K0 likes8 downloads4mo agoHugging Face21JoyboyBrian /RED6k-toy-anothertextn<1K0 likes7 downloads1y agoHugging Face22devrim /goodwiki_long_toytextn<1K0 likes6 downloads1y agoHugging Face23mpattanaik7 /luvorae-support-toytextn<1K0 likes6 downloads8mo agoHugging Face24novastar112 /toy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k ToyMaze2D Hard All-Step Future-Rollout COT This dataset is generated from the local VisGym ToyMaze2D maze_2d/hard environment. Rows: train/: 500,000 gzip-compressed JSONL rows. test/: 100 gzip-compressed JSONL rows. Each row is a full trajectory conversation. Every user turn stores a prompt and one JPEG image item with image_prev, image, and image_next; image_prev == image is validated for every step, and the final step has image_next == image. Two non-stop move steps per… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/toy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k.tabularimage-to-text100K<n<1M0 likes6 downloads5mo agoHugging Face25teetone /qwen3-8b-math-depth2-toytext1K<n<10K0 likes6 downloads2mo agoHugging Face26srinjoyMukherjee /raft_toy_big_datasetThis is a toy dataset for raft pipleine. It contains 133 examples. The columns are : category, question, context and answer. textn<1K0 likes5 downloads2y agoHugging Face27Toyyu /ruozhiba-llama3-tttext1K<n<10K0 likes4 downloads2y agoHugging Face28ex0pired /multiturn_toysetting_stage2textn<1K0 likes4 downloads8mo agoHugging Face29nishitneema /raft_toy_130_nomictextn<1K0 likes3 downloads2y agoHugging Face30nishitneema /raft_toy_130_dragontextn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.