CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01reach-vb /random-imagesimagen<1K0 likes23k downloads1y agoHugging Face02research-backup /qa_squadshifts_synthetic_randomTBA 0 likes17k downloads4y agoHugging Face03sunLry /robotwin-random-47-500 RoboTwin Randomized 47 Tasks, 500 Episodes Each This dataset contains a 500-episode subset for each of 47 randomized RoboTwin tasks. Episodes 0 through 499 were selected from each task. Structure <task>-demo_randomized-1000/ episode_<n>/ episode_<n>.hdf5 instructions.json Each HDF5 episode contains robot actions, joint positions, arm metadata, and three encoded camera streams: action observations/qpos observations/left_arm_dim… See the full description on the dataset page: https://huggingface.co/datasets/sunLry/robotwin-random-47-500.robotics0 likes16k downloads3mo agoHugging Face04randing2000 /TARGO0 likes9.8k downloads3mo agoHugging Face05hbXNov /hle_math_exact_match_no_image_int_answer_random128imagen<1K0 likes6.3k downloads2y agoHugging Face06StarVLA /RoboTwin-Randomizedvideo10K<n<100K0 likes5.2k downloads8mo agoHugging Face07NurshatMenglik /Paired_Compressible_Boussinesq_Flow_Simulation_with_Random_Temperature_BCs Paired Compressible / Boussinesq Flow with Random Temperature BCs 📄 Paper: A Neural Surrogate Approach for Simulating Natural Convection Problems (arXiv:2606.25259) — Nurshat Menglik, Alex Shao, David Hyde. 10,000 matched pairs of 2D natural-convection simulations of the differentially heated square cavity under randomized wall-temperature boundary conditions. Each sample solves the same problem twice — once with the Boussinesq model and once with the fully compressible model —… See the full description on the dataset page: https://huggingface.co/datasets/NurshatMenglik/Paired_Compressible_Boussinesq_Flow_Simulation_with_Random_Temperature_BCs.image0 likes4.7k downloads1mo agoHugging Face08tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes4.7k downloads3y agoHugging Face09Stage-jh-monitor /total-300-random-jh-epoch4 total-300-random-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.440625 Valid samples: 320/320 tabularn<1K0 likes4.6k downloads16d agoHugging Face10Ravimal-Ranathunga01 /jacbuilder-eval-data0 likes4k downloads57m agoHugging Face11Ran0 /JieZi JieZi (解字) 💻 Project &nbsp;·&nbsp; 📦 Dataset &nbsp;·&nbsp; 🌐 Demo JieZi (解字) is a large-scale, expert-audited visual question answering (VQA) dataset dedicated to ancient Chinese character exegesis. It pairs high-quality character glyph images with fine-grained expert annotations across nine paleographic tasks—including headword recognition, etymology, structural analysis, glyph evolution, and component function—providing a rigorous benchmark for multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Ran0/JieZi.imagevisual-question-answering100K<n<1M2 likes3.9k downloads17d agoHugging Face12mpg-ranch /leafy_spurge Background This dataset comprises 1.3 cm resolution aerial images of grasslands in western Montana, USA, captured by a commercial drone. Many scenes contain leafy spurge (Euphorbia esula), introduced to North America, now widespread in rangeland ecosystems, which is highly invasive and damaging to crop production and biodiversity. Technicians surveyed 1000 points in the study area, noting spurge presence or absence, and recorded each point’s position with precision global… See the full description on the dataset page: https://huggingface.co/datasets/mpg-ranch/leafy_spurge.image1K<n<10K6 likes2.9k downloads1y agoHugging Face13dougalldeepmind /2026-09-11-dh-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch Delegated-harm evaluation with corrected scoring of saved rollouts field value experiment Delegated-harm evaluation with corrected scoring of saved rollouts date_generated 2026-09-11 constitution none source_repo teaching_claude_why_replication @ d627d0587a2980069b7700e72f727dae594c9f49 models {"hf_path": "dougalldeepmind/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch", "base_model": "Qwen/Qwen3.6-27B", "adapter": true… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-11-dh-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch.0 likes2.7k downloads14d agoHugging Face14erickfm /melee-ranked-replays Melee Ranked Replays Anonymized Slippi ranked replays (platinum+) from Super Smash Bros. Melee, sharded by character and rank pair. Built for behavior-cloning and other replay-driven ML work on Melee — notably MIMIC. Contents Raw .slp files grouped into tarballs by (character, rank_pair, source_archive), organized into per-character folders: {CHAR}/ {CHAR}_{rank_pair}_a{N}.tar.gz metadata/ metadata_a{N}.json Characters (25): BOWSER, CPTFALCON, DK, DOC, FALCO… See the full description on the dataset page: https://huggingface.co/datasets/erickfm/melee-ranked-replays.textreinforcement-learning1M<n<10M1 likes2.6k downloads3mo agoHugging Face15Zaid /mmlu-random-Atext10K<n<100K0 likes2.5k downloads2y agoHugging Face16rand0nmr /Wan-Syn_77x448x832_600ktext100K<n<1M0 likes2.4k downloads9mo agoHugging Face17castorini /rank_llm_datatextn<1K3 likes2.1k downloads11mo agoHugging Face18data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2k downloads3y agoHugging Face19Zaid /mmlu-random-2text10K<n<100K0 likes2k downloads2y agoHugging Face20ibm-research /argument_quality_ranking_30k Dataset Card for Argument-Quality-Ranking-30k Dataset Dataset Summary Argument Quality Ranking The dataset contains 30,497 crowd-sourced arguments for 71 debatable topics labeled for quality and stance, split into train, validation and test sets. The dataset was originally published as part of our paper: A Large-scale Dataset for Argument Quality Ranking: Construction and Analysis. Argument Topic This subset contains 9,487 of the arguments only with… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/argument_quality_ranking_30k.tabulartext-classification10K<n<100K13 likes1.8k downloads3y agoHugging Face21InstaDeepAI /genomics-long-range-benchmarkDataset for benchmark of genomic deep learning models.15 likes1.8k downloads2y agoHugging Face22Zaid /mmlu-random-1text10K<n<100K0 likes1.7k downloads2y agoHugging Face23cambridgeltl /vsr_random VSR: Visual Spatial Reasoning This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper]. Usage from datasets import load_dataset data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"} dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files) Note that the image files still need to be downloaded separately. See data/ for details. Go to our github repo for more introductions. Citation If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.imagetext-classification10K<n<100K4 likes1.6k downloads4y agoHugging Face24dougalldeepmind /2026-09-12-dh-qwen36-lora-table2-9284-nonmoral-deliberation-684-rank-64-dynbatch Delegated-harm evaluation with corrected scoring of saved rollouts field value experiment Delegated-harm evaluation with corrected scoring of saved rollouts date_generated 2026-09-12 constitution none source_repo teaching_claude_why_replication @ 7cb72cf09f15fb58839862e1ef1c6930562b3cca models {"hf_path": "dougalldeepmind/2026-09-02-qwen36-lora-table2-9284-nonmoral-deliberation-684-rank-64-dynbatch", "base_model": "Qwen/Qwen3.6-27B", "adapter": true, "mode":… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-12-dh-qwen36-lora-table2-9284-nonmoral-deliberation-684-rank-64-dynbatch.0 likes1.5k downloads13d agoHugging Face25BlidReview /steady-rans-generalization Steady-RANS cross-family generalization dataset Data for the paper "Towards generalized flow field prediction: one model across unseen object families" (under double blind review; this account is anonymous for that reason). Trained checkpoints and evaluation code are in the companion model repo: steady-rans-surrogates. Steady incompressible k-omega SST (OpenFOAM simpleFoam) external flow around 855 distinct shapes (17 scripted parametric families plus 40 ModelNet object… See the full description on the dataset page: https://huggingface.co/datasets/BlidReview/steady-rans-generalization.3d1K<n<10K0 likes1.5k downloads1mo agoHugging Face26random123123 /BrushDatatext10K<n<100K14 likes1.4k downloads2y agoHugging Face27ranWang /UN_Historical_PDF_Article_Text_Corpus python dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="train") or dataset = load_dataset("ranWang/UN_Historical_PDF_Article_Text_Corpus", split="randomTest") lang_list = ["ar", "en", "es", "fr", "ru", "zh"] for row in dataset: # 获取pdf文章内容 for lang in lang_list: # type == str lang_match_file_content = row[lang] # 如果按页分割 lang_match_file_pages_content = lang_match_file_content.split("\n----\n") text100K<n<1M2 likes1.4k downloads3y agoHugging Face28acozma /imagenet-1k-rand_blur Dataset Card for "imagenet-1k-rand_blur" More Information needed image100K<n<1M3 likes1.4k downloads3y agoHugging Face29yguooo /newyorker_caption_ranking New Yorker Caption Ranking Dataset Dataset Descriptions Homepage: https://nextml.github.io/caption-contest-data/ Repository: https://github.com/yguooo/cartoon-caption-generation Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning Point of Contact: yguo@cs.wisc.edu Dataset Summary We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.imagetext-generation1M<n<10M6 likes1.4k downloads2y agoHugging Face30Ahad690 /app-rank-anchors App Rank Anchors Community-federated public app-store calibration anchors for the AppScope open app-intelligence stack. Each row is a public fact — a segment + rank + observed download flow — derived from the public Google Play realInstalls delta over a time window paired with an app's chart rank in that window. Pooling these anchors across self-hosting contributors lets the Garg–Telang download estimator calibrate absolute scale (scale_b) per (platform, category, country)… See the full description on the dataset page: https://huggingface.co/datasets/Ahad690/app-rank-anchors.n<1K0 likes1.3k downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.