CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ryanmarten /OpenThoughts-1k-sample [!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.text1K<n<10K60 likes1.3m downloads1y agoHugging Face02transferable-samplers /many-peptides-md [!IMPORTANT] Critical Update The original 8AA TICA models within subsampled_trajectories/*/8AA/*.npz employed a CA-only atom selection. These models are not valid for comparison to results in our paper. Updated files (uploaded 15/12/2025) now contain corrected models. If you previously downloaded this dataset, please re-download to ensure accurate results. Note: Codebase references to tica_features_ca must now be replaced with tica_features. This was resolved in our codebase by PR #26. Note:… See the full description on the dataset page: https://huggingface.co/datasets/transferable-samplers/many-peptides-md.11 likes884k downloads9mo agoHugging Face03malcolmrey /samplesimage1K<n<10K10 likes99k downloads5h agoHugging Face04Samsoup /cosmos_qatext10K<n<100K0 likes72k downloads2y agoHugging Face05samilod9 /vaz0 likes70k downloads28d agoHugging Face06m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes57k downloads2y agoHugging Face07samilod9 /nmnj0 likes21k downloads28d agoHugging Face08axolotl-ai-co /evolkit-logprobs-pipeline-75k-v2-sampletextn<1K1 likes17k downloads2y agoHugging Face09Vchitect /VBench_sampled_videogated VBench Sampled Video 1K<n<10K4 likes15k downloads29d agoHugging Face10icyCreater /SpreadsheetBench-2-sample SpreadsheetBench 2 Dataset SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business and financial workflows. The release contains 321 tasks in four categories: Debugging, Financial Model, Template, and Visualization. Overview Directory Task type Tasks Input workbooks Gold workbooks Debugging Spreadsheet debugging and error correction 100 100 10 Financial_Model Completion of multi-sheet financial models 100 100 20 Template Completion of… See the full description on the dataset page: https://huggingface.co/datasets/icyCreater/SpreadsheetBench-2-sample.0 likes14k downloads2mo agoHugging Face11Vchitect /VBench-2.0_sampled_videos Sample Videos of VBench-2.0 This dataset is used in the paper:👉 arXiv:2503.21755 video10K<n<100K0 likes14k downloads1y agoHugging Face12zeahub /camus-sample CAMUS Sample - 2-D Echocardiographic Ultrasound Dataset This is a sample subset of the full CAMUS dataset, provided for demonstration and testing purposes. It contains 6 files (1 patient per split). For the full dataset (500 patients), see: zeahub/camus. This dataset is a zea-format (HDF5) conversion of the CAMUS dataset for multi-structure segmentation in 2-D echocardiography. Property Value Modality 2-D transthoracic echocardiography Patients 500 Views… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus-sample.image-segmentation0 likes13k downloads2mo agoHugging Face13hf-internal-testing /dummy-audio-samplesaudion<1K0 likes12k downloads15h agoHugging Face14blueyo0 /SA-Med3D-140K SA-Med3D-140K [github] Dataset Summary SA-Med3D-140K is a large-scale, multi-modal, multi-anatomical volumetric medical image segmentation dataset. It was created to facilitate the development of general-purpose foundation models for 3D medical image segmentation. The dataset comprises 21,729 3D medical images and 143,518 corresponding masks. It was gathered from a combination of 70 public datasets and 8,128 privately licensed annotated cases from 24 hospitals.… See the full description on the dataset page: https://huggingface.co/datasets/blueyo0/SA-Med3D-140K.11 likes12k downloads1y agoHugging Face15sayakpaul /sample-datasetsimagen<1K1 likes11k downloads2mo agoHugging Face16ScalingIntelligence /kernelbench-samples KernelBench Samples Samples from experiments for KernelBench, described in our arxiv Learn more about KernelBench from our Paper Github Repo The samples are organized as such baseline_eval (Section 4 Baseline) repeated_sampling (Section 5.1.1 Repeated Sampling) iterative_refinement (Section 5.1.2 Iterative Refinement of Generations) Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}. The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.3 likes10k downloads2y agoHugging Face17TIACentre /TIAToolBox_Remote_Samples LICENSE No re-distribution allowed. Purpose This repository contains publicly available samples used by the TIAToolBox for testing purposes. Some of these images have been downloaded from [OpenSlide] for code verification purposes. GitHub Repository: [TIAToolBox] other0 likes9.3k downloads16d agoHugging Face18Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8.9k downloads1y agoHugging Face19moonshine-ai /audio_samples_1kaudio0 likes8.1k downloads6mo agoHugging Face20samuelt7 /csf0 likes8k downloads3mo agoHugging Face21Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes7.7k downloads2mo agoHugging Face22RMT-team /babilong-1k-samples BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs Preprint is on arXiv and code for LLM evaluation is available on GitHub. BABILong Leaderboard with top-performing long-context models. bAbI + Books = BABILong BABILong is a novel generative benchmark for evaluating the performance of NLP models in processing arbitrarily long documents with distributed facts. It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.text10K<n<100K4 likes7.6k downloads2y agoHugging Face23samyakjain /MSRBackupsReasoningCheckpointstabular10M<n<100M0 likes6.7k downloads1y agoHugging Face24paxini /Omnisharing_DB_SampleData Overview The embodied intelligence industry is currently facing significant development challenges. The most critical issue is the lack of high-quality data, particularly omnimodal data that integrates force and tactile sensing. The PaXini introduces the PX OmniSharing Dataset, built on the PaXini Super EID Factory, enabling large-scale, high-fidelity human data collection across diverse tasks and scenarios. The dataset includes multi-dimensional tactile data, multi-view visual… See the full description on the dataset page: https://huggingface.co/datasets/paxini/Omnisharing_DB_SampleData.8 likes6.5k downloads4mo agoHugging Face25samuelt7 /lsd0 likes6k downloads1y agoHugging Face26sam-guided-vlas /eval-resultsvideo100K<n<1M0 likes6k downloads9d agoHugging Face27Sam04 /au30_tra Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Sam04/au30_tra.0 likes5.9k downloads7mo agoHugging Face28knkarthick /samsum Dataset Card for SAMSum Corpus Dataset Description Links Homepage: hhttps://arxiv.org/abs/1911.12237v2 Repository: https://arxiv.org/abs/1911.12237v2 Paper: https://arxiv.org/abs/1911.12237v2 Point of Contact: https://huggingface.co/knkarthick Dataset Summary The SAMSum dataset contains about 16k messenger-like conversations with summaries. Conversations were created and written down by linguists fluent in English. Linguists were asked to… See the full description on the dataset page: https://huggingface.co/datasets/knkarthick/samsum.textsummarization10K<n<100K44 likes5.7k downloads1y agoHugging Face29agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.5k downloads2y agoHugging Face30P2SAMAPA /p2-etf-samba-models0 likes5.4k downloads17h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.