CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ryanmarten /OpenThoughts-1k-sample [!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.text1K<n<10K60 likes1.3m downloads1y agoHugging Face02transferable-samplers /many-peptides-md [!IMPORTANT] Critical Update The original 8AA TICA models within subsampled_trajectories/*/8AA/*.npz employed a CA-only atom selection. These models are not valid for comparison to results in our paper. Updated files (uploaded 15/12/2025) now contain corrected models. If you previously downloaded this dataset, please re-download to ensure accurate results. Note: Codebase references to tica_features_ca must now be replaced with tica_features. This was resolved in our codebase by PR #26. Note:… See the full description on the dataset page: https://huggingface.co/datasets/transferable-samplers/many-peptides-md.11 likes884k downloads9mo agoHugging Face03malcolmrey /samplesimage1K<n<10K10 likes99k downloads7h agoHugging Face04m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes57k downloads2y agoHugging Face05axolotl-ai-co /evolkit-logprobs-pipeline-75k-v2-sampletextn<1K1 likes17k downloads2y agoHugging Face06Vchitect /VBench_sampled_videogated VBench Sampled Video 1K<n<10K4 likes15k downloads1mo agoHugging Face07icyCreater /SpreadsheetBench-2-sample SpreadsheetBench 2 Dataset SpreadsheetBench 2 evaluates spreadsheet agents on end-to-end business and financial workflows. The release contains 321 tasks in four categories: Debugging, Financial Model, Template, and Visualization. Overview Directory Task type Tasks Input workbooks Gold workbooks Debugging Spreadsheet debugging and error correction 100 100 10 Financial_Model Completion of multi-sheet financial models 100 100 20 Template Completion of… See the full description on the dataset page: https://huggingface.co/datasets/icyCreater/SpreadsheetBench-2-sample.0 likes14k downloads2mo agoHugging Face08Vchitect /VBench-2.0_sampled_videos Sample Videos of VBench-2.0 This dataset is used in the paper:👉 arXiv:2503.21755 video10K<n<100K0 likes14k downloads1y agoHugging Face09zeahub /camus-sample CAMUS Sample - 2-D Echocardiographic Ultrasound Dataset This is a sample subset of the full CAMUS dataset, provided for demonstration and testing purposes. It contains 6 files (1 patient per split). For the full dataset (500 patients), see: zeahub/camus. This dataset is a zea-format (HDF5) conversion of the CAMUS dataset for multi-structure segmentation in 2-D echocardiography. Property Value Modality 2-D transthoracic echocardiography Patients 500 Views… See the full description on the dataset page: https://huggingface.co/datasets/zeahub/camus-sample.image-segmentation0 likes13k downloads2mo agoHugging Face10hf-internal-testing /dummy-audio-samplesaudion<1K0 likes12k downloads17h agoHugging Face11sayakpaul /sample-datasetsimagen<1K1 likes11k downloads2mo agoHugging Face12ScalingIntelligence /kernelbench-samples KernelBench Samples Samples from experiments for KernelBench, described in our arxiv Learn more about KernelBench from our Paper Github Repo The samples are organized as such baseline_eval (Section 4 Baseline) repeated_sampling (Section 5.1.1 Repeated Sampling) iterative_refinement (Section 5.1.2 Iterative Refinement of Generations) Within each folder, we organize the results by /level/model/problem_{id}/sample_{id}. The inner most .json file contains the generated kernel and… See the full description on the dataset page: https://huggingface.co/datasets/ScalingIntelligence/kernelbench-samples.3 likes10k downloads2y agoHugging Face13TIACentre /TIAToolBox_Remote_Samples LICENSE No re-distribution allowed. Purpose This repository contains publicly available samples used by the TIAToolBox for testing purposes. Some of these images have been downloaded from [OpenSlide] for code verification purposes. GitHub Repository: [TIAToolBox] other0 likes9.3k downloads16d agoHugging Face14Research-EAI /essential-web-1t-sample-fdc-partitioned 🌐 Essential-Web: FDC Level-2 Partitioned Dataset 📋 Dataset Description This dataset contains a 1 trillion token sample from Essential-Web, partitioned by Free Decimal Correspondence (FDC) level-2 categories. Essential-Web is a 24-trillion-token web dataset with extensive document-level metadata designed to enable rapid dataset curation through SQL-like filtering. 🔍 Free Decimal Correspondence (FDC) The FDC taxonomy is an open classification system… See the full description on the dataset page: https://huggingface.co/datasets/Research-EAI/essential-web-1t-sample-fdc-partitioned.text100M<n<1B5 likes8.9k downloads1y agoHugging Face15moonshine-ai /audio_samples_1kaudio0 likes8.1k downloads6mo agoHugging Face16RMT-team /babilong-1k-samples BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs Preprint is on arXiv and code for LLM evaluation is available on GitHub. BABILong Leaderboard with top-performing long-context models. bAbI + Books = BABILong BABILong is a novel generative benchmark for evaluating the performance of NLP models in processing arbitrarily long documents with distributed facts. It contains 9 configs, corresponding to different sequence lengths in tokens: 0k… See the full description on the dataset page: https://huggingface.co/datasets/RMT-team/babilong-1k-samples.text10K<n<100K4 likes7.6k downloads2y agoHugging Face17paxini /Omnisharing_DB_SampleData Overview The embodied intelligence industry is currently facing significant development challenges. The most critical issue is the lack of high-quality data, particularly omnimodal data that integrates force and tactile sensing. The PaXini introduces the PX OmniSharing Dataset, built on the PaXini Super EID Factory, enabling large-scale, high-fidelity human data collection across diverse tasks and scenarios. The dataset includes multi-dimensional tactile data, multi-view visual… See the full description on the dataset page: https://huggingface.co/datasets/paxini/Omnisharing_DB_SampleData.8 likes6.5k downloads4mo agoHugging Face18agentlans /common-crawl-sample Common Crawl sample A small unofficial random subset of the famous Common Crawl dataset. 60 random segment WET files were downloaded from Common Crawl on 2024-05-12. Lines between 500 and 5000 characters long (inclusive) were kept. Only unique texts were kept. No other filtering. Languages Each text was assigned to one of the language codes using the GCLD3 Python package. The Chinese texts were classified as either simplified, traditional, or Cantonese using the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/common-crawl-sample.texttext-generation1M<n<10M8 likes5.5k downloads2y agoHugging Face19HCAI-Lab-GT /dolma3-6t-sample-100000-docs dolma3-6t-sample-100000-docs Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/). Layout HCAI-Lab/dolma3-6t-sample-100000-docs/ ├── bin_summary.csv ├── sample_contract.json └── worker_NNNN/ └── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker) Dual-access… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-100000-docs.0 likes5.1k downloads4mo agoHugging Face20hngl /swebench-verified-sample-100-Qwen3-30B-evaltext10K<n<100K0 likes4.9k downloads10mo agoHugging Face21unileon-robotics /community-benign-samples This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE. MALWARE-SAMPLES DATASET Disclaimer: This repository contains benign samples with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). This README file explains how dataset is structured, its metadata, safe use as well… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/community-benign-samples.2 likes4.5k downloads2mo agoHugging Face22unileon-robotics /community-suspicious-samples This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE. MALWARE-SAMPLES DATASET Disclaimer: This repository may contain real samples of malware that can be executed (.exe) and artifacts related with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). DO NOT execute any of… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/community-suspicious-samples.image1K<n<10K0 likes4.3k downloads2mo agoHugging Face23asenion-ai /sampled-local-resumes sampled-local-resumes This dataset contains synthetic resume data sampled from local folders (20% sample from each folder). License This dataset is released under the Apache License 2.0. Please see the LICENSE and NOTICE files for details. Attribution Copyright 2025 Fairly AI Inc. dba Asenion This dataset includes data released by Fairly AI Inc. dba Asenion under the Apache License, Version 2.0. You may obtain a copy of the License at:… See the full description on the dataset page: https://huggingface.co/datasets/asenion-ai/sampled-local-resumes.text-generation1K<n<10K0 likes4.3k downloads1y agoHugging Face24EleutherAI /rpj-v2-sampleThis is a mirror of the sample-10B subset of RedPajama-Data-V2 which we have re-uploaded in order to resolve issues with the original download script. Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rpj-v2-sample.texttext-generation1M<n<10M2 likes4k downloads2y agoHugging Face25eustlb /audio-samplesaudion<1K0 likes3.7k downloads9mo agoHugging Face26stanford-cs336 /owt-sampleThese files were created with the following script: from datasets import load_dataset from tqdm import tqdm import io dataset = load_dataset("Skylion007/openwebtext")['train'] split_dataset = dataset.train_test_split(train_size=2400000, test_size=60000, seed=0) with io.open('data/owt_train.txt','w') as fopen: listout = [] for data in tqdm(split_dataset['train']): listout.append(data['text']+'<|endoftext|>') if len(listout) > 1000: _ =… See the full description on the dataset page: https://huggingface.co/datasets/stanford-cs336/owt-sample.text10M<n<100M7 likes3.4k downloads2y agoHugging Face27RoboSynChallenge /cobotmagic_Sim_sample_loading0 likes3.4k downloads2mo agoHugging Face28stablellama /Qwen-Image-2512_samplesThis dataset is a highly diverse set of high quality images generated with Qwen Image 2512. Possible uses Regularization images for training models based on Qwen Image 2512 Quality testing Data source The images were created in ComfyUI with the bf16 version of Qwen Image 2512. For each prompt were four images generated, all are (without any cherry picking) included in the corresponding dataset directories. bf16 - full model weights 1328x1328 pixels - native resolution… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Qwen-Image-2512_samples.texttext-to-image1K<n<10K3 likes3.1k downloads8mo agoHugging Face29hf-internal-testing /cats_vs_dogs_sampleimagen<1K1 likes3k downloads1y agoHugging Face30bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes3k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.