CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01m-a-p /FineFineWeb-sample FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-sample.tabulartext-classification100M<n<1B4 likes57k downloads2y agoHugging Face02samyakjain /MSRBackupsReasoningCheckpointstabular10M<n<100M0 likes6.7k downloads1y agoHugging Face03olm /olm-CC-MAIN-2022-21-sampling-ratio-0.14775510204 Dataset Card for OLM May 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the May 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes4k downloads4y agoHugging Face04olm /olm-CC-MAIN-2022-27-sampling-ratio-0.16142697881 Dataset Card for OLM June/July 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the June/July 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.8k downloads4y agoHugging Face05olm /olm-CC-MAIN-2017-22-sampling-ratio-0.16178770949 Dataset Card for OLM May 2017 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the May 2017 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M0 likes2.5k downloads4y agoHugging Face06olm /olm-CC-MAIN-2022-33-sampling-ratio-0.20 Dataset Card for OLM August 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 20% of the August 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes2.5k downloads4y agoHugging Face07sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.3k downloads9mo agoHugging Face08stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes1.9k downloads22d agoHugging Face09unitedideas /practice-radar-behavioral-health-npi-sample New behavioral-health organization NPIs — weekly NPPES sample A 15-row public sample from a weekly, reproducible selection of newly enumerated Type 2 behavioral-health organizations in the U.S. Centers for Medicare & Medicaid Services National Plan and Provider Enumeration System (NPPES). Edition at a glance Measured period: July 6–12, 2026 New Type 2 organizations screened: 2,722 Behavioral-health organizations selected: 486 States and territories represented:… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/practice-radar-behavioral-health-npi-sample.tabularn<1K0 likes1.8k downloads2mo agoHugging Face10TheFinAI /dolma3_300B_sampletabular100M<n<1B0 likes1.7k downloads4mo agoHugging Face11sammyliu /qwen3-8b-activations-l20-l36 Qwen3 8B Activations for Layers 20 and 36 This dataset contains assistant-token residual activations harvested from Qwen/Qwen3-8B over 980000 training conversations from lmsys/lmsys-chat-1m. We only generated for Layer 20 and 36 because each one costs 2TB and we simply cannot afford to store more :) You can use this dataset to train SAEs, linear probes, other mech interp models etc, for Qwen3 8B. We picked Qwen3 8B because this is a small part of a larger experiment to use feature… See the full description on the dataset page: https://huggingface.co/datasets/sammyliu/qwen3-8b-activations-l20-l36.tabular100M<n<1B0 likes1.7k downloads5mo agoHugging Face12olm /olm-CC-MAIN-2022-49-sampling-ratio-olm-0.15114822547 Dataset Card for OLM November/December 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 15% of the November/December 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabulartext-generation10M<n<100M3 likes1.5k downloads4y agoHugging Face13Samsoup /DSCodeBench DSCodeBench Task-grouped, multidimensional code-generation quality estimation data derived from DSCodeBench. Dataset contents The release contains 24,972 complete artifact rows from 999 tasks. The source commit is e75ef26fedea7415bdffd3e1cbff95ddad89e7e2. Each row contains the task instruction, released 200-case test generator, generated Python code, generator identity, sandbox execution context, the independently collected 200-element correctness vector, and four… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/DSCodeBench.tabular10K<n<100K0 likes1.4k downloads2mo agoHugging Face14smartcat /Amazon_Sample_Metadata_2023 Dataset Card for Dataset Name Original datasets can be found on: https://amazon-reviews-2023.github.io/ Dataset Details This dataset was made as sample of several datasets from the link above. Dataset Description This dataset is a curated sample derived from seven filtered Amazon product category datasets(Amazon All Beauty, Amazon Fashion, Sports and Outdoors, Health and Personal Care, Amazon Clothing Shoes and Jewlery, Baby Products and Beauty and Personal… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/Amazon_Sample_Metadata_2023.tabular1M<n<10M1 likes1.4k downloads2y agoHugging Face15olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-seed-69tabular10M<n<100M1 likes1.2k downloads4y agoHugging Face16sammshen /lmcache-agentic-traces LMCache Agentic Dataset Collection A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache. Motivation Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.tabulartext-generation10K<n<100K15 likes1.2k downloads4mo agoHugging Face17blab-jhu /dclm-refinedweb-600m-sampletabular100M<n<1B0 likes1.1k downloads1mo agoHugging Face18SBMM75 /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/SBMM75/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes870 downloads10d agoHugging Face19olm /olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295 Dataset Card for OLM September/October 2022 Common Crawl Cleaned and deduplicated pretraining dataset, created with the OLM repo here from 16% of the September/October 2022 Common Crawl snapshot. Note: last_modified_timestamp was parsed from whatever a website returned in it's Last-Modified header; there are likely a small number of outliers that are incorrect, so we recommend removing the outliers before doing statistics with last_modified_timestamp. tabular10M<n<100M1 likes798 downloads4y agoHugging Face20VAST-AI /AniGen-Sample-Dataset AniGen Sample Data This directory is a compact example subset of the AniGen training dataset. What Is Included 10 examples 10 unique raw assets Full cross-modal files for each example A subset metadata.csv with 10 rows The retained directory layout follows the core structure of the reference test set: raw/ renders/ renders_cond/ skeleton/ voxels/ features/ metadata.csv statistics.txt latents/ (encoded by the trained slat auto-encoder) ss_latents/ (encoded by the… See the full description on the dataset page: https://huggingface.co/datasets/VAST-AI/AniGen-Sample-Dataset.imagen<1K1 likes725 downloads5mo agoHugging Face21b-remy /jade-samples-10000x1 JADE amortized posterior samples — 10,000 observations x 1 draw Noisy weak-lensing convergence observations paired with joint posterior draws of (convergence field, cosmology) from the amortized conditional diffusion model of JADE. [!IMPORTANT] This dataset is not a product of arXiv:2606.31988. It was generated afterwards, with the same trained model, to support posterior calibration diagnostics that do not appear in the paper. No number in the paper was computed from it, and… See the full description on the dataset page: https://huggingface.co/datasets/b-remy/jade-samples-10000x1.tabular10K<n<100K0 likes684 downloads12d agoHugging Face22snimu /fineweb-edu-sample-10BT-tiktokenizedtabular1M<n<10M0 likes649 downloads2y agoHugging Face23AINovice2005 /carbon-cpu-enriched-sequences-sampledtabular1M<n<10M0 likes649 downloads27d agoHugging Face24Pclanglais /tokenized_sampletabular1M<n<10M0 likes626 downloads2y agoHugging Face25SamuelChien821 /semikongbench-100 SemiKongBench-100 One hundred independently authored synthetic semiconductor operations tasks: ten workflows across ten frozen fabs. Each task includes 20 agent-visible assets, a task-local SQLite world, 38 provider-shaped tools, an oracle trajectory, before/after snapshots, and a deterministic 100-point verifier. The model leaderboard is intentionally empty until a complete version-pinned 100-task model run exists. Release qualification executes oracle, replay, and six… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/semikongbench-100.tabularreinforcement-learningn<1K0 likes625 downloads20d agoHugging Face26science-of-finetuning /fineweb-1m-sampletabular1M<n<10M1 likes619 downloads2y agoHugging Face27samsitol /so101_cubeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101", "total_episodes": 50, "total_frames": 29698, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsitol/so101_cube.tabularrobotics10K<n<100K0 likes612 downloads1y agoHugging Face28samuki-hf /thinking-rollouts thinking-rollouts Unconstrained rollouts from thinking (chain-of-thought) models on DS-1000 and LiveCodeBench, CoT saved verbatim alongside the final answer. Format per genlm/rollouts issue #5; schema is a superset of temperature-sweep-data. Hive-partitioned Parquet, thinking_mode folded into the model tag: rollouts/domain=<dataset>/model=<tag>/temp=<temp>/data.parquet (tags like qwen3-8b-think, qwen3-1.7b-nothink). 100 samples/instance. Columns: model, thinking_mode, temp… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/thinking-rollouts.tabulartext-generation10M<n<100M2 likes608 downloads2mo agoHugging Face29enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes564 downloads2y agoHugging Face30Samarth0710 /reviewarena ReviewArena ReviewArena accompanies the NeurIPS Evaluations & Datasets submission ReviewArena: A Large-Scale Cross-Conference Dataset and Benchmark for LLM Peer Review. This release is a large, multi-conference corpus of peer-reviewed papers + their reviews + author rebuttals + acceptance decisions, harvested from OpenReview and aligned with OCR'd full-text markdown of each paper PDF where available. 51,529 papers 196,099 reviews 558,785 OCR'd PDF pages (markdown inlined… See the full description on the dataset page: https://huggingface.co/datasets/Samarth0710/reviewarena.tabulartext-generation10K<n<100K2 likes549 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.