CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yelp /yelp_review_full Dataset Card for YelpReviewFull Dataset Summary The Yelp reviews dataset consists of reviews from Yelp. It is extracted from the Yelp Dataset Challenge 2015 data. Supported Tasks and Leaderboards text-classification, sentiment-classification: The dataset is mainly used for text classification: given the text, predict the sentiment. Languages The reviews were mainly written in english. Dataset Structure Data Instances A… See the full description on the dataset page: https://huggingface.co/datasets/Yelp/yelp_review_full.texttext-classification100K<n<1M149 likes21k downloads3y agoHugging Face02PromptEval /PromptEval_MMLU_full MMLU Multi-Prompt Evaluation Data Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.tabularquestion-answering10M<n<100M3 likes8.2k downloads2y agoHugging Face03KMK040412 /aitw-processed-labeled-full AiTW Processed Full with App Labels This repository contains a full processed Android in the Wild (AiTW) mirror together with an app-labeled step index, official split assignment by episode_id, major-app statistics, and a ready-to-train Gmail subset. Why This Exists AiTW is large and not easy to navigate by app. The original labels contain useful fields such as goal_info, current_activity, and action coordinates, but users often need extra processing before they… See the full description on the dataset page: https://huggingface.co/datasets/KMK040412/aitw-processed-labeled-full.imageimage-text-to-text1M<n<10M1 likes8.1k downloads4mo agoHugging Face04MRSHREY197 /icrm-hitek-full-db-mixed ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/MRSHREY197/icrm-hitek-full-db-mixed.text1B<n<10B0 likes6.7k downloads20d agoHugging Face05weikaih /vsi-bench-qa-v3-hm3d-fullimage10K<n<100K0 likes4.9k downloads1y agoHugging Face06devil-69 /ICMR-HITEK-FULL-MIXED-DBgatedtext1B<n<10B0 likes4.4k downloads20h agoHugging Face07benjamin-paine /free-music-archive-full FMA: A Dataset for Music Analysis Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson. International Society for Music Information Retrieval Conference (ISMIR), 2017. We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.audioaudio-to-audio100K<n<1M20 likes3.7k downloads2y agoHugging Face08acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3.4k downloads3y agoHugging Face09wannabeyour /icrm-hitek-fulldbtext1B<n<10B0 likes2.8k downloads1mo agoHugging Face10TfqDeadlox636 /icrm-hitek-fulldbtext1B<n<10B3 likes2.7k downloads1mo agoHugging Face11yuxuanw8 /EO1H-313K-full100K<n<1M0 likes2.7k downloads1y agoHugging Face12SwayStar123 /preprocessed_DCAE-f64_1024_pd12m-fulltext1M<n<10M0 likes2.5k downloads1y agoHugging Face13DatologyAI /DatBench-Full DatBench: Discriminative, Faithful, and Efficient VLM Evaluations DatBench is a curated evaluation suite for vision–language models (VLMs) designed to be faithful, discriminative, and efficient. 📄 DatBench: Discriminative, Faithful, and Efficient VLM Evaluationshttps://arxiv.org/abs/2601.02316 Modern VLM benchmarks often overestimate model capability due to multiple-choice inflation, language-only shortcuts, annotation noise, and redundant low-signal samples. DatBench reframes… See the full description on the dataset page: https://huggingface.co/datasets/DatologyAI/DatBench-Full.image100K<n<1M77 likes2.2k downloads5mo agoHugging Face14benjamin-paine /freesound-laion-640k-commercial-16khz-full About this Repository This repository is the training split of the complete FreeSound LAION 640k dataset, limited only to licenses that permit commercial works, resampled to 16khz using torchaudio.transforms.Resample. This is ideal for use cases where a variety of audio is desired but fidelity and labels are unnecessary, such as background audio for augmenting other datasets. Dataset Versions You are looking at the full dataset which contains 403,146 unique sounds… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k-commercial-16khz-full.audioaudio-to-audio100K<n<1M2 likes2.1k downloads2y agoHugging Face15sauravsingh2111 /icrm-hitek-full-db-mixed ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/sauravsingh2111/icrm-hitek-full-db-mixed.text1B<n<10B0 likes2k downloads23d agoHugging Face16ENSEONG /full-math-private-n256-Qwen2.5-3B-Instruct-bontabular100K<n<1M0 likes1.7k downloads6mo agoHugging Face17sckeptic /icrm-hitek-full-db-mixed ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/sckeptic/icrm-hitek-full-db-mixed.text1B<n<10B0 likes1.7k downloads29d agoHugging Face18mmarone /fineweb-edu-full-metadata[WIP] FineWeb-Edu with Metadata This repo contains 3 versions of the FineWeb-Edu v1 dataset: fwedu1-metaonly/ fwedu1-text-content-zstd/ fineweb-edu-1.0.0-meta-and-text/ These are all joinable via the hash column, which is xxhash64 in pyspark, calculated on the text column. This hash is unique for all instances in the dataset. For convenience, this join is done for you in the third table fwedu1-metaonly is just the metadata of the data exactly as it comes from the FineWeb-Edu v1… See the full description on the dataset page: https://huggingface.co/datasets/mmarone/fineweb-edu-full-metadata.tabular100M<n<1B0 likes1.5k downloads1y agoHugging Face19davanstrien /bl_books_flickr_fullimage1M<n<10M1 likes1.5k downloads4y agoHugging Face20ENSEONG /full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes1.4k downloads5mo agoHugging Face21Craige113 /icrm-hitek-full-db-mixed ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/Craige113/icrm-hitek-full-db-mixed.text1B<n<10B0 likes1.2k downloads1mo agoHugging Face22ENSEONG /preprocessed-full-math-private-n256-Llama-3.2-3B-Instruct-bontabular100K<n<1M0 likes1.2k downloads5mo agoHugging Face23AhmedSaman /tadabur-lora-data-fullaudio10K<n<100K0 likes1.1k downloads29d agoHugging Face24phonix-db /phonix-full Phonix: Database for Anharmonic Phonon Interactions About Phonix is a first-principles database of anharmonic phonon interactions. Phonix v2026.03.28 targets approximately 20,000 inorganic, non-metallic, and non-magnetic compounds drawn from the Materials Project (v2022.10.28) and Phonondb (v2018-04-07), including thermal conductivity data for over 6,800 materials. The database was generated using the automated workflow software auto-kappa (GitHub), which integrates… See the full description on the dataset page: https://huggingface.co/datasets/phonix-db/phonix-full.tabular10K<n<100K3 likes1.1k downloads6mo agoHugging Face25idkdd /icrm-hitek-full-db-mixed ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet). Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent). Contents The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/idkdd/icrm-hitek-full-db-mixed.text1B<n<10B0 likes1.1k downloads24d agoHugging Face26eKaiva /fullteegeetext100M<n<1B1 likes1.1k downloads27d agoHugging Face27Koa-Chang /TissueMNIST-224-full-gpt5nano-with-vlm-features TissueMNIST 224 Full Train Val with GPT-5-nano VLM Features The full TissueMNIST train and validation splits with categorical morphology features generated by GPT-5-nano. Test is included as the full TissueMNIST passthrough split with null vlm_model_name and placeholder vlm_feature values for schema consistency. This dataset is derived from the official MedMNIST TissueMNIST 224px data. The VLM feature labels are categorical privileged-information annotations for CS231N VLM-LUPI… See the full description on the dataset page: https://huggingface.co/datasets/Koa-Chang/TissueMNIST-224-full-gpt5nano-with-vlm-features.text100K<n<1M0 likes995 downloads4mo agoHugging Face28ed-donner /items_fulltabular100K<n<1M10 likes972 downloads10mo agoHugging Face29lerobot /full_foldingThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "openarms_follower", "total_episodes": 5688, "total_frames": 14129038, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:5688" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/full_folding.tabularrobotics10M<n<100M6 likes946 downloads7mo agoHugging Face30ClaudiaRichard /mbti_classification_dataset_fullPoststabular1K<n<10K1 likes933 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.