CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /ai2_arc Dataset Card for "ai2_arc" Dataset Summary A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.textquestion-answering1K<n<10K402 likes868k downloads3y agoHugging Face02SakanaAI /AI-CUDA-Engineer-Archive The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.tabular10K<n<100K228 likes129k downloads2y agoHugging Face03artur-muratov /multilingual-speech-commands-15lang Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.audio1M<n<10M16 likes80k downloads1y agoHugging Face04secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B448 likes75k downloads6d agoHugging Face05picbreeder-vlm /picbreeder-vlm-archive Picbreeder-VLM Archive Every image evolved by the swarm of vision-language-model "breeders" in In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models (GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the lineage graphs, and the analysis artifacts behind the paper and blog. The original Picbreeder (Secretan et al., 2008) let crowds of humans collaboratively evolve images from CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.imageimage-to-text100K<n<1M14 likes66k downloads3mo agoHugging Face06arcinstitute /State-Parse-FilteredThe single cell RNA-seq dataset with human PBMC samples was sourced from Parse Biosciences [1]. [1] Performance of Evercode™ WT v3 in Human Immune Cells (PBMCs), https://www.parsebiosciences.com/datasets/performance-of-evercode-wt-v3-in-human-immune-cells-pbmcs/; Parse Biosciences, Seattle, USA; accessed 05/27/2025. Certain uses of this data may require a license from Parse Biosciences, Inc. textn<1K0 likes63k downloads4mo agoHugging Face07farhanhubble /jfk-archives Dataset Card for JFK Archives This dataset is a collection of all records pertaining to the assassination of the US president, John F. Kennedy, released until April 2025 through archives.org by the US government. Dataset Details Dataset Description The original data downloaded from archives.org consists of 56,300 scanned documents in PDF format, released until April 2025. The files are organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.textquestion-answering10K<n<100K0 likes42k downloads1y agoHugging Face08argilla /databricks-dolly-15k-curated-en Guidelines In this dataset, you will find a collection of records that show a category, an instruction, a context and a response to that instruction. The aim of the project is to correct the instructions, intput and responses to make sure they are of the highest quality and that they match the task category that they belong to. All three texts should be clear and include real information. In addition, the response should be as complete but concise as possible. To curate the dataset… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-en.text10K<n<100K45 likes35k downloads3y agoHugging Face09MohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes33k downloads2y agoHugging Face10argilla /ultrafeedback-binarized-preferences-cleaned UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback-binarized-preferences/README.md. Differences with argilla/ultrafeedback-binarized-preferences… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned.tabulartext-generation10K<n<100K165 likes27k downloads3y agoHugging Face11mteb /arguana ArguAna An MTEB dataset Massive Text Embedding Benchmark ArguAna: Retrieval of the Best Counterargument without Prior Topic Knowledge Task category Retrieval (text-to-text) Domains Social, Web, Written Reference ACL Source datasets: mteb/arguana How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("ArguAna") evaluator = mteb.MTEB([task]) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arguana.texttext-retrieval10K<n<100K7 likes27k downloads5mo agoHugging Face12arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face13arsaporta /symile-m3 Dataset Card for Symile-M3 Symile-M3 is a multilingual dataset of (audio, image, text) samples. The dataset is specifically designed to test a model's ability to capture higher-order information between three distinct high-dimensional data types: by incorporating multiple languages, we construct a task where text and audio are both needed to predict the image, and where, importantly, neither text nor audio alone would suffice. Paper: https://arxiv.org/abs/2411.01053 GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/arsaporta/symile-m3.audiozero-shot-classification10M<n<100M8 likes24k downloads2y agoHugging Face14argilla /distilabel-capybara-dpo-7k-binarized Capybara-DPO 7K binarized A DPO dataset built with distilabel atop the awesome LDJnr/Capybara This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models. Why? Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.tabularquestion-answering1K<n<10K184 likes23k downloads2y agoHugging Face15argilla /distilabel-intel-orca-dpo-pairs distilabel Orca Pairs for DPO The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved. Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.text10K<n<100K182 likes23k downloads1y agoHugging Face16RoganInglis /vllm-control-arena vLLM Main Tasks Dataset AI coding tasks generated from vLLM git commits Dataset Description This dataset contains 6801 coding tasks automatically generated from git commits in the vLLM repository. Each task represents a real-world coding challenge derived from actual development work. Dataset Structure The dataset contains the following columns: commit_hash: The git commit hash parent_hash: The parent commit hash commit_title: The original commit… See the full description on the dataset page: https://huggingface.co/datasets/RoganInglis/vllm-control-arena.tabulartext-generation1K<n<10K0 likes21k downloads1y agoHugging Face17aicrowd /arc-whestbench-public-2026 Organized by: Alignment Research Center (ARC), AIcrowd WhestBench 2026: ARC White-Box Estimation Challenge WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs. This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.tabularother1K<n<10K0 likes20k downloads27d agoHugging Face18scholarweave /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/scholarweave/arxiv-latex.texttext-generation1M<n<10M147 likes20k downloads8d agoHugging Face19ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M157 likes19k downloads9d agoHugging Face20MMInstruction /ArxivCap Dataset Card for ArxivCap Data Instances Example-1 of single (image, caption) pairs "......" stands for omitted parts. { 'src': 'arXiv_src_2112_060/2112.08947', 'meta': { 'meta_from_kaggle': { 'journey': '', 'license': 'http://arxiv.org/licenses/nonexclusive-distrib/1.0/', 'categories': 'cs.ET' }, 'meta_from_s2': { 'citationCount': 8… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/ArxivCap.imageimage-to-text100K<n<1M58 likes19k downloads2y agoHugging Face21Matthijs /cmu-arctic-xvectors Speaker embeddings extracted from CMU ARCTIC There is one .npy file for each utterance in the dataset, 7931 files in total. The speaker embeddings are 512-element X-vectors. The CMU ARCTIC dataset divides the utterances among the following speakers: bdl (US male) slt (US female) jmk (Canadian male) awb (Scottish male) rms (US male) clb (US female) ksp (Indian male) The X-vectors were extracted using this script, which uses the speechbrain/spkrec-xvect-voxceleb model. Usage: from… See the full description on the dataset page: https://huggingface.co/datasets/Matthijs/cmu-arctic-xvectors.texttext-to-speech1K<n<10K64 likes19k downloads4y agoHugging Face22allenai /art Dataset Card for "art" Dataset Summary ART consists of over 20k commonsense narrative contexts and 200k explanations. The Abductive Natural Language Inference Dataset from AI2. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances anli Size of downloaded dataset files: 5.12 MB Size of the generated dataset: 34.36 MB Total amount of disk used: 39.48… See the full description on the dataset page: https://huggingface.co/datasets/allenai/art.textmultiple-choice100K<n<1M9 likes18k downloads3y agoHugging Face23ArmelR /the-pile-splitted Dataset description The pile is an 800GB dataset of english text designed by EleutherAI to train large-scale language models. The original version of the dataset can be found here. The dataset is divided into 22 smaller high-quality datasets. For more information each of them, please refer to the datasheet for the pile. However, the current version of the dataset, available on the Hub, is not splitted accordingly. We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.text10M<n<100M23 likes17k downloads3y agoHugging Face24argilla /ultrafeedback-binarized-preferences-cleaned-kto UltraFeedback - Binarized using the Average of Preference Ratings (Cleaned) KTO A KTO signal transformed version of the highly loved UltraFeedback Binarized Preferences Cleaned, the preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback This dataset represents a new iteration on top of argilla/ultrafeedback-binarized-preferences, and is the recommended and preferred dataset by Argilla to use from now on when fine-tuning on UltraFeedback. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/argilla/ultrafeedback-binarized-preferences-cleaned-kto.texttext-generation100K<n<1M10 likes16k downloads3y agoHugging Face25ArtificialAnalysis /big_bench_audio Artificial Analysis Big Bench Audio Dataset Summary Big Bench Audio is an audio version of a subset of Big Bench Hard questions. The dataset can be used for evaluating the reasoning capabilities of models that support audio input. The dataset includes 1000 audio recordings for all questions from the following Big Bench Hard categories. Descriptions are taken from Suzgun et al. (2022): Formal Fallacies Syllogisms Negation (Formal Fallacies) - 250 questions Given a context… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio.audioaudio-to-audio1K<n<10K39 likes14k downloads2y agoHugging Face26oneHFR /3d-front-arimage10K<n<100K0 likes13k downloads1y agoHugging Face27maxwellinked /time-lapse-artifacts Time-Lapse Artifacts 873 indexed video files document one artist's traditional drawing practice. The recorded finish dates span September 17, 2024 through September 20, 2026; nine Pre-Standard dates remain unknown. Standardized acquisition began July 13, 2025. The current indexes contain 2,196,054,134,482 indexed video bytes (approximately 2.20 TB). The recordings began as personal practice documentation and a durable record of manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.tabular1K<n<10K3 likes12k downloads2d agoHugging Face28garrethlee /comprehensive-arithmetic-problemstext1M<n<10M0 likes11k downloads5mo agoHugging Face29argilla /apigen-function-calling Dataset card for argilla/apigen-function-calling This dataset is a merge of argilla/Synth-APIGen-v0.1 and Salesforce/xlam-function-calling-60k, making over 100K function calling examples following the APIGen recipe. Prepare for training This version is not ready to do fine tuning, but you can run a script like prepare_for_sft.py to prepare it, and run the same recipe that can be found in argilla/Llama-3.2-1B-Instruct-APIGen-FC-v0.1#training-procedure. Modify the prompt… See the full description on the dataset page: https://huggingface.co/datasets/argilla/apigen-function-calling.texttext-generation100K<n<1M19 likes11k downloads2y agoHugging Face30mteb /arena-resultsThis dataset contains the saved results from MTEB-Arena tabular1K<n<10K4 likes9k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.