CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tau /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.textquestion-answering10K<n<100K155 likes301k downloads3y agoHugging Face02extraordinarylab /commonsense-qatext10K<n<100K0 likes21k downloads11mo agoHugging Face03coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.3k downloads8mo agoHugging Face04CUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes3k downloads9mo agoHugging Face05PleIAs /Medical-Commons Medical-Commons Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias. It includes three different collection: International scientific collection of 2M articles from OpenAlex. French scientific collection of XM articles, reports and PhD theses from French institutional repositories. Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.tabular1M<n<10M2 likes2.3k downloads2y agoHugging Face06trace-commons /agent-traces Trace Commons — Agent Traces Trace Commons is one open, public dataset of coding-agent sessions — the back-and-forth between a developer and an AI coding agent, including prompts, model responses, tool calls, and command output — contributed voluntarily as an open resource for studying, evaluating, and building on how these agents actually work. Every trace here was donated only from a public, open-source repository, was anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.tabulartext-generationn<1K35 likes2k downloads3mo agoHugging Face07Open-Eval-Commons /OpenEval OpenEval An open-source, item-centered evaluation repository toward the open science of AI evaluation. This official dataset is maintained by the Open Eval Commons project. For using or contributing to OpenEval (thank you!), please refer to our detailed documentation. 🌐 OpenEval Homepage | 📦 GitHub Repository 🏗️Dataset Structure Currently, the data are split into three tables for storage efficiency: bench, where bench entries are indexed by the field… See the full description on the dataset page: https://huggingface.co/datasets/Open-Eval-Commons/OpenEval.text10M<n<100M9 likes1.6k downloads17d agoHugging Face08tasksource /commonsense_qa_2.0https://github.com/allenai/csqa2 @article{talmor2022commonsenseqa, title={CommonsenseQA 2.0: Exposing the limits of AI through gamification}, author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan}, journal={arXiv preprint arXiv:2201.05320}, year={2022} } textquestion-answering10K<n<100K4 likes1.5k downloads3y agoHugging Face09ziiio /CommonSketch CommonSketch Dataset Summary CommonSketch is a semantically annotated sketch dataset introduced in the paper SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering. The dataset contains 23,100 human-drawn sketches across 300 object classes. Each sketch is paired with a fine-grained caption and element-level commonsense annotations for evaluating sketch abstraction and semantic recognizability. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ziiio/CommonSketch.imageimage-classification10K<n<100K0 likes1.4k downloads5mo agoHugging Face10PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.2k downloads3mo agoHugging Face11commonsense-index-dev /DemoFeedbacktextn<1K0 likes1k downloads2y agoHugging Face12bcv-commons /aligned-mwe aligned_mwe — multi-word target expressions per lexeme Where lexeme-alignments is one row per surface token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia", בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.tabulartranslation1M<n<10M0 likes883 downloads8d agoHugging Face13bcv-commons /target-stopwords target-stopwords Per-language function-word lists, induced from that language's own Bible text — frequency + dispersion (the classic corpus-linguistics stopword-induction recipe), then RESCUED against the language's own alignment output + a source-anchored content signal so genuinely frequent CONTENT words ("God", "Lord") are never dropped. A candidate word is rescued out of the list (judged a real content word, not a function word) only when all four hold — see the… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/target-stopwords.texttext-classification100K<n<1M0 likes825 downloads8d agoHugging Face14jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes822 downloads1y agoHugging Face15zwhe99 /commonsense_170khttps://github.com/AGI-Edgerunners/LLM-Adapters/blob/main/ft-training_set/commonsense_170k.json text100K<n<1M5 likes820 downloads2y agoHugging Face16bcv-commons /senses-attested senses_attested — attested target renderings per lexeme sense The empirical evidence layer produced for shoresh (bcv-query data-contract): for a lexeme in a disambiguated (binyan, sense), which target-language words attest it, with counts. It is the supply that fills shoresh's senses_i18n/_gaps demand and cross-checks the llm_strongs_glosses predictions — it does not replace shoresh's curated senses_i18n/<iso>.tsv; consumed as an HF Parquet dataset. Schema (per row)… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/senses-attested.tabular10M<n<100M0 likes660 downloads8d agoHugging Face17Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes584 downloads1y agoHugging Face18Pclanglais /EU-Science-Commonstabular1M<n<10M0 likes534 downloads4mo agoHugging Face19bcv-commons /lexeme-alignments lexeme-alignments — surface → original-language lexeme (Strong's-bridged) For each language, the attested mapping from target surface word-forms → the original-language lexeme they render, mined by the aligner. Lexeme-anchored, provenance-honest, additive — the design principles are in docs/publishing-principles.md. One language per partition, for consumption by bcv-commons and downstream tools. The language: list above tracks the published partitions; the authoritative list is… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/lexeme-alignments.tabulartranslation10M<n<100M0 likes477 downloads8d agoHugging Face20jinaai /wikimedia-commons-maps_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-maps reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-maps_beir.image1K<n<10K0 likes473 downloads1y agoHugging Face21Lots-of-LoRAs /task828_copa_commonsense_cause_effect Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task828_copa_commonsense_cause_effect Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task828_copa_commonsense_cause_effect.texttext-generationn<1K0 likes459 downloads2y agoHugging Face22Mitsua /safe-commons-pd-3m Safe Commons PD 3M This is a balanced and safe-to-use public domain / CC0 images dataset. All images and texts come from Wikimedia Commons and Wikidata with strict filtering. Images license is either Public Domain or CC0 (varies by image). Texts license is either CC0 or CC BY-SA (varies by caption source). No synthetic data (AI generated images or captions) is in the dataset. To build this dataset, we tried to avoid any knowledge leaks from existing pre-trained models at the… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/safe-commons-pd-3m.imagetext-to-image1M<n<10M6 likes367 downloads2y agoHugging Face23aiintelligentsystems /vel_commons_wikidata Visual Entity Linking: Wikimedia Commons & Wikidata This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict. Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/ uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.image100K<n<1M5 likes358 downloads2y agoHugging Face24Rijgersberg /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.tabulartext-generation10M<n<100M6 likes330 downloads2y agoHugging Face25dutta18 /commonsense_corpus4.7Mtext1M<n<10M1 likes305 downloads3y agoHugging Face26Lots-of-LoRAs /task116_com2sense_commonsense_reasoning Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task116_com2sense_commonsense_reasoning Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task116_com2sense_commonsense_reasoning.texttext-generation1K<n<10K0 likes301 downloads2y agoHugging Face27Lots-of-LoRAs /task827_copa_commonsense_reasoning Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task827_copa_commonsense_reasoning Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task827_copa_commonsense_reasoning.texttext-generationn<1K0 likes294 downloads2y agoHugging Face28KomeijiForce /CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering10K<n<100K0 likes270 downloads3y agoHugging Face29dm-petrov /youtube-commons-small 📺 YouTube-Commons-Small 📺 This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license. Dataset Description This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes. Features The dataset includes the following information for each video: Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.tabulartext-generation100K<n<1M1 likes248 downloads1y agoHugging Face30multi-domain-reasoning /commonsense_qatext1K<n<10K2 likes247 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.