CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tau /commonsense_qa Dataset Card for "commonsense_qa" Dataset Summary CommonsenseQA is a new multiple-choice question answering dataset that requires different types of commonsense knowledge to predict the correct answers . It contains 12,102 questions with one correct answer and four distractor answers. The dataset is provided in two major training/validation/testing set splits: "Random split" which is the main evaluation split, and "Question token split", see paper for details.… See the full description on the dataset page: https://huggingface.co/datasets/tau/commonsense_qa.textquestion-answering10K<n<100K155 likes302k downloads3y agoHugging Face02extraordinarylab /commonsense-qatext10K<n<100K0 likes21k downloads11mo agoHugging Face03PleIAs /YouTube-Commons 📺 YouTube-Commons 📺 YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC-By license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). In total, this represents nearly 45 billion words (44,811,518,375). All the videos where shared on YouTube with a CC-BY license: the dataset provide all the necessary provenance information… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/YouTube-Commons.text-generation396 likes7.3k downloads2y agoHugging Face04bcv-commons /compact-alignments compact-alignments — per-verse, per-book, content-addressed The token-position companion to lexeme-alignments (which is aggregated/type-level and can't tell you what happened in any one verse). This dataset restores position: for a given edition's Bible book, which Hebrew/Greek content word aligned to which target-text token, verse by verse. The authoritative list of what's published is always manifest.json, not this file. Original-language source editions (needed… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/compact-alignments.translation0 likes3.4k downloads7d agoHugging Face05coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.3k downloads8mo agoHugging Face06CUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes3k downloads9mo agoHugging Face07PleIAs /Medical-Commons Medical-Commons Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias. It includes three different collection: International scientific collection of 2M articles from OpenAlex. French scientific collection of XM articles, reports and PhD theses from French institutional repositories. Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion. The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.tabular1M<n<10M2 likes2.3k downloads2y agoHugging Face08trace-commons /agent-traces Trace Commons — Agent Traces Trace Commons is one open, public dataset of coding-agent sessions — the back-and-forth between a developer and an AI coding agent, including prompts, model responses, tool calls, and command output — contributed voluntarily as an open resource for studying, evaluating, and building on how these agents actually work. Every trace here was donated only from a public, open-source repository, was anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.tabulartext-generationn<1K35 likes2k downloads3mo agoHugging Face09Open-Eval-Commons /OpenEval OpenEval An open-source, item-centered evaluation repository toward the open science of AI evaluation. This official dataset is maintained by the Open Eval Commons project. For using or contributing to OpenEval (thank you!), please refer to our detailed documentation. 🌐 OpenEval Homepage | 📦 GitHub Repository 🏗️Dataset Structure Currently, the data are split into three tables for storage efficiency: bench, where bench entries are indexed by the field… See the full description on the dataset page: https://huggingface.co/datasets/Open-Eval-Commons/OpenEval.text10M<n<100M9 likes1.6k downloads16d agoHugging Face10video-reasoning /physical-commonsense0 likes1.6k downloads1y agoHugging Face11tasksource /commonsense_qa_2.0https://github.com/allenai/csqa2 @article{talmor2022commonsenseqa, title={CommonsenseQA 2.0: Exposing the limits of AI through gamification}, author={Talmor, Alon and Yoran, Ori and Bras, Ronan Le and Bhagavatula, Chandra and Goldberg, Yoav and Choi, Yejin and Berant, Jonathan}, journal={arXiv preprint arXiv:2201.05320}, year={2022} } textquestion-answering10K<n<100K4 likes1.6k downloads3y agoHugging Face12ziiio /CommonSketch CommonSketch Dataset Summary CommonSketch is a semantically annotated sketch dataset introduced in the paper SEA: Evaluating Sketch Abstraction Efficiency via Element-level Commonsense Visual Question Answering. The dataset contains 23,100 human-drawn sketches across 300 object classes. Each sketch is paired with a fine-grained caption and element-level commonsense annotations for evaluating sketch abstraction and semantic recognizability. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/ziiio/CommonSketch.imageimage-classification10K<n<100K0 likes1.4k downloads5mo agoHugging Face13PleIAs /French-Science-Commons French Science Commons French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres. Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.tabular10M<n<100M26 likes1.2k downloads3mo agoHugging Face14commonsense-index-dev /DemoFeedbacktextn<1K0 likes1k downloads2y agoHugging Face15bcv-commons /aligned-mwe aligned_mwe — multi-word target expressions per lexeme Where lexeme-alignments is one row per surface token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia", בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.tabulartranslation1M<n<10M0 likes884 downloads8d agoHugging Face16bcv-commons /target-stopwords target-stopwords Per-language function-word lists, induced from that language's own Bible text — frequency + dispersion (the classic corpus-linguistics stopword-induction recipe), then RESCUED against the language's own alignment output + a source-anchored content signal so genuinely frequent CONTENT words ("God", "Lord") are never dropped. A candidate word is rescued out of the list (judged a real content word, not a function word) only when all four hold — see the… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/target-stopwords.texttext-classification100K<n<1M0 likes825 downloads7d agoHugging Face17jinaai /wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.image1K<n<10K0 likes804 downloads1y agoHugging Face18zwhe99 /commonsense_170khttps://github.com/AGI-Edgerunners/LLM-Adapters/blob/main/ft-training_set/commonsense_170k.json text100K<n<1M5 likes773 downloads2y agoHugging Face19bcv-commons /target-morphology target-morphology Per-language unsupervised morphology models — productive suffixes, prefixes, and a stem lexicon, each learned MDL-free ("Linguistica"-style: a suffix is productive if it attaches to many paradigm stems) from that language's own Bible text. No labels, no pretrained model, no download — so it runs on any language with a translation, including those with zero LLM/encoder coverage. stem(word) strips one productive affix when the remainder is a known stem; inflected… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/target-morphology.token-classification0 likes765 downloads15d agoHugging Face20bcv-commons /senses-attested senses_attested — attested target renderings per lexeme sense The empirical evidence layer produced for shoresh (bcv-query data-contract): for a lexeme in a disambiguated (binyan, sense), which target-language words attest it, with counts. It is the supply that fills shoresh's senses_i18n/_gaps demand and cross-checks the llm_strongs_glosses predictions — it does not replace shoresh's curated senses_i18n/<iso>.tsv; consumed as an HF Parquet dataset. Schema (per row)… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/senses-attested.tabular10M<n<100M0 likes668 downloads8d agoHugging Face21gautijha37 /YouTube-Commons 📺 YouTube-Commons 📺 YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC-By license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). In total, this represents nearly 45 billion words (44,811,518,375). All the videos where shared on YouTube with a CC-BY license: the dataset provide all the necessary provenance… See the full description on the dataset page: https://huggingface.co/datasets/gautijha37/YouTube-Commons.text-generation0 likes538 downloads1y agoHugging Face22Pclanglais /EU-Science-Commonstabular1M<n<10M0 likes527 downloads4mo agoHugging Face23bcv-commons /lexeme-alignments lexeme-alignments — surface → original-language lexeme (Strong's-bridged) For each language, the attested mapping from target surface word-forms → the original-language lexeme they render, mined by the aligner. Lexeme-anchored, provenance-honest, additive — the design principles are in docs/publishing-principles.md. One language per partition, for consumption by bcv-commons and downstream tools. The language: list above tracks the published partitions; the authoritative list is… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/lexeme-alignments.tabulartranslation10M<n<100M0 likes509 downloads8d agoHugging Face24Lots-of-LoRAs /task828_copa_commonsense_cause_effect Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task828_copa_commonsense_cause_effect Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task828_copa_commonsense_cause_effect.texttext-generationn<1K0 likes460 downloads2y agoHugging Face25jinaai /wikimedia-commons-maps_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-maps reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-maps_beir.image1K<n<10K0 likes460 downloads1y agoHugging Face26Rijgersberg /YouTube-Commons-nl-audio YouTube Commons NL Audio This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions, all under a CC BY 4.0 license. It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB. Source The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.audioautomatic-speech-recognition10K<n<100K1 likes424 downloads1y agoHugging Face27Jerrylz /common_sense_reasoninggatedThis is the repository containing LoRA checkpoints tuned on common sense reasoning datasets with 0.5B foundation model, served as training data for DnD. 2 likes404 downloads10mo agoHugging Face28Mitsua /safe-commons-pd-3m Safe Commons PD 3M This is a balanced and safe-to-use public domain / CC0 images dataset. All images and texts come from Wikimedia Commons and Wikidata with strict filtering. Images license is either Public Domain or CC0 (varies by image). Texts license is either CC0 or CC BY-SA (varies by caption source). No synthetic data (AI generated images or captions) is in the dataset. To build this dataset, we tried to avoid any knowledge leaks from existing pre-trained models at the… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/safe-commons-pd-3m.imagetext-to-image1M<n<10M6 likes369 downloads2y agoHugging Face29aiintelligentsystems /vel_commons_wikidata Visual Entity Linking: Wikimedia Commons & Wikidata This dataset allows to train and evaluate ML models that link Wikimedia Commons images to the Wikidata items they depict. Disclaimer: All images contained in this dataset are generally assumed to be freely usable (as intended for Wikimedia Commons). Each image's license and author/ uploader is - to the best of our ability - reported in its metadata (see section Dataset Structure). If you want your image's attribution changed or the… See the full description on the dataset page: https://huggingface.co/datasets/aiintelligentsystems/vel_commons_wikidata.image100K<n<1M5 likes358 downloads2y agoHugging Face30Rijgersberg /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.tabulartext-generation10M<n<100M6 likes330 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.