CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /m-ArenaHard-v2.0 Dataset Card for m-ArenaHard-v2.0 This dataset is used in the paper When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs. Dataset Details The m-ArenaHard-v2.0 dataset is a multilingual LLM evaluation set. This is built on the LMarena (formerly LMSYS) arena-hard-auto-v2.0 test dataset. This dataset(containing 750 prompts) was filtered to "english" only prompts using the papluca/xlm-roberta-base-language-detection model resulting… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/m-ArenaHard-v2.0.texttext-generation10K<n<100K7 likes650 downloads5mo agoHugging Face02Suzhen /CodeChat-V2.0 CodeChat: Developer–LLM Conversations Dataset Paper: https://arxiv.org/abs/2509.10402 GitHub: https://github.com/Software-Evolution-Analytics-Lab-SEAL/CodeChat CodeChat_2 is a large-scale dataset comprising 587,568 real-world developer–LLM conversations, derived from the WildChat dataset. Dataset Overview Field 👉V1.0 V2.0 Records 82,845 conversations 587,568 conversations Code 368,506 code snippets 2,252,399 code snippets Languages 20+… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/CodeChat-V2.0.texttext-generation100K<n<1M1 likes351 downloads2mo agoHugging Face03mjbommar /opengloss-v2.0-qa-pairs Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — QA Pairs Question/answer pairs written per sense and answerable only from that sense's own stored text — its gloss, its examples, its entry's encyclopedia article and etymology — with every source labelled by an id the answer has to cite. Uncited… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-qa-pairs.textquestion-answering100K<n<1M0 likes143 downloads19d agoHugging Face04mjbommar /opengloss-v2.0-encyclopedia Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Encyclopedia The long-form entry-level prose of OpenGloss v2.0, one row per rendition. The encyclopedia config holds the 300–500-word article about each headword, written at up to five reading levels; the explanation config holds the shorter "why… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-encyclopedia.tabulartext-generation100K<n<1M0 likes124 downloads19d agoHugging Face05mjbommar /opengloss-v2.0-contrasts Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Contrasts For every synonym, antonym or confusable_with edge whose far end resolves to a sense that is actually in the release, one 60–120 word paragraph saying how the two terms actually differ — the register that separates them, the axis they… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-contrasts.texttext-generation10K<n<100K0 likes122 downloads19d agoHugging Face06mjbommar /opengloss-v2.0-pretrain Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Pretrain The release rendered as continuous prose for language-model pretraining or continued pretraining: four document templates per entry — a dictionary entry, a thesaurus entry, an encyclopedia article and a usage note — written as plain text… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-pretrain.texttext-generation100K<n<1M0 likes121 downloads19d agoHugging Face07mjbommar /opengloss-v2.0-definitions Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Definitions The flat definition view: one row for every stored rendition of every live sense's definition, the canonical (neutral, plain) gloss included. This is the reading-level and register grading of OpenGloss v2.0 laid out one row at a time… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-definitions.tabulartext-generation1M<n<10M0 likes117 downloads19d agoHugging Face08mjbommar /opengloss-v2.0-queries Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Queries Synthetic search queries written per sense, in eight styles — keyword, question, conversational, constraint, role, example-based, step-by-step and directive — with each sense's sibling senses in the prompt so the queries discriminate… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-queries.texttext-retrieval1M<n<10M0 likes114 downloads19d agoHugging Face09mjbommar /opengloss-v2.0-etymology Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Etymology Structured word histories for OpenGloss v2.0: one row per entry that has an etymology, with a prose summary and the ordered trail of source languages, each segment carrying its language, ISO 639-3 code where one applies, attested form… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-etymology.texttext-generation10K<n<100K0 likes112 downloads19d agoHugging Face10mjbommar /opengloss-v2.0-lexicon Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Lexicon The entry-level view of OpenGloss v2.0: one row per lexeme, with everything that belongs to the entry rather than to one of its meanings — the kind discriminator, per-POS morphology, structured etymology, the lexical explanation, the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-lexicon.tabulartext-generation10K<n<100K0 likes111 downloads19d agoHugging Face11mjbommar /opengloss-v2.0-senses Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Senses The sense-level view of OpenGloss v2.0 and the repo most consumers want: one row per live sense, with its canonical gloss, its eight reading-level and register renditions, its sense-tagged example sentences with headword character spans, its… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-senses.tabulartext-generation100K<n<1M0 likes103 downloads19d agoHugging Face12Locutusque /hyperion-v2.0 Hyperion v2.0 Introduction Hyperion is a comprehensive question answering and conversational dataset designed to promote advancements in AI research with a particular emphasis on reasoning and understanding in scientific domains such as science, medicine, mathematics, and computer science. It integrates data from a wide array of datasets, facilitating the development of models capable of handling complex inquiries and instructions. Dataset Description Hyperion… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v2.0.texttext-generation1M<n<10M6 likes84 downloads3y agoHugging Face13pythainlp /han-instruct-dataset-v2.0 Dataset Card for Han Instruct Dataset v2.0 The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset. 🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collect all Thai instruct dataset that made by human and our old model. The dataset can use to train Instruction Following model like ChatGPT or other. Many question are collect from Reference desk at Thai wikipedia. Data sources: Reference desk… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v2.0.texttext-generation1K<n<10K3 likes57 downloads6d agoHugging Face14cyberagent /AdParaphrase-v2.0 AdParaphrase v2.0 This repository contains data for our paper "AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset" (ACL2025 Findings). Overview AdParaphrase v2.0 is a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generating attractive ad texts. Compared with AdParaphrase v1.0, this dataset is 20 times larger… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/AdParaphrase-v2.0.tabulartext-generation10K<n<100K1 likes32 downloads1y agoHugging Face15ukcli /Cybersecurity-Dataset-Fenrir-v2.0 Cybersecurity Defense Instruction-Tuning Dataset (v2.0) Created by Alican Kiraz TL;DR A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Cybersecurity-Dataset-Fenrir-v2.0.texttext-generation10K<n<100K0 likes32 downloads6mo agoHugging Face16PhSecX /Cybersecurity-Dataset-Fenrir-v2.0 Cybersecurity Defense Instruction-Tuning Dataset (v2.0) Created by Alican Kiraz TL;DR A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/PhSecX/Cybersecurity-Dataset-Fenrir-v2.0.texttext-generation10K<n<100K0 likes29 downloads6mo agoHugging Face17invinciblejha01 /Cybersecurity-Dataset-Fenrir-v2.0 Cybersecurity Defense Instruction-Tuning Dataset (v2.0) Created by Alican Kiraz TL;DR A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Cybersecurity-Dataset-Fenrir-v2.0.texttext-generation10K<n<100K0 likes26 downloads6mo agoHugging Face18Felladrin /ChatML-hercules-v2.0Locutusque/hercules-v2.0 in ChatML format. Python code used for conversion: from datasets import load_dataset import pandas from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained( pretrained_model_name_or_path="Felladrin/Llama-160M-Chat-v1" ) dataset = load_dataset("Locutusque/hercules-v2.0", split="train") def format(columns): messages = [] conversation = columns["conversations"] for i in range(len(conversation)): message =… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-hercules-v2.0.textquestion-answering1M<n<10M1 likes23 downloads3y agoHugging Face19Zud0 /Cybersecurity-Dataset-Fenrir-v2.0 Cybersecurity Defense Instruction-Tuning Dataset (v2.0) Created by Alican Kiraz TL;DR A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training. Apache-2.0 licensed and production-ready. Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security. 1  What’s… See the full description on the dataset page: https://huggingface.co/datasets/Zud0/Cybersecurity-Dataset-Fenrir-v2.0.texttext-generation10K<n<100K0 likes18 downloads6mo agoHugging Face20Adeptschneider /CiviVox-Swahili-text-corpus-v2.0 Swahili Text Dataset Overview This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Dataset Details Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset Content The dataset consists of two main columns: id: A unique identifier for each text entry text:… See the full description on the dataset page: https://huggingface.co/datasets/Adeptschneider/CiviVox-Swahili-text-corpus-v2.0.texttext-generation1M<n<10M0 likes14 downloads2y agoHugging Face21niqqyniqqy /CiviVox-Swahili-text-corpus-v2.0 Swahili Text Dataset Overview This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language. Dataset Details Source: AfriBERTa Corpus (Swahili subset) Language: Swahili Size: 1.54M Format: Hugging Face Dataset Content The dataset consists of two main columns: id: A unique identifier for each text entry text:… See the full description on the dataset page: https://huggingface.co/datasets/niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0.texttext-generation1M<n<10M0 likes14 downloads6mo agoHugging Face22DhruvParth /Mistral-7B-Instruct-v2.0-PairRM-DPO-Datasettexttext-generationn<1K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.