datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
m-ArenaHard-v2.0
Dataset Card for m-ArenaHard-v2.0
This dataset is used in the paper When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs.
Dataset Details
The m-ArenaHard-v2.0 dataset is a multilingual LLM evaluation set. This is built on the LMarena (formerly LMSYS) arena-hard-auto-v2.0 test dataset.
This dataset(containing 750 prompts) was filtered to "english" only prompts using the papluca/xlm-roberta-base-language-detection model resulting… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/m-ArenaHard-v2.0.CodeChat-V2.0
CodeChat: Developer–LLM Conversations Dataset
Paper: https://arxiv.org/abs/2509.10402
GitHub: https://github.com/Software-Evolution-Analytics-Lab-SEAL/CodeChat
CodeChat_2 is a large-scale dataset comprising 587,568 real-world developer–LLM conversations, derived from the WildChat dataset.
Dataset Overview
Field
👉V1.0
V2.0
Records
82,845 conversations
587,568 conversations
Code
368,506 code snippets
2,252,399 code snippets
Languages
20+… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/CodeChat-V2.0.opengloss-v2.0-qa-pairs
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — QA Pairs
Question/answer pairs written per sense and answerable only from that sense's own stored text — its gloss, its examples, its entry's encyclopedia article and etymology — with every source labelled by an id the answer has to cite. Uncited… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-qa-pairs.opengloss-v2.0-encyclopedia
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Encyclopedia
The long-form entry-level prose of OpenGloss v2.0, one row per rendition. The encyclopedia config holds the 300–500-word article about each headword, written at up to five reading levels; the explanation config holds the shorter "why… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-encyclopedia.opengloss-v2.0-contrasts
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Contrasts
For every synonym, antonym or confusable_with edge whose far end resolves to a sense that is actually in the release, one 60–120 word paragraph saying how the two terms actually differ — the register that separates them, the axis they… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-contrasts.opengloss-v2.0-pretrain
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Pretrain
The release rendered as continuous prose for language-model pretraining or continued pretraining: four document templates per entry — a dictionary entry, a thesaurus entry, an encyclopedia article and a usage note — written as plain text… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-pretrain.opengloss-v2.0-definitions
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Definitions
The flat definition view: one row for every stored rendition of every live sense's definition, the canonical (neutral, plain) gloss included. This is the reading-level and register grading of OpenGloss v2.0 laid out one row at a time… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-definitions.opengloss-v2.0-queries
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Queries
Synthetic search queries written per sense, in eight styles — keyword, question, conversational, constraint, role, example-based, step-by-step and directive — with each sense's sibling senses in the prompt so the queries discriminate… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-queries.opengloss-v2.0-etymology
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Etymology
Structured word histories for OpenGloss v2.0: one row per entry that has an etymology, with a prose summary and the ordered trail of source languages, each segment carrying its language, ISO 639-3 code where one applies, attested form… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-etymology.opengloss-v2.0-lexicon
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Lexicon
The entry-level view of OpenGloss v2.0: one row per lexeme, with everything that belongs to the entry rather than to one of its meanings — the kind discriminator, per-POS morphology, structured etymology, the lexical explanation, the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-lexicon.opengloss-v2.0-senses
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Senses
The sense-level view of OpenGloss v2.0 and the repo most consumers want: one row per live sense, with its canonical gloss, its eight reading-level and register renditions, its sense-tagged example sentences with headword character spans, its… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-senses.hyperion-v2.0
Hyperion v2.0
Introduction
Hyperion is a comprehensive question answering and conversational dataset designed to promote advancements in AI research with a particular emphasis on reasoning and understanding in scientific domains such as science, medicine, mathematics, and computer science. It integrates data from a wide array of datasets, facilitating the development of models capable of handling complex inquiries and instructions.
Dataset Description
Hyperion… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v2.0.han-instruct-dataset-v2.0
Dataset Card for Han Instruct Dataset v2.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collect all Thai instruct dataset that made by human and our old model. The dataset can use to train Instruction Following model like ChatGPT or other.
Many question are collect from Reference desk at Thai wikipedia.
Data sources:
Reference desk… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v2.0.AdParaphrase-v2.0
AdParaphrase v2.0
This repository contains data for our paper "AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset" (ACL2025 Findings).
Overview
AdParaphrase v2.0 is a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generating attractive ad texts. Compared with AdParaphrase v1.0, this dataset is 20 times larger… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/AdParaphrase-v2.0.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/Cybersecurity-Dataset-Fenrir-v2.0.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/PhSecX/Cybersecurity-Dataset-Fenrir-v2.0.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/invinciblejha01/Cybersecurity-Dataset-Fenrir-v2.0.ChatML-hercules-v2.0Locutusque/hercules-v2.0 in ChatML format.
Python code used for conversion:
from datasets import load_dataset
import pandas
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
pretrained_model_name_or_path="Felladrin/Llama-160M-Chat-v1"
)
dataset = load_dataset("Locutusque/hercules-v2.0", split="train")
def format(columns):
messages = []
conversation = columns["conversations"]
for i in range(len(conversation)):
message =… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-hercules-v2.0.Cybersecurity-Dataset-Fenrir-v2.0
Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 83,920 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/Zud0/Cybersecurity-Dataset-Fenrir-v2.0.CiviVox-Swahili-text-corpus-v2.0
Swahili Text Dataset
Overview
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Dataset Details
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: 1.54M
Format: Hugging Face Dataset
Content
The dataset consists of two main columns:
id: A unique identifier for each text entry
text:… See the full description on the dataset page: https://huggingface.co/datasets/Adeptschneider/CiviVox-Swahili-text-corpus-v2.0.CiviVox-Swahili-text-corpus-v2.0
Swahili Text Dataset
Overview
This dataset contains a comprehensive collection of Swahili text data, derived from the AfriBERTa Corpus. It provides a rich resource for natural language processing tasks focused on the Swahili language.
Dataset Details
Source: AfriBERTa Corpus (Swahili subset)
Language: Swahili
Size: 1.54M
Format: Hugging Face Dataset
Content
The dataset consists of two main columns:
id: A unique identifier for each text entry
text:… See the full description on the dataset page: https://huggingface.co/datasets/niqqyniqqy/CiviVox-Swahili-text-corpus-v2.0.Mistral-7B-Instruct-v2.0-PairRM-DPO-Dataset
