datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval.
The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.pg-en
Overview
Property
Value
Source
Project Gutenberg (English catalog)
Snapshot
2026-07-02-18-47-04
Total files
50871
Total Tokens (BPE)
~7.14 billion
Total directories
86522
Primary format
Plain text (.txt)
Bulk format
Parquet, JSONL
Repository Structure
.
├── app.py # Application entry point
├── main.py # Main pipeline / processing script
├── d.py… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/pg-en.GridCorpus_9M_Sudoku_Puzzles_Enriched
╔══════════════════════════════════════════════════════════════════════╗
║ ║
║ G R I D C O R P U S ║
║ ║
║ "004300209005009001070060043..." ║
║ │ ║
║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.general-master-en-202608
General · Master · English · 2026-08
English pretraining text, assembled from three public sources, cleaned with one
character-level cleaner, and filtered for repetition.
109,337,531 documents and 468,064,046,462 characters.
Composition
Config
Documents
Characters
What it is
fineweb-edu-dedup
65,010,430
297,544,916,118
Web text an educational classifier kept
cosmopedia-v2
38,591,146
144,011,993,012
Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.ds-tf1-en-3m
📚 DS-TF1-EN-3M: A Dataset of 3M Moral Fables
DS-TF1-EN-3M is a large-scale synthetic dataset of 3 million English moral fables, each crafted using small, instruction-tuned language models (~8B parameters). Every story follows a canonical narrative structure and is designed with pedagogical clarity in mind.
🔗 Project Resources
Codebase: github.com/klusai/tinyfabulist
📊 Dataset Summary
Metric
Average
Total
Input Tokens
181.53
544,596,141
Output… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-tf1-en-3m.carbon-cpu-enriched-sequences
carbon-cpu-enriched-sequences
A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences
and row-level features for quality analysis, GPU enrichment and embedding generation.
Information of Features
Feature
Type
Description
record_id
string
NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity.
begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.Wikipedia-EN-FA-Accessibility-Bridge
Wikipedia EN-FA Accessibility Bridge
Current, attributable English and Persian Wikipedia article snapshots for pages
created during a 69-day recency window, plus an EN↔FA
counterpart index and static accessibility signals.
The reproducible full baseline is the official 2026-08-01 Wikimedia dump:
6,289,549 English articles without Persian, 129,821 Persian
articles without English, and 22,277,907 namespace-0 pages in the
combined parity index. Redirects are retained in the parity… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge.openthoughts3-en-ar-midtrain
openthoughts3-en-ar-midtrain
Arabic translation of the OpenThoughts3_1.2M split of smoltalk2 (config Mid): long mathematical reasoning traces with <think> blocks, in a two-message user/assistant format. Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 1,135,104 source rows are present, none dropped.
The pipeline segments each message into prose and verbatim blocks (code, LaTeX, tables, and inline non-translatables are masked and never sent to the model)… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/openthoughts3-en-ar-midtrain.financial-english-source-corpus
Financial English Source Corpus
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. This
version preserves the final pre-split source rows.
Derived 1280-token split versions are available separately:
financial-english-source-corpus-qwen35-1280
financial-english-source-corpus-gemma4-e2b-1280
Dataset
Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.nemotron-mc-en-ar-midtrain
nemotron-mc-en-ar-midtrain
Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.nemotron-r1-en-ar-midtrain
nemotron-r1-en-ar-midtrain
Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.financial-english-source-corpus-qwen35-1280
Financial English Source Corpus Qwen35 1280
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. The
uploaded Parquet files are already prepared with the 1280-token source split
used by the downstream training pipeline.
This split version is derived from the pre-split
Financial English Source Corpus
by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.entailmentbank
EntailmentBank
EntailmentBank is a dataset of multistep entailment trees for open-domain science question answering. Each example links a question and answer to a structured proof: a tree of multi-premise entailment steps from known facts, through intermediate conclusions, to a hypothesis.
This repository contains the EMNLP 2021 v2 release in JSONL format, with four configs:
Config
Description
task1
Generate an entailment-tree proof from gold supporting facts
task2… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/entailmentbank.gutenberg-en-v1-clean
gutenberg - clean
dataset_info:
- config_name: default
features:
- name: text
dtype: string
- name: label
dtype: string
- name: score
dtype: float64
- name: sha256dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 3384868097
num_examples: 9978
- name: validation
num_bytes: 195405579
num_examples: 574
- name: test
num_bytes: 189439446
num_examples: 565
download_size: 2317462261
dataset_size:… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/gutenberg-en-v1-clean.Logics-SWE-Env-2.5K
Logics-SWE-Env-2.5K
2,553 software engineering task instances · 1,771 repositories · 4 programming languages
🤗 Related model: Logics-SWE-Qwen3.6-27B
📄 Paper: One to More, More to One
💻 GitHub: AgenticBigBang
Overview
What is this dataset?
Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.chess-sft-eval
Chess SFT Eval & Benchmark
Held-out evaluation splits and a frozen benchmark for the
Chess SFT training pipeline.
Every FEN in these files is excluded from training data via a blocklist to guarantee
zero contamination.
Eval examples
13,000
Benchmark examples
13,000
Splits
9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate)
Format
JSONL
Training companion
Chess-Nut-Engine/chess-sft-data
How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.Synthetic-JP-EN-Coding-Dataset-801k
Synthetic-JP-EN-Coding-Dataset-801k
Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。
日本語: 173849件
英語: 627413件
元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。
nvidia/Nemotron-4-340B-Instruct
microsoft/Phi-3-medium-4k-instruct
mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.omnimcp_enterprise_dataops_lakehouse_village_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_enterprise_dataops_lakehouse_village_teaser.exp-pool-encyclopedic-dolma2-tokenized
Locus EXP Encyclopedic - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.pool-encyclopedic
locus-1 encyclopedic pool
The encyclopedic pool for locus-1 - one of seven cluster
datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a
mixture is a query rather than a rebuild.
6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c
Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0)
Build settings digest: cbd87eae472c6b30
Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.opengloss-v2.0-encyclopedia
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Encyclopedia
The long-form entry-level prose of OpenGloss v2.0, one row per rendition. The encyclopedia config holds the 300–500-word article about each headword, written at up to five reading levels; the explanation config holds the shorter "why… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-encyclopedia.english-daily-dialogues-10k
English Daily Dialogues 10K
A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark.
Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.code-tutorials-en
Dataset Card for "code-tutorials-en"
en only
100 words or more
reading ease of 50 or more
DatasetDict({
train: Dataset({
features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],
num_rows: 223162
})
validation: Dataset({
features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],
num_rows: 5873
})
test: Dataset({
features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.llama2-high-entropy-prompts
High-entropy prompts for suffix-based backdoor detection
Prompts on which base meta-llama/Llama-2-7b-hf has high predictive
entropy, built to give a suffix-optimization backdoor detector measurable
headroom: a clean model should stay uncertain on these prompts, while a poisoned
model driven by a trigger-like suffix should collapse to low entropy. Prompts
where the base model is already confident cannot separate the two.
How the prompts were made
Short prefixes… See the full description on the dataset page: https://huggingface.co/datasets/Alookhoshk/llama2-high-entropy-prompts.opengloss-v2.1-encyclopedia
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Encyclopedia
The long-form entry-level prose of OpenGloss v2.1, one row per rendition. The encyclopedia config holds the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-encyclopedia.english-distillation-3.5m
English Distillation 3.5M
English-language distillation corpus containing 3,537,636 rows across 41 Parquet shards.
The uploaded Parquet files are the preserved English corpus used for distillation work.
morphbench-en-task6-definition-v2
MorphBench-EN Task 6 — Complex-Word Definition (expanded v2)
Given a morphologically complex English word, generate its dictionary
definition: define word=<word> -> → gloss. Derivations come from UniMorph
(eng.derivations.tsv), glosses from Wiktionary.
This is the expanded rebuild of the task5a_definition config in
yuanxin112/morphbench-en
(train 6,651 → 14,971), covering 60 derivational affixes / 32 function labels
instead of the original 20 / 13.
Splits
Splits… See the full description on the dataset page: https://huggingface.co/datasets/yuanxin112/morphbench-en-task6-definition-v2.entity-native-agent-sessions
Entity-Native vs File-Native Agent Sessions on SWE-bench Verified
Full session logs from a controlled A/B experiment measuring how a coding agent's
retrieval substrate changes its behaviour, cost, and success rate on real
software-engineering tasks.
Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the
same repository state. The only difference is how the agent is allowed to find code.
Arm
Label
Tools available
A
file-native
Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.
