CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01doctor-ghezelbaash /dr-saeid-ghezelbaash-entity-data Dr. Saeed Ghezelbash Public Knowledge Graph A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval. The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.textquestion-answering1K<n<10K1 likes2k downloads8h agoHugging Face02AdhyanshVerma /pg-en Overview Property Value Source Project Gutenberg (English catalog) Snapshot 2026-07-02-18-47-04 Total files 50871 Total Tokens (BPE) ~7.14 billion Total directories 86522 Primary format Plain text (.txt) Bulk format Parquet, JSONL Repository Structure . ├── app.py # Application entry point ├── main.py # Main pipeline / processing script ├── d.py… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/pg-en.tabulartext-generation10K<n<100K1 likes1.4k downloads18d agoHugging Face03beta3 /GridCorpus_9M_Sudoku_Puzzles_Enriched ╔══════════════════════════════════════════════════════════════════════╗ ║ ║ ║ G R I D C O R P U S ║ ║ ║ ║ "004300209005009001070060043..." ║ ║ │ ║ ║ ▼… See the full description on the dataset page: https://huggingface.co/datasets/beta3/GridCorpus_9M_Sudoku_Puzzles_Enriched.tabularfeature-extraction1M<n<10M1 likes1.3k downloads7mo agoHugging Face04yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes992 downloads6mo agoHugging Face05Brainquiver /general-master-en-202608 General · Master · English · 2026-08 English pretraining text, assembled from three public sources, cleaned with one character-level cleaner, and filtered for repetition. 109,337,531 documents and 468,064,046,462 characters. Composition Config Documents Characters What it is fineweb-edu-dedup 65,010,430 297,544,916,118 Web text an educational classifier kept cosmopedia-v2 38,591,146 144,011,993,012 Synthetic prose from a seeded generator… See the full description on the dataset page: https://huggingface.co/datasets/Brainquiver/general-master-en-202608.tabulartext-generation100M<n<1B1 likes897 downloads25d agoHugging Face06klusai /ds-tf1-en-3m 📚 DS-TF1-EN-3M: A Dataset of 3M Moral Fables DS-TF1-EN-3M is a large-scale synthetic dataset of 3 million English moral fables, each crafted using small, instruction-tuned language models (~8B parameters). Every story follows a canonical narrative structure and is designed with pedagogical clarity in mind. 🔗 Project Resources Codebase: github.com/klusai/tinyfabulist 📊 Dataset Summary Metric Average Total Input Tokens 181.53 544,596,141 Output… See the full description on the dataset page: https://huggingface.co/datasets/klusai/ds-tf1-en-3m.tabulartext-generation1M<n<10M4 likes744 downloads1y agoHugging Face07AINovice2005 /carbon-cpu-enriched-sequences carbon-cpu-enriched-sequences A CPU-enriched subset of the carbon pretraining corpus (eukaryote_generator), combining original source fields with normalized sequences and row-level features for quality analysis, GPU enrichment and embedding generation. Information of Features Feature Type Description record_id string NCBI Identifier linking the row back to the source genomic record. It provides the primary record-level identity. begin_of_sequence… See the full description on the dataset page: https://huggingface.co/datasets/AINovice2005/carbon-cpu-enriched-sequences.tabulartext-generation10M<n<100M0 likes719 downloads8d agoHugging Face08Reza2kn /Wikipedia-EN-FA-Accessibility-Bridge Wikipedia EN-FA Accessibility Bridge Current, attributable English and Persian Wikipedia article snapshots for pages created during a 69-day recency window, plus an EN↔FA counterpart index and static accessibility signals. The reproducible full baseline is the official 2026-08-01 Wikimedia dump: 6,289,549 English articles without Persian, 129,821 Persian articles without English, and 22,277,907 namespace-0 pages in the combined parity index. Redirects are retained in the parity… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge.tabulartranslation10M<n<100M1 likes672 downloads1mo agoHugging Face09SultanR /openthoughts3-en-ar-midtrain openthoughts3-en-ar-midtrain Arabic translation of the OpenThoughts3_1.2M split of smoltalk2 (config Mid): long mathematical reasoning traces with <think> blocks, in a two-message user/assistant format. Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 1,135,104 source rows are present, none dropped. The pipeline segments each message into prose and verbatim blocks (code, LaTeX, tables, and inline non-translatables are masked and never sent to the model)… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/openthoughts3-en-ar-midtrain.tabulartext-generation1M<n<10M0 likes637 downloads1mo agoHugging Face10alwaysgood /financial-english-source-corpus Financial English Source Corpus This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. This version preserves the final pre-split source rows. Derived 1280-token split versions are available separately: financial-english-source-corpus-qwen35-1280 financial-english-source-corpus-gemma4-e2b-1280 Dataset Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.tabulartext-generation1M<n<10M0 likes528 downloads14d agoHugging Face11SultanR /nemotron-mc-en-ar-midtrain nemotron-mc-en-ar-midtrain Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.tabulartext-generation10M<n<100M0 likes416 downloads1mo agoHugging Face12SultanR /nemotron-r1-en-ar-midtrain nemotron-r1-en-ar-midtrain Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.tabulartext-generation1M<n<10M0 likes349 downloads1mo agoHugging Face13alwaysgood /financial-english-source-corpus-qwen35-1280 Financial English Source Corpus Qwen35 1280 This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. The uploaded Parquet files are already prepared with the 1280-token source split used by the downstream training pipeline. This split version is derived from the pre-split Financial English Source Corpus by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.tabulartext-generation1M<n<10M0 likes253 downloads3mo agoHugging Face14sxiong /entailmentbank EntailmentBank EntailmentBank is a dataset of multistep entailment trees for open-domain science question answering. Each example links a question and answer to a structured proof: a tree of multi-premise entailment steps from known facts, through intermediate conclusions, to a hypothesis. This repository contains the EMNLP 2021 v2 release in JSONL format, with four configs: Config Description task1 Generate an entailment-tree proof from gold supporting facts task2… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/entailmentbank.tabularquestion-answering10K<n<100K1 likes240 downloads3mo agoHugging Face15BEE-spoke-data /gutenberg-en-v1-clean gutenberg - clean dataset_info: - config_name: default features: - name: text dtype: string - name: label dtype: string - name: score dtype: float64 - name: sha256dtype: string - name: word_count dtype: int64 splits: - name: train num_bytes: 3384868097 num_examples: 9978 - name: validation num_bytes: 195405579 num_examples: 574 - name: test num_bytes: 189439446 num_examples: 565 download_size: 2317462261 dataset_size:… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/gutenberg-en-v1-clean.tabulartext-generation10K<n<100K4 likes220 downloads9mo agoHugging Face16Logics-MLLM /Logics-SWE-Env-2.5K Logics-SWE-Env-2.5K 2,553 software engineering task instances · 1,771 repositories · 4 programming languages 🤗 Related model: Logics-SWE-Qwen3.6-27B 📄 Paper: One to More, More to One 💻 GitHub: AgenticBigBang Overview What is this dataset? Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.tabulartext-generation1K<n<10K3 likes220 downloads22h agoHugging Face17Chess-Nut-Engine /chess-sft-eval Chess SFT Eval & Benchmark Held-out evaluation splits and a frozen benchmark for the Chess SFT training pipeline. Every FEN in these files is excluded from training data via a blocklist to guarantee zero contamination. Eval examples 13,000 Benchmark examples 13,000 Splits 9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate) Format JSONL Training companion Chess-Nut-Engine/chess-sft-data How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.tabulartext-generation10K<n<100K0 likes190 downloads6mo agoHugging Face18Aratako /Synthetic-JP-EN-Coding-Dataset-801k Synthetic-JP-EN-Coding-Dataset-801k Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。 日本語: 173849件 英語: 627413件 元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.tabulartext-generation100K<n<1M17 likes163 downloads2y agoHugging Face19emgena /omnimcp_enterprise_dataops_lakehouse_village_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_enterprise_dataops_lakehouse_village_teaser.tabulartext-generationn<1K0 likes162 downloads7d agoHugging Face20placeholderlabs /exp-pool-encyclopedic-dolma2-tokenized Locus EXP Encyclopedic - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-encyclopedic-dolma2-tokenized.tabulartext-generation1M<n<10M0 likes136 downloads1mo agoHugging Face21placeholderlabs /pool-encyclopedic locus-1 encyclopedic pool The encyclopedic pool for locus-1 - one of seven cluster datasets that are mixed into the pretraining corpus. Every pool shares one schema, so a mixture is a query rather than a rebuild. 6,498,683 documents, 7,314,613,330 tokens under allenai/dolma2-tokenizer@5292e5d6c0f40b67cc765fe41bec991cf4345b5c Built from: HuggingFaceFW/finewiki (en @ 8bd13e72e6a0) Build settings digest: cbd87eae472c6b30 Fingerprint compatibility digest: 6ee2a61c2f91d6a1… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pool-encyclopedic.tabulartext-generation10M<n<100M0 likes123 downloads2mo agoHugging Face22mjbommar /opengloss-v2.0-encyclopedia Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility. OpenGloss v2.0 — Encyclopedia The long-form entry-level prose of OpenGloss v2.0, one row per rendition. The encyclopedia config holds the 300–500-word article about each headword, written at up to five reading levels; the explanation config holds the shorter "why… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-encyclopedia.tabulartext-generation100K<n<1M0 likes121 downloads17d agoHugging Face233nesdeniz /english-daily-dialogues-10k English Daily Dialogues 10K A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tabulartext-generation10K<n<100K2 likes120 downloads1mo agoHugging Face24BEE-spoke-data /code-tutorials-en Dataset Card for "code-tutorials-en" en only 100 words or more reading ease of 50 or more DatasetDict({ train: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 223162 }) validation: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 5873 }) test: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.tabulartext-generation100K<n<1M1 likes113 downloads9mo agoHugging Face25Alookhoshk /llama2-high-entropy-prompts High-entropy prompts for suffix-based backdoor detection Prompts on which base meta-llama/Llama-2-7b-hf has high predictive entropy, built to give a suffix-optimization backdoor detector measurable headroom: a clean model should stay uncertain on these prompts, while a poisoned model driven by a trigger-like suffix should collapse to low entropy. Prompts where the base model is already confident cannot separate the two. How the prompts were made Short prefixes… See the full description on the dataset page: https://huggingface.co/datasets/Alookhoshk/llama2-high-entropy-prompts.tabulartext-generation1K<n<10K0 likes109 downloads17d agoHugging Face26mjbommar /opengloss-v2.1-encyclopedia Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility. OpenGloss v2.1 — Encyclopedia The long-form entry-level prose of OpenGloss v2.1, one row per rendition. The encyclopedia config holds the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-encyclopedia.tabulartext-generation100K<n<1M0 likes108 downloads15d agoHugging Face27nmsofficial /english-distillation-3.5m English Distillation 3.5M English-language distillation corpus containing 3,537,636 rows across 41 Parquet shards. The uploaded Parquet files are the preserved English corpus used for distillation work. tabulartext-generation1M<n<10M0 likes108 downloads12d agoHugging Face28yuanxin112 /morphbench-en-task6-definition-v2 MorphBench-EN Task 6 — Complex-Word Definition (expanded v2) Given a morphologically complex English word, generate its dictionary definition: define word=<word> -> → gloss. Derivations come from UniMorph (eng.derivations.tsv), glosses from Wiktionary. This is the expanded rebuild of the task5a_definition config in yuanxin112/morphbench-en (train 6,651 → 14,971), covering 60 derivational affixes / 32 function labels instead of the original 20 / 13. Splits Splits… See the full description on the dataset page: https://huggingface.co/datasets/yuanxin112/morphbench-en-task6-definition-v2.tabulartext-generation10K<n<100K0 likes107 downloads2mo agoHugging Face29rs545837 /entity-native-agent-sessions Entity-Native vs File-Native Agent Sessions on SWE-bench Verified Full session logs from a controlled A/B experiment measuring how a coding agent's retrieval substrate changes its behaviour, cost, and success rate on real software-engineering tasks. Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the same repository state. The only difference is how the agent is allowed to find code. Arm Label Tools available A file-native Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.tabulartext-generationn<1K0 likes102 downloads22d agoHugging Face30Nicolas-BZRD /English_French_Songs_Lyrics_Translation_Original Original Songs Lyrics with French Translation Dataset Summary Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French. Details of the number of songs by language of origin can be found in the table below: Original language Number of songs en 75786 fr 18486 es 1743 it 803 de 691 sw 529 ko 193 id 169 pt 142 no 122 fi 113 sv 70 hr 53 so 43 ca 41 tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.tabulartranslation10K<n<100K16 likes99 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.