datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.financial-english-source-corpus
Financial English Source Corpus
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. This
version preserves the final pre-split source rows.
Derived 1280-token split versions are available separately:
financial-english-source-corpus-qwen35-1280
financial-english-source-corpus-gemma4-e2b-1280
Dataset
Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.financial-english-source-corpus-qwen35-1280
Financial English Source Corpus Qwen35 1280
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. The
uploaded Parquet files are already prepared with the 1280-token source split
used by the downstream training pipeline.
This split version is derived from the pre-split
Financial English Source Corpus
by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.chess-sft-eval
Chess SFT Eval & Benchmark
Held-out evaluation splits and a frozen benchmark for the
Chess SFT training pipeline.
Every FEN in these files is excluded from training data via a blocklist to guarantee
zero contamination.
Eval examples
13,000
Benchmark examples
13,000
Splits
9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate)
Format
JSONL
Training companion
Chess-Nut-Engine/chess-sft-data
How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.english-daily-dialogues-10k
English Daily Dialogues 10K
A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark.
Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.english-distillation-3.5m
English Distillation 3.5M
English-language distillation corpus containing 3,537,636 rows across 41 Parquet shards.
The uploaded Parquet files are the preserved English corpus used for distillation work.
English_French_Songs_Lyrics_Translation_Original
Original Songs Lyrics with French Translation
Dataset Summary
Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French.
Details of the number of songs by language of origin can be found in the table below:
Original language
Number of songs
en
75786
fr
18486
es
1743
it
803
de
691
sw
529
ko
193
id
169
pt
142
no
122
fi
113
sv
70
hr
53
so
43
ca
41
tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.wiki_paragraphs_english
WIKI Paragraphs English
A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format.
Features
Multiple splits for different use cases
Random shuffle with Fisher-Yates algorithm
Structured format with text and metadata
Size-varied validation/test sets (100 to 10k samples)
Splits Overview
Split Name
Samples
Typical Usage
train
1,000,000
Primary training data
validation
10,000
Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.amazon-esci-english-smallMath_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.DBNL-public-qa-english-translationreasoning_english
Norwegian Reasoning
A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true.
A total of 30.036 tasks where generated.
Of these a total of 6745 tasks had the exact correct answer.
This was split into test=250, validation=250 and train=6245
chess-sft-corpus-4x-eval
Chess SFT Eval and Benchmark
Held-out evaluation splits and a frozen benchmark for
Chess-Nut-Engine/chess-sft-corpus-4x.
Every FEN in these files is excluded from generated training data (the
blocklist is game-scoped: sibling positions of eval games are excluded too).
Frozen from the 4x corpus generation run of 2026-07-06 (generator revision 3cd161b1078cdfa6598fba939f40250072adb524)
Benchmark: 13,000 frozen examples across 9 splits; eval splits share the game-scoped blocklist… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-corpus-4x-eval.bible-parallel-english
Parallel Bible — English Translations and Ancient Versions
A verse-aligned parallel corpus of the Protestant Bible in seventeen English
translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta
for the New Testament.
Looking for every language? This repository is a curated English set,
chosen for spread across translation families and small enough to load whole.
For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses —
see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.chief-engineer-deliberation
Chief Engineer — Deliberation Traces
Turn-by-turn multi-persona deliberation from The Chief Engineer, a small local
Gemma agent built for the HF Build Small hackathon (Backyard AI). Where the
lesson ledger
records what the agent learned, this records how it reasons: the argument between
the personas on each job. It grows two ways: a reproducible static export
(make deliberation) and live turns logged on every run of the Space (gated on
HF_TOKEN; config + agent reasoning only… See the full description on the dataset page: https://huggingface.co/datasets/kylebrodeur/chief-engineer-deliberation.tg-ru-engagement
📨 Telegram Russian Posts — Engagement Dataset
Коллекция русскоязычных постов из Telegram-каналов с метриками вовлечённости (просмотры, репосты). Датасет предназначен для задач fine-tuning языковых моделей, предсказания вирусности и классификации контента.
Dataset Summary
Язык
Русский
Записей
70 448
Период
Октябрь 2021 — Июль 2026
Источник
Telegram-каналы (парсинг)
Средняя длина текста
274 символа
Медиана просмотров
145 109
Макс.… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/tg-ru-engagement.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tatar-english-russian-corpus
Dataset Card: Tatar-English-Russian Parallel Corpus
Dataset Details
Dataset Description
This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources:
KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added
yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.SlimOrca-Dedup-English-UzbekThis is an Uzbek translated version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.
It is a single parquet file.
Check here for cleaned Uzbek only slim Orca dataset: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned
prompt-engineering-fr
Prompt Engineering FR - Techniques, Evaluation et Gestion du Contexte
Dataset bilingue complet sur le Prompt Engineering, l'evaluation de LLM et la gestion de la fenetre de contexte.
Cree par AYI NEDJIMI Consultants - Expertise en Intelligence Artificielle et Transformation Digitale.
Description
Ce dataset couvre l'ensemble des techniques modernes de prompt engineering, les benchmarks et metriques d'evaluation de LLM, ainsi que les strategies de gestion de la fenetre de… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/prompt-engineering-fr.prompt-engineering-en
Prompt Engineering EN - Techniques, Evaluation & Context Management
Comprehensive bilingual dataset on Prompt Engineering, LLM evaluation, and context window management.
Created by AYI NEDJIMI Consultants - Expertise in Artificial Intelligence and Digital Transformation.
Description
This dataset covers all modern prompt engineering techniques, LLM evaluation benchmarks and metrics, and context window management strategies. It is based on three reference articles:
Prompt… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/prompt-engineering-en.apollo_english_guidelines_translated_to_dutch_with_nllb200
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
english_islamqainfo
Dataset Card for English Islam QA Info
Dataset Description
The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice.
Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.Magpie-Qwen2-Pro-200K-English-koTranslated Magpie-Align/Magpie-Qwen2-Pro-200K-English using nayohan/llama3-instrucTrans-enko-8b.
For this dataset, we only used data that is 5000 characters or less in length and has language of English.
Thanks for @Magpie-Align and @nayohan.
@misc{xu2024magpie,
title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing},
author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Qwen2-Pro-200K-English-ko.Quran_English_Myanmar_Parrelel_Corpus
Quran English-Myanmar Parallel Corpus
Description
This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers.
English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali.
Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.EnglishtoFrench-Translation-Dataset
English–French Translation Dataset (SFT / LoRA Ready)
A clean, structured dataset of 50,000 English–French sentence pairs designed
for supervised fine-tuning (SFT) of large language models, LoRA adapters, and
general machine translation tasks.
Overview
Property
Value
Language pair
English → French
Total rows
50,000
Train split
45,000 (90%)
Validation split
2,500 (5%)
Test split
2,500 (5%)
Format
CSV (Alpaca-style prompt format)
License
CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.tamil-english-corpus
Tamil-English Retrieval Corpus
A high-quality multilingual retrieval corpus constructed from the Mozhi Tamil Corpus and machine-translated into English using IndicTrans2.
Dataset Summary
This dataset contains Tamil documents paired with English translations.
The corpus was created by filtering high-quality documents from the Mozhi Tamil Corpus and translating them using AI4Bharat's IndicTrans2 translation model.
The resulting corpus is intended to support:… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tamil-english-corpus.Motivation_Employee_Engagement_Content_1
Motivation Employee Engagement Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_1.apollo_english_guidelines_translated_to_dutch_with_marianmt
Data description
Apollo corpus, English guidelines translated to Dutch using MariaNMT.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.
