CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes992 downloads7mo agoHugging Face02alwaysgood /financial-english-source-corpus Financial English Source Corpus This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. This version preserves the final pre-split source rows. Derived 1280-token split versions are available separately: financial-english-source-corpus-qwen35-1280 financial-english-source-corpus-gemma4-e2b-1280 Dataset Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.tabulartext-generation1M<n<10M0 likes528 downloads14d agoHugging Face03alwaysgood /financial-english-source-corpus-qwen35-1280 Financial English Source Corpus Qwen35 1280 This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. The uploaded Parquet files are already prepared with the 1280-token source split used by the downstream training pipeline. This split version is derived from the pre-split Financial English Source Corpus by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.tabulartext-generation1M<n<10M0 likes253 downloads3mo agoHugging Face04Chess-Nut-Engine /chess-sft-eval Chess SFT Eval & Benchmark Held-out evaluation splits and a frozen benchmark for the Chess SFT training pipeline. Every FEN in these files is excluded from training data via a blocklist to guarantee zero contamination. Eval examples 13,000 Benchmark examples 13,000 Splits 9 (perception, rules, tactics, evaluation, openings, endgames, planning, chess960, mate) Format JSONL Training companion Chess-Nut-Engine/chess-sft-data How eval and benchmark differ… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-eval.tabulartext-generation10K<n<100K0 likes190 downloads6mo agoHugging Face053nesdeniz /english-daily-dialogues-10k English Daily Dialogues 10K A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tabulartext-generation10K<n<100K2 likes120 downloads1mo agoHugging Face06nmsofficial /english-distillation-3.5m English Distillation 3.5M English-language distillation corpus containing 3,537,636 rows across 41 Parquet shards. The uploaded Parquet files are the preserved English corpus used for distillation work. tabulartext-generation1M<n<10M0 likes108 downloads13d agoHugging Face07Nicolas-BZRD /English_French_Songs_Lyrics_Translation_Original Original Songs Lyrics with French Translation Dataset Summary Dataset of 99289 songs containing their metadata (author, album, release date, song number), original lyrics and lyrics translated into French. Details of the number of songs by language of origin can be found in the table below: Original language Number of songs en 75786 fr 18486 es 1743 it 803 de 691 sw 529 ko 193 id 169 pt 142 no 122 fi 113 sv 70 hr 53 so 43 ca 41 tl… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Songs_Lyrics_Translation_Original.tabulartranslation10K<n<100K16 likes99 downloads3y agoHugging Face08pere /wiki_paragraphs_english WIKI Paragraphs English A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.tabulartext-generation1M<n<10M0 likes84 downloads2y agoHugging Face09thebajajra /amazon-esci-english-smalltabulartoken-classification100K<n<1M1 likes81 downloads1y agoHugging Face10miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes77 downloads1y agoHugging Face11nscharrenberg /DBNL-public-qa-english-translationtabularquestion-answering10K<n<100K0 likes54 downloads1y agoHugging Face12pere /reasoning_english Norwegian Reasoning A reasoning dataset made by DeepSeek R1. The reasoning data is made from punctuation-restoration tasks from Wikipedia. We have stored the reasoning in cases where the output is 100% true. A total of 30.036 tasks where generated. Of these a total of 6745 tasks had the exact correct answer. This was split into test=250, validation=250 and train=6245 tabulartext-generation1K<n<10K0 likes53 downloads2y agoHugging Face13Chess-Nut-Engine /chess-sft-corpus-4x-eval Chess SFT Eval and Benchmark Held-out evaluation splits and a frozen benchmark for Chess-Nut-Engine/chess-sft-corpus-4x. Every FEN in these files is excluded from generated training data (the blocklist is game-scoped: sibling positions of eval games are excluded too). Frozen from the 4x corpus generation run of 2026-07-06 (generator revision 3cd161b1078cdfa6598fba939f40250072adb524) Benchmark: 13,000 frozen examples across 9 splits; eval splits share the game-scoped blocklist… See the full description on the dataset page: https://huggingface.co/datasets/Chess-Nut-Engine/chess-sft-corpus-4x-eval.tabulartext-generation10K<n<100K0 likes51 downloads3mo agoHugging Face14sermonindex /bible-parallel-english Parallel Bible — English Translations and Ancient Versions A verse-aligned parallel corpus of the Protestant Bible in seventeen English translations, spanning 1599 to 2022, plus the Latin Vulgate and Syriac Peshitta for the New Testament. Looking for every language? This repository is a curated English set, chosen for spread across translation families and small enough to load whole. For the full corpus — 1,253 translations in 1,004 languages, 14.4M verses — see… See the full description on the dataset page: https://huggingface.co/datasets/sermonindex/bible-parallel-english.tabulartext-generation100K<n<1M0 likes50 downloads14d agoHugging Face15kylebrodeur /chief-engineer-deliberation Chief Engineer — Deliberation Traces Turn-by-turn multi-persona deliberation from The Chief Engineer, a small local Gemma agent built for the HF Build Small hackathon (Backyard AI). Where the lesson ledger records what the agent learned, this records how it reasons: the argument between the personas on each job. It grows two ways: a reproducible static export (make deliberation) and live turns logged on every run of the Space (gated on HF_TOKEN; config + agent reasoning only… See the full description on the dataset page: https://huggingface.co/datasets/kylebrodeur/chief-engineer-deliberation.tabulartext-generationn<1K0 likes48 downloads26d agoHugging Face16ru-dataset /tg-ru-engagement 📨 Telegram Russian Posts — Engagement Dataset Коллекция русскоязычных постов из Telegram-каналов с метриками вовлечённости (просмотры, репосты). Датасет предназначен для задач fine-tuning языковых моделей, предсказания вирусности и классификации контента. Dataset Summary Язык Русский Записей 70 448 Период Октябрь 2021 — Июль 2026 Источник Telegram-каналы (парсинг) Средняя длина текста 274 символа Медиана просмотров 145 109 Макс.… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/tg-ru-engagement.tabulartext-generation10K<n<100K1 likes47 downloads2mo agoHugging Face17miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes41 downloads1y agoHugging Face18TatarNLPWorld /tatar-english-russian-corpusgated Dataset Card: Tatar-English-Russian Parallel Corpus Dataset Details Dataset Description This dataset is a parallel corpus containing 14,983 sentences in three languages: Tatar, English, and Russian. It combines two distinct sources: KickItLikeShika/english-tatar-translation (7,746 entries) – an existing English-Tatar dataset with Russian translations added yasalma/tt-en-language-corpus (7,615 entries) – another English-Tatar corpus with newly added… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-english-russian-corpus.tabulartranslation10K<n<100K0 likes32 downloads1mo agoHugging Face19MLDataScientist /SlimOrca-Dedup-English-UzbekThis is an Uzbek translated version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup. It is a single parquet file. Check here for cleaned Uzbek only slim Orca dataset: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned tabulartext-classification1M<n<10M2 likes31 downloads2y agoHugging Face20AYI-NEDJIMI /prompt-engineering-fr Prompt Engineering FR - Techniques, Evaluation et Gestion du Contexte Dataset bilingue complet sur le Prompt Engineering, l'evaluation de LLM et la gestion de la fenetre de contexte. Cree par AYI NEDJIMI Consultants - Expertise en Intelligence Artificielle et Transformation Digitale. Description Ce dataset couvre l'ensemble des techniques modernes de prompt engineering, les benchmarks et metriques d'evaluation de LLM, ainsi que les strategies de gestion de la fenetre de… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/prompt-engineering-fr.tabularquestion-answeringn<1K0 likes27 downloads7mo agoHugging Face21AYI-NEDJIMI /prompt-engineering-en Prompt Engineering EN - Techniques, Evaluation & Context Management Comprehensive bilingual dataset on Prompt Engineering, LLM evaluation, and context window management. Created by AYI NEDJIMI Consultants - Expertise in Artificial Intelligence and Digital Transformation. Description This dataset covers all modern prompt engineering techniques, LLM evaluation benchmarks and metrics, and context window management strategies. It is based on three reference articles: Prompt… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/prompt-engineering-en.tabularquestion-answeringn<1K0 likes27 downloads7mo agoHugging Face22UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face23kingkaung /english_islamqainfo Dataset Card for English Islam QA Info Dataset Description The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice. Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.tabulartable-question-answering10K<n<100K5 likes24 downloads2y agoHugging Face24youjunhyeok /Magpie-Qwen2-Pro-200K-English-koTranslated Magpie-Align/Magpie-Qwen2-Pro-200K-English using nayohan/llama3-instrucTrans-enko-8b. For this dataset, we only used data that is 5000 characters or less in length and has language of English. Thanks for @Magpie-Align and @nayohan. @misc{xu2024magpie, title={Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing}, author={Zhangchen Xu and Fengqing Jiang and Luyao Niu and Yuntian Deng and Radha Poovendran and Yejin Choi and Bill Yuchen Lin}… See the full description on the dataset page: https://huggingface.co/datasets/youjunhyeok/Magpie-Qwen2-Pro-200K-English-ko.tabulartext-generation100K<n<1M1 likes22 downloads2y agoHugging Face25kingkaung /Quran_English_Myanmar_Parrelel_Corpus Quran English-Myanmar Parallel Corpus Description This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers. English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali. Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.tabulartranslation1K<n<10K0 likes22 downloads2y agoHugging Face26ChaoticEconomist /EnglishtoFrench-Translation-Dataset English–French Translation Dataset (SFT / LoRA Ready) A clean, structured dataset of 50,000 English–French sentence pairs designed for supervised fine-tuning (SFT) of large language models, LoRA adapters, and general machine translation tasks. Overview Property Value Language pair English → French Total rows 50,000 Train split 45,000 (90%) Validation split 2,500 (5%) Test split 2,500 (5%) Format CSV (Alpaca-style prompt format) License CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.tabulartranslation10K<n<100K0 likes19 downloads4mo agoHugging Face27nlpctx /tamil-english-corpus Tamil-English Retrieval Corpus A high-quality multilingual retrieval corpus constructed from the Mozhi Tamil Corpus and machine-translated into English using IndicTrans2. Dataset Summary This dataset contains Tamil documents paired with English translations. The corpus was created by filtering high-quality documents from the Mozhi Tamil Corpus and translating them using AI4Bharat's IndicTrans2 translation model. The resulting corpus is intended to support:… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tamil-english-corpus.tabulartext-retrieval1K<n<10K0 likes19 downloads3mo agoHugging Face28PhillyMac /Motivation_Employee_Engagement_Content_1 Motivation Employee Engagement Content 1 This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications. Dataset Structure Each record contains: text: The content text source_url: Original source URL source_title: Title of the source document source_domain: Domain of the source license_type: License classification (e.g. public_domain, cc_by, cc_by_sa) attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Motivation_Employee_Engagement_Content_1.tabulartext-generation1K<n<10K0 likes18 downloads6mo agoHugging Face29UMCU /apollo_english_guidelines_translated_to_dutch_with_marianmt Data description Apollo corpus, English guidelines translated to Dutch using MariaNMT. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes17 downloads2y agoHugging Face30miscovery /arabic_egypt_english_world_facts 🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized) The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.tabularquestion-answering10K<n<100K13 likes17 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.