CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shareAI /ShareGPT-Chinese-English-90k ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss) Features: Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.question-answering10K<n<100K288 likes3k downloads9mo agoHugging Face02yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes992 downloads7mo agoHugging Face03Joemn /TamilNadu-State-board-books-2025-Tamil-and-english-versions Tamil Nadu School Textbooks — Tamil and English Structured Text This dataset contains text extracted from 312 Tamil Nadu State Board school textbooks for Standards 1–12. It covers Tamil- and English-medium books and provides each retained book in two forms: structured JSON with book metadata, ordered sections, typed content blocks, source references, extraction statistics, and curation provenance; Markdown for reading, inspection, and downstream text processing. The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.texttext-retrievaln<1K1 likes574 downloads2mo agoHugging Face04mtimur /distill-gpt4-eng-chat Description Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases: either has non-english letters and special symbols either has http-links either has html blocks either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used ) textquestion-answering100K<n<1M2 likes534 downloads2y agoHugging Face05pandalin98 /EngineMT-QA EngineMT-QA Dataset Overview EngineMT-QA is a large-scale, multi-task, multimodal dataset for Time-Series Question Answering (Time-Series QA). It enables research on aligning multivariate time-series signals with natural language through four key cognitive tasks: Understanding Perception Reasoning Decision-Making The dataset is built on N-CMAPSS, simulating real-world aero-engine operational and maintenance scenarios. It supports the development and evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/pandalin98/EngineMT-QA.question-answering100K<n<1M4 likes331 downloads1y agoHugging Face06KisanVaani /agriculture-qa-english-only Dataset Card for Dataset Name This dataset contains question-answer pairs related to agriculture. The dataset can be used for tasks such as question answering, information retrieval, and natural language understanding in the agricultural domain. The questions cover various aspects of agriculture, including crop production, animal husbandry, soil management, and farming practices. Dataset Details he dataset is structured as a collection of JSON files, with each file… See the full description on the dataset page: https://huggingface.co/datasets/KisanVaani/agriculture-qa-english-only.textquestion-answering10K<n<100K26 likes293 downloads2y agoHugging Face07NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes222 downloads2mo agoHugging Face08STEM-AI-mtl /Electrical-engineering To the electrical engineering community This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes. Authors STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec textquestion-answering1K<n<10K61 likes193 downloads2y agoHugging Face09Mxode /PolyDevTasks-Chinese_English_German PolyDevTasks: 多语言软件开发智能任务 💻 Github 仓库 简体中文 | English | Deutsch 介绍 我们发布了 PolyDevTasks,这是一个包含超过 38 万条真实编码任务指令的数据集,涵盖 3 种自然语言(中文、英文、德语)和 8 种编程语言(C、C#、C++、Go、Java、JavaScript、Python、Rust)。不同于翻译或模板化的数据集,每条指令都是独立编写的,体现了特定语言与生态的习惯用法(例如 Go 的并发、C# 的 LINQ、UNIX I/O),并强调智能体的行为特征,如工具使用、网络操作、文件处理和优雅退出。PolyDevTasks 专为训练和评测 Agents 与 LLMs 在端到端软件工作流和跨语言泛化能力上的表现而设计。 数据统计 📂 文件级统计(每个 NL/PL 文件) 文件 数目 instruction均长 response均长 zh/c.jsonl 24,590 78.90 6,374.07… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/PolyDevTasks-Chinese_English_German.texttext-generation100K<n<1M3 likes191 downloads1y agoHugging Face10IsmaelMousa /engsaf Engineering Short Answer Feedback A collection of real short-answer responses from engineering exams across multiple engineering domains. Background In recent years, there has been a growing interest in using Artificial Intelligence (AI) to automate student assessment in education. Among different types of assessments, summative assessments play a crucial role in evaluating a student's understanding level of a course. Such examinations often involve short-answer… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/engsaf.textquestion-answering1K<n<10K0 likes176 downloads6mo agoHugging Face11EngMuhammadAtef /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M1 likes162 downloads5mo agoHugging Face12jiangjiechen /ekar_english Dataset Card for ekar_english Dataset Summary New!(9/18/2022) E-KAR v1.1 is officially released (at the main branch), with a higher-quality English dataset! In v1.1, we further improve the Chinese-to-English translation quality of the English E-KAR, with over 600 problems and over 1,000 explanations manually adjusted. You can still find previous version (as in the paper) in the v1.0 branch in the repo. For more information please refer to… See the full description on the dataset page: https://huggingface.co/datasets/jiangjiechen/ekar_english.textquestion-answering1K<n<10K4 likes152 downloads4y agoHugging Face13whyamanbhardwaj /Scholarly-Epistemic-Engine Dataset Card for Scholarly-Epistemic-Engine: arXiv cs.AI Corpus and Embeddings This dataset contains the processed text, metadata, and semantic vector embeddings of approximately 90,000 scholarly articles from the arXiv Computer Science - Artificial Intelligence (cs.AI) category, spanning from 1993 to December 2024. It is designed to support Retrieval-Augmented Generation (RAG) systems and semantic knowledge discovery. Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/whyamanbhardwaj/Scholarly-Epistemic-Engine.textquestion-answering100K<n<1M2 likes141 downloads4mo agoHugging Face14KaifengGGG /WenYanWen_English_Parallel Dataset Card for WenYanWen_English_Parallel Dataset Summary The WenYanWen_English_Parallel dataset is a multilingual parallel corpus in Classical Chinese (Wenyanwen), modern Chinese, and English. The Classical Chinese and modern Chinese parts are sourced from the NiuTrans/Classical-Modern dataset, while the corresponding English translations are generated using Gemini Pro. Data Fields info: A string representing the title or source information of the text.… See the full description on the dataset page: https://huggingface.co/datasets/KaifengGGG/WenYanWen_English_Parallel.texttranslation1M<n<10M11 likes134 downloads2y agoHugging Face15TunstallTensor /unreal-engine-5.7-qagated Unreal Engine 5.7 Instruction-Tuning Dataset Dataset Description This dataset contains 122,199 high-quality, synthetic Question and Answer pairs specifically designed for instruction-tuning Large Language Models (LLMs) to become expert coding and architectural assistants for Unreal Engine 5.7. Because Unreal Engine frequently deprecates older APIs (from UE4 to UE5) and introduces massive paradigm shifts (like Nanite, Lumen, and World Partition), standard… See the full description on the dataset page: https://huggingface.co/datasets/TunstallTensor/unreal-engine-5.7-qa.textquestion-answering100K<n<1M19 likes111 downloads5mo agoHugging Face16CJJones /LLM_Electrical_Engineering_Educational_Synthetic_DialogDataset Card for LLM_Electrical_Engineering_Educational_Synthetic_Dialog The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. Dataset Description The LLM_Electrical_Engineering_Educational_Synthetic_Dialog dataset contains AI-generated conversational interactions designed for training large language models in electrical engineering education. This synthetic dialogue corpus simulates tutor-student… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/LLM_Electrical_Engineering_Educational_Synthetic_Dialog.textquestion-answering10K<n<100K3 likes108 downloads7mo agoHugging Face17shreyansh12183 /shreyansh-hinglish-english-stem-500k 🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations. 📖 Overview In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.textquestion-answering100K<n<1M0 likes104 downloads3d agoHugging Face18atenglens /taiwanese_english_translationThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.question-answering4 likes97 downloads3y agoHugging Face19mznaser /ERI-Engineering-Reasoning-and-Instruction ERI — Engineering Reasoning and Instruction 53,429 instruction–response pairs spanning nine engineering disciplines, seven question types and three difficulty levels, for instruction-tuning and evaluating language models on engineering reasoning. from datasets import load_dataset eri = load_dataset("mznaser/ERI-Engineering-Reasoning-and-Instruction") civil = eri.filter(lambda r: r["field"] == "civil_engineering") Schema Every record has six fields: Field… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/ERI-Engineering-Reasoning-and-Instruction.texttext-generation10K<n<100K0 likes90 downloads1mo agoHugging Face20SonexaAI /Small-Life-Dataset-ru-eng Russian-English Dialogue Dataset 🎯 Overview A comprehensive bilingual dialogue dataset containing 50,000 high-quality question-answer pairs in Russian and English. The dataset is balanced across two main categories: programming/technical topics and general conversation. Dataset Statistics: 📊 Total Dialogues: 50,000 🇷🇺 Russian: 25,135 (50.3%) 🇬🇧 English: 24,865 (49.7%) 💻 Coding Topics: 25,056 (50.1%) 💬 General Conversation: 24,944 (49.9%) 📑… See the full description on the dataset page: https://huggingface.co/datasets/SonexaAI/Small-Life-Dataset-ru-eng.question-answering2 likes88 downloads7d agoHugging Face21laion /nemotron-terminal-software_engineering nemotron-terminal-software_engineering Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "software_engineering". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-software_engineering.textquestion-answering10K<n<100K0 likes81 downloads5mo agoHugging Face22miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes77 downloads1y agoHugging Face23InfoBayAI /English-STEM-QA-MCQ-DatasetgatedDataset Description: This dataset is a large-scale collection of English STEM Question Answering (QA) data, containing 4,388,206 question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, reasoning, problem-solving, and educational learning in English. The dataset consists of multiple-choice question answering (MCQA) samples across core STEM domains including Physics, Mathematics, Chemistry, Biology, and General… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-STEM-QA-MCQ-Dataset.textquestion-answeringn<1K0 likes67 downloads7d agoHugging Face24onkanat /sdr-engineers-dataset 📡 Software-Defined Radio (SDR) for Engineers Dataset A comprehensive, high-quality instruction-tuning, preference optimization (DPO), and RAG dataset compiled from the textbook "Software-Defined Radio for Engineers" (2018) by Travis F. Collins, Robin Getz, Di Pu, and Alexander M. Wyglinski (Artech House). This dataset covers Software-Defined Radio (SDR) architecture, digital signal processing (DSP) fundamentals, probability in communications, digital modulations (PAM, QAM, PSK)… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/sdr-engineers-dataset.texttext-generation1K<n<10K0 likes67 downloads2mo agoHugging Face25DuoNeural /ml-ai-engineer-sft DuoNeural ML/AI Engineer SFT Dataset A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments. Why this dataset exists Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.texttext-generation1K<n<10K1 likes66 downloads3mo agoHugging Face26mst-ai /nuclear_eng_HW_dataset Automated Grading of Handwritten STEM Homework: A Five-Stage Pipeline for Nuclear Engineering 📄 Read the full paper (PDF) Contents automatic_homework_grader_nuclear.pdf — Full technical report describing the five-stage pipeline data/grading_records.parquet — Structured grading records for each student submission data/results.parquet — Evaluation results and error taxonomy LLM Grading of Handwritten STEM Homework (Nuclear Physics) Per-question… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/nuclear_eng_HW_dataset.tabulartext-classification1K<n<10K0 likes62 downloads3mo agoHugging Face27noxneural /lilium_albanicum_eng_alb Lilium Albanicum Eng-Alb Task Categories: Translation Question-Answering Conversational Languages: English (en), Albanian (sq) Size Categories: 100K < n < 1M Dataset Card for "Lilium Albanicum" Dataset Summary The Lilium Albanicum dataset is a comprehensive English-Albanian and Albanian-English parallel corpus. The dataset includes original translations and extended synthetic Q&A pairs, which are designed to support and optimize LLM translation… See the full description on the dataset page: https://huggingface.co/datasets/noxneural/lilium_albanicum_eng_alb.tabulartranslation100K<n<1M4 likes56 downloads2y agoHugging Face28NLP-FBK /e3c-crf-english Dataset description Here we realease the dataset to perform the Case Report Forms filling task obtained from The European Clinical Case Corpus as described in the paper Converting Annotated Clinical Cases into Structured Case Report Forms presented at the BioNLP workshop at ACL 2025. The dataset is composed by a set patients with related clinical_note that describe their history and conditions. Each patient is uniquely identified by the document_id column. The task consists of… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/e3c-crf-english.textquestion-answering1K<n<10K0 likes56 downloads10mo agoHugging Face29NKODATALABS /patent-engineering-diagrams-preview Patent & Engineering Diagram Reasoning — Authentic Engineering Preview NKO Data Labs · design-partner market probe · target commercial release: 25k+ figure-linked examples A multimodal engineering-data concept linking technical figures to structured component/concept and description context for diagram understanding and technical reasoning. Preview status This repository now contains a small authentic source-backed engineering-diagram preview using selected U.S.… See the full description on the dataset page: https://huggingface.co/datasets/NKODATALABS/patent-engineering-diagrams-preview.question-answeringn<1K0 likes56 downloads1mo agoHugging Face30nscharrenberg /DBNL-public-qa-english-translationtabularquestion-answering10K<n<100K0 likes54 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.