CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01NoirZangetsu /Flutter-Code-with-Questions-Dataset-English 🧠 Flutter Code with Questions Dataset (English) This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development. 📂 Dataset Structure The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes: A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.textquestion-answering1K<n<10K3 likes222 downloads2mo agoHugging Face02IsmaelMousa /engsaf Engineering Short Answer Feedback A collection of real short-answer responses from engineering exams across multiple engineering domains. Background In recent years, there has been a growing interest in using Artificial Intelligence (AI) to automate student assessment in education. Among different types of assessments, summative assessments play a crucial role in evaluating a student's understanding level of a course. Such examinations often involve short-answer… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/engsaf.textquestion-answering1K<n<10K0 likes176 downloads6mo agoHugging Face03Sankar-2910 /genz-to-english GenZ-to-English Translation Dataset A high-quality text-to-text dataset for translating Gen Z slang into clear, standard English. The dataset is designed for training and evaluating language models that convert modern internet slang into natural, readable English while preserving the original meaning. Overview This dataset contains 300k++ curated translation pairs covering a wide range of contemporary internet slang. It includes expressions commonly found across… See the full description on the dataset page: https://huggingface.co/datasets/Sankar-2910/genz-to-english.texttranslation100K<n<1M1 likes138 downloads13d agoHugging Face04Ehtisham1328 /urdu-idioms-with-english-translationtexttranslation1K<n<10K5 likes77 downloads3y agoHugging Face05miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes77 downloads1y agoHugging Face06Programmer-RD-AI /sinhala-english-singlish-translation Sinhala–English–Singlish Translation Dataset A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations. 📋 Table of Contents Dataset Overview Installation Quick Start Dataset Structure Usage Examples Citation License Credits Dataset Overview Description: 34,500 aligned triplets of Sinhala (native script) English (human translation) Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.texttranslation10K<n<100K3 likes46 downloads1y agoHugging Face07miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes41 downloads1y agoHugging Face08ReliableAI /Irish-English-Parallel-Collection UCCIX's English-Irish Parallel Textual Corpus Dataset Summary This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR. This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data. Dataset Sources Source Description Statistics Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.texttext-generation10K<n<100K1 likes33 downloads2y agoHugging Face09Aipresso /cleaned-english-prompts Cleaned English Prompts Dataset Dataset Description A cleaned dataset containing English prompts and their corresponding responses. This dataset is designed for training conversational AI models and language models. Dataset Summary Columns: Questions and Response Language: English Size: 1,000-10,000 examples Format: CSV Cleaning: Data has been processed and cleaned for training Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/cleaned-english-prompts.textquestion-answering1M<n<10M0 likes29 downloads11mo agoHugging Face10p208p2002 /csl-electrical-engineering csl-electrical-engineering" 由CSL數據集分割出來的電機工程(Electrical Engineering)子集,提供簡繁兩種版本。 from datasets import load_dataset dataset = load_dataset("p208p2002/csl-electrical-engineering","zh-cn") dataset = load_dataset("p208p2002/csl-electrical-engineering","zh-tw") textsummarization10K<n<100K4 likes27 downloads3y agoHugging Face11kingkaung /english_islamqainfo Dataset Card for English Islam QA Info Dataset Description The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice. Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.tabulartable-question-answering10K<n<100K5 likes24 downloads2y agoHugging Face12kingkaung /Quran_English_Myanmar_Parrelel_Corpus Quran English-Myanmar Parallel Corpus Description This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers. English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali. Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.tabulartranslation1K<n<10K0 likes22 downloads2y agoHugging Face13ChaoticEconomist /EnglishtoFrench-Translation-Dataset English–French Translation Dataset (SFT / LoRA Ready) A clean, structured dataset of 50,000 English–French sentence pairs designed for supervised fine-tuning (SFT) of large language models, LoRA adapters, and general machine translation tasks. Overview Property Value Language pair English → French Total rows 50,000 Train split 45,000 (90%) Validation split 2,500 (5%) Test split 2,500 (5%) Format CSV (Alpaca-style prompt format) License CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.tabulartranslation10K<n<100K0 likes19 downloads4mo agoHugging Face14miscovery /arabic_egypt_english_world_facts 🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized) The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.tabularquestion-answering10K<n<100K13 likes17 downloads1y agoHugging Face15ZombitX64 /OpenSubtitles-Thai-English OpenSubtitles-Thai-English: YouTube Subtitle Parallel Dataset (en-th) ชุดข้อมูลนี้เป็นชุดข้อมูลแปลภาษาอังกฤษ-ไทย (en-th) ที่ได้จากซับไตเติล YouTube โดยผ่านกระบวนการ clean, dedup, และ alignment เพื่อให้เหมาะกับงาน NLP/ML เช่น การฝึกโมเดลแปลภาษา การสร้าง embedding หรือ fine-tune LLM This dataset contains English-Thai (en-th) parallel sentences extracted from YouTube subtitles, cleaned, deduplicated, and aligned for NLP/ML tasks such as machine translation, embedding, or LLM… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/OpenSubtitles-Thai-English.texttext-generation100K<n<1M0 likes12 downloads11mo agoHugging Face16bekan /english_karakalpak_parallel_corpus_v1 English-Karakalpak Parallel Corpus (en-kaa) Dataset Description English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.texttranslation10K<n<100K2 likes11 downloads10mo agoHugging Face17ScoutieAutoML /scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.tabulartext-classification10K<n<100K0 likes10 downloads2y agoHugging Face18KnoxDevelopers /english_luo_sentence_pair_dataset English_Luo_sentence_pair_dataset 31,055 pair sentences (verses) extracted from the English and Luo Bibles. This dataset has not been verified by any Luo or any person who knows and understands both English & Luo languages. This dataset has not been audited/cleaned very well and may still contain some noise. texttranslation10K<n<100K0 likes10 downloads2mo agoHugging Face19Devavrat28 /English-Marathi_Evaluation Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Devavrat28/English-Marathi_Evaluation.texttranslationn<1K0 likes9 downloads1y agoHugging Face20ahhany /engd_researchestexttext-generation10K<n<100K0 likes8 downloads3y agoHugging Face21bekan /english_karakalpak_pairs_parallel_corpus_v2_8907 English-Karakalpak Parallel Corpus v2 (8.9K) Dataset Description English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa). This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This resource… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.texttranslation1K<n<10K1 likes4 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.