datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.engsaf
Engineering Short Answer Feedback
A collection of real short-answer responses from engineering exams across multiple engineering domains.
Background
In recent years, there has been a growing interest in using Artificial Intelligence (AI) to automate student assessment in education.
Among different types of assessments, summative assessments play a crucial role in evaluating a student's understanding level of a course.
Such examinations often involve short-answer… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/engsaf.genz-to-english
GenZ-to-English Translation Dataset
A high-quality text-to-text dataset for translating Gen Z slang into clear, standard English.
The dataset is designed for training and evaluating language models that convert modern internet slang into natural, readable English while preserving the original meaning.
Overview
This dataset contains 300k++ curated translation pairs covering a wide range of contemporary internet slang.
It includes expressions commonly found across… See the full description on the dataset page: https://huggingface.co/datasets/Sankar-2910/genz-to-english.urdu-idioms-with-english-translationMath_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.Irish-English-Parallel-Collection
UCCIX's English-Irish Parallel Textual Corpus
Dataset Summary
This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR.
This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data.
Dataset Sources
Source
Description
Statistics
Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.cleaned-english-prompts
Cleaned English Prompts Dataset
Dataset Description
A cleaned dataset containing English prompts and their corresponding responses. This dataset is designed for training conversational AI models and language models.
Dataset Summary
Columns: Questions and Response
Language: English
Size: 1,000-10,000 examples
Format: CSV
Cleaning: Data has been processed and cleaned for training
Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/cleaned-english-prompts.csl-electrical-engineering
csl-electrical-engineering"
由CSL數據集分割出來的電機工程(Electrical Engineering)子集,提供簡繁兩種版本。
from datasets import load_dataset
dataset = load_dataset("p208p2002/csl-electrical-engineering","zh-cn")
dataset = load_dataset("p208p2002/csl-electrical-engineering","zh-tw")
english_islamqainfo
Dataset Card for English Islam QA Info
Dataset Description
The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice.
Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.Quran_English_Myanmar_Parrelel_Corpus
Quran English-Myanmar Parallel Corpus
Description
This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers.
English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali.
Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.EnglishtoFrench-Translation-Dataset
English–French Translation Dataset (SFT / LoRA Ready)
A clean, structured dataset of 50,000 English–French sentence pairs designed
for supervised fine-tuning (SFT) of large language models, LoRA adapters, and
general machine translation tasks.
Overview
Property
Value
Language pair
English → French
Total rows
50,000
Train split
45,000 (90%)
Validation split
2,500 (5%)
Test split
2,500 (5%)
Format
CSV (Alpaca-style prompt format)
License
CC… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/EnglishtoFrench-Translation-Dataset.arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.OpenSubtitles-Thai-English
OpenSubtitles-Thai-English: YouTube Subtitle Parallel Dataset (en-th)
ชุดข้อมูลนี้เป็นชุดข้อมูลแปลภาษาอังกฤษ-ไทย (en-th) ที่ได้จากซับไตเติล YouTube โดยผ่านกระบวนการ clean, dedup, และ alignment เพื่อให้เหมาะกับงาน NLP/ML เช่น การฝึกโมเดลแปลภาษา การสร้าง embedding หรือ fine-tune LLM
This dataset contains English-Thai (en-th) parallel sentences extracted from YouTube subtitles, cleaned, deduplicated, and aligned for NLP/ML tasks such as machine translation, embedding, or LLM… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/OpenSubtitles-Thai-English.english_karakalpak_parallel_corpus_v1
English-Karakalpak Parallel Corpus (en-kaa)
Dataset Description
English-Karakalpak Parallel Corpus is a high-quality dataset containing 10,441 aligned sentence pairs in English and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. The corpus utilizes the official… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_parallel_corpus_v1.scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning English, this dataset contains grammar, syntax, spelling and punctuation rules, as well as English words with Russian translations.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_english_russian_dictionary_grammar_spelling_vectorized.english_luo_sentence_pair_dataset
English_Luo_sentence_pair_dataset
31,055 pair sentences (verses) extracted from the English and Luo Bibles.
This dataset has not been verified by any Luo or any person who knows and understands both English & Luo languages.
This dataset has not been audited/cleaned very well and may still contain some noise.
English-Marathi_Evaluation
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Devavrat28/English-Marathi_Evaluation.engd_researchesenglish_karakalpak_pairs_parallel_corpus_v2_8907
English-Karakalpak Parallel Corpus v2 (8.9K)
Dataset Description
English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This resource… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.
