datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.nuclear_eng_HW_dataset
Automated Grading of Handwritten STEM Homework: A Five-Stage Pipeline for Nuclear Engineering
📄 Read the full paper (PDF)
Contents
automatic_homework_grader_nuclear.pdf — Full technical report describing the five-stage pipeline
data/grading_records.parquet — Structured grading records for each student submission
data/results.parquet — Evaluation results and error taxonomy
LLM Grading of Handwritten STEM Homework (Nuclear Physics)
Per-question… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/nuclear_eng_HW_dataset.DBNL-public-qa-english-translationlilium_albanicum_eng_alb
Lilium Albanicum Eng-Alb
Task Categories:
Translation
Question-Answering
Conversational
Languages: English (en), Albanian (sq)
Size Categories: 100K < n < 1M
Dataset Card for "Lilium Albanicum"
Dataset Summary
The Lilium Albanicum dataset is a comprehensive English-Albanian and Albanian-English parallel corpus. The dataset includes original translations and extended synthetic Q&A pairs, which are designed to support and optimize LLM translation… See the full description on the dataset page: https://huggingface.co/datasets/noxneural/lilium_albanicum_eng_alb.Website_Traffic_and_EngagementGeneral_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.prompt-engineering-fr
Prompt Engineering FR - Techniques, Evaluation et Gestion du Contexte
Dataset bilingue complet sur le Prompt Engineering, l'evaluation de LLM et la gestion de la fenetre de contexte.
Cree par AYI NEDJIMI Consultants - Expertise en Intelligence Artificielle et Transformation Digitale.
Description
Ce dataset couvre l'ensemble des techniques modernes de prompt engineering, les benchmarks et metriques d'evaluation de LLM, ainsi que les strategies de gestion de la fenetre de… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/prompt-engineering-fr.prompt-engineering-en
Prompt Engineering EN - Techniques, Evaluation & Context Management
Comprehensive bilingual dataset on Prompt Engineering, LLM evaluation, and context window management.
Created by AYI NEDJIMI Consultants - Expertise in Artificial Intelligence and Digital Transformation.
Description
This dataset covers all modern prompt engineering techniques, LLM evaluation benchmarks and metrics, and context window management strategies. It is based on three reference articles:
Prompt… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/prompt-engineering-en.SlimOrca-Dedup-English-UzbekThis is an Uzbek translated version of https://huggingface.co/datasets/Open-Orca/SlimOrca-Dedup.
It is a single parquet file.
Check here for cleaned Uzbek only slim Orca dataset: https://huggingface.co/datasets/MLDataScientist/SlimOrca-Dedup-Uzbek-cleaned
tamil-english-corpus
Tamil-English Retrieval Corpus
A high-quality multilingual retrieval corpus constructed from the Mozhi Tamil Corpus and machine-translated into English using IndicTrans2.
Dataset Summary
This dataset contains Tamil documents paired with English translations.
The corpus was created by filtering high-quality documents from the Mozhi Tamil Corpus and translating them using AI4Bharat's IndicTrans2 translation model.
The resulting corpus is intended to support:… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/tamil-english-corpus.arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.engram-eval
engram evaluation data
The evaluation data behind Typed Decisions in Agent Memory: Where They Help, Where They Don't, and What It Costs
(Rishabh Sharma, 2026, doi:10.5281/zenodo.22948964; version 1: doi:10.5281/zenodo.22941758): update sets
that extend LoCoMo with fact changes, labeled contradiction pairs, the relation decisions escalated to an LLM,
and every scored answer from the paper's runs. Code: the engram repository (bench/make_hf_dataset.py builds this
directory from the… See the full description on the dataset page: https://huggingface.co/datasets/ris3abh-11/engram-eval.
