datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dr-saeid-ghezelbaash-entity-data
Dr. Saeed Ghezelbash Public Knowledge Graph
A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval.
The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.EnokiQA
EnokiQA
EnokiQA is an annotated dataset for fine-grained hallucination detection in long-form question answering. Each example contains a factual question, a no-context LLM answer, the full Wikipedia article used as verification evidence, sentence-grouped factual triples, and per-triple NLI and hallucination probabilities.
The dataset is dual-granularity: every hallucination label is attached to a claim (an extracted triple) and projected to a character span of the answer. The… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/EnokiQA.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.entailmentbank
EntailmentBank
EntailmentBank is a dataset of multistep entailment trees for open-domain science question answering. Each example links a question and answer to a structured proof: a tree of multi-premise entailment steps from known facts, through intermediate conclusions, to a hypothesis.
This repository contains the EMNLP 2021 v2 release in JSONL format, with four configs:
Config
Description
task1
Generate an entailment-tree proof from gold supporting facts
task2… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/entailmentbank.enwikivoyage-retrieval-202605
English Wikivoyage Retrieval 2026-05
I like travel datasets because they are about real places, real constraints,
and the small practical questions people ask before they go somewhere. I am
sharing this English Wikivoyage retrieval corpus in that spirit: as honest
work from a researcher-builder who wants to explore the world, make the
pipeline inspectable, and let other people reuse or challenge the choices.
This is not presented as a finished travel product or a private… See the full description on the dataset page: https://huggingface.co/datasets/vincentss/enwikivoyage-retrieval-202605.rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
gpqa-open-ended
GPQA Open-Ended
An open-ended reformulation of the GPQA (Graduate-Level Google-Proof Questions) benchmark. All 546 questions have been converted from multiple-choice to free-response format, preserving the original difficulty and domain expertise requirements while removing the ability to eliminate answers or pattern-match against option structure.
Why open-ended?
MCQ benchmarks have a ceiling problem for scalable oversight research: a non-expert judge who cannot solve… See the full description on the dataset page: https://huggingface.co/datasets/joanvelja/gpqa-open-ended.granola-entity-questions
GRANOLA Entity Questions Dataset Card
Dataset details
Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions)
Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.sunnah_ar_en_dataset
Dataset Card for Dataset Name
Dataset Card for Hadiths 14 Books Collection
This dataset contains a comprehensive bilingual (Arabic-English) collection of hadiths from 14 major authenticated books of Islamic tradition. It includes over 50762 narrations with complete metadata, organized by book, chapter, and narrator. Each hadith is presented in both its original Arabic text and English translation, making it ideal for cross-lingual NLP tasks, Islamic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/gurgutan/sunnah_ar_en_dataset.dfm11-fineinstructions-en
DFM11 FineInstructions English
This dataset contains 500,000 grounded English instruction/answer pairs and a
corresponding 500,000 grounded multi-turn conversations generated with the
FineInstructions pipeline. Every chat has a pair_id matching exactly one row
in the pairs configuration. The chats configuration contains 183,978
legacy continuations and 316,022 continuations balanced over 20 controlled
interaction modes.
Grounding documents came from ten filtered Common Pile… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm11-fineinstructions-en.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.ru-en-code-curriculum
RuEn Code Curriculum
RuEn Code Curriculum is a curated Russian-English dataset for continued pretraining (CPT) and supervised fine-tuning (SFT) of small code-oriented language models.
This public release contains only records classified as redistributable. Local-training-only web and code sources used by the internal curriculum are intentionally excluded.
Dataset summary
Configuration
Split
Records
Tokens
sft
train
53,278
13,997,239
sft
reserve
19,225… See the full description on the dataset page: https://huggingface.co/datasets/sup2ch/ru-en-code-curriculum.code-environnement
Code de l'environnement, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-environnement.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.X-SVAMP_en_zh_ko_it_es
X-SVAMP
🤗 Paper | 📖 arXiv
Dataset Description
X-SVAMP is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish).
It is intended to evaluate the math reasoning abilities of LLMs. The dataset is translated by GPT-4-turbo from the original English-version SVAMP.
In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-SVAMP_en_zh_ko_it_es.enamed-2025-answers
ENAMED 2025 Answer Keys & Errata
Official answer key (Gabarito Oficial) and scoring metadata for INEP's ENAMED 2025 exam.
Dataset Summary
100 Answer Records matching IDs in gustavokch/enamed-2025.
Errata & Annulment Tracking: Includes is_annulled flag and errata notes.
Usage
from datasets import load_dataset
answers = load_dataset("gustavokch/enamed-2025-answers", split="test")
print(answers[0])
nuclear_eng_HW_dataset
Automated Grading of Handwritten STEM Homework: A Five-Stage Pipeline for Nuclear Engineering
📄 Read the full paper (PDF)
Contents
automatic_homework_grader_nuclear.pdf — Full technical report describing the five-stage pipeline
data/grading_records.parquet — Structured grading records for each student submission
data/results.parquet — Evaluation results and error taxonomy
LLM Grading of Handwritten STEM Homework (Nuclear Physics)
Per-question… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/nuclear_eng_HW_dataset.synthetic-enterprise-operations-pack
Solstice Synthetic Enterprise Operations Pack (Sample)
A curated synthetic internal company dataset spanning engineering, task systems, collaboration, CRM, support, incidents, documents, and account-health workflows. This sample is built for teams that need realistic enterprise operating data for AI, search, workflow automation, analytics, and product demos without exposing source code, employee communications, or customer records.
Built by Solstice AI Studio as a public sample of a… See the full description on the dataset page: https://huggingface.co/datasets/solsticestudioai/synthetic-enterprise-operations-pack.DBNL-public-qa-english-translationGlobal_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.lilium_albanicum_eng_alb
Lilium Albanicum Eng-Alb
Task Categories:
Translation
Question-Answering
Conversational
Languages: English (en), Albanian (sq)
Size Categories: 100K < n < 1M
Dataset Card for "Lilium Albanicum"
Dataset Summary
The Lilium Albanicum dataset is a comprehensive English-Albanian and Albanian-English parallel corpus. The dataset includes original translations and extended synthetic Q&A pairs, which are designed to support and optimize LLM translation… See the full description on the dataset page: https://huggingface.co/datasets/noxneural/lilium_albanicum_eng_alb.enamed-2025-booklet-2-answers
ENAMED 2025 (Caderno 2) Answer Keys & Official Gabarito Definitivo
Official final answer key (Gabarito Definitivo) and scoring metadata for INEP's ENAMED 2025 exam (Booklet 2 / Caderno 02).
Dataset Summary
100 Answer Records matching IDs in gustavokch/enamed-2025-booklet-2.
Gabarito Definitivo & Annulments: Incorporates all 10 official post-appeal question annulments (is_annulled: true, answer_index: -1).
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gustavokch/enamed-2025-booklet-2-answers.turkuaz-rag
Turkuaz-RAG: A Novel Turkish Multi-Context Retrieval Benchmark
Turkuaz-RAG is the first benchmark specifically created for evaluating multi-context retrieval tasks in Turkish. It addresses a major gap in low-resource language research by providing multi-context questions, answers, and corresponding contexts.
Description of Benchmark
Languages: Turkish
Size: ~2,500 triplets (question, contexts, answer)
Context Sources: Turkish news articles from MLSUM
Question Types:… See the full description on the dataset page: https://huggingface.co/datasets/eneSadi/turkuaz-rag.Website_Traffic_and_EngagementGeneral_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.EntropyMath-Gen-v1
EntropyMath-Generated-v1
EntropyMath-Generated-v1 is a quality-gated, statement-unique generated mathematical reasoning evaluation dataset. It contains 934 problems exported from the EntropyMath generation framework, each paired with a statement, an answer, a solution, a verification_code consistency-evidence field, lineage and provenance metadata, and content hashes.
This Hugging Face entry (huggingface.co/datasets/sgmlc1234/EntropyMath-Gen-v1) is the primary hosting location for… See the full description on the dataset page: https://huggingface.co/datasets/sgmlc1234/EntropyMath-Gen-v1.crypto-education-en-corpus
Crypto Education Corpus (EN)
A curated educational corpus about cryptocurrency and blockchain technology, designed for building and evaluating RAG (Retrieval-Augmented Generation) systems.
Source Distribution
Source
Documents
Share
iqwiki.com
2,379
68.2%
academy.binance.com
581
16.7%
kraken.com
183
5.2%
coinbase.com
125
3.6%
gemini.com
94
2.7%
ethereum.org
74
2.1%
investopedia.com
51
1.5%
Word Count Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/kskada/crypto-education-en-corpus.MMLU-ProX_EN_Cleaned
MMLU-ProX English Cleaned
Dataset Description
This is a cleaned version of the English subset from MMLU-ProX (arXiv:2503.10497),
a comprehensive multilingual benchmark for evaluating large language models. The original MMLU-ProX dataset
contains 11,829 questions across 29 languages, built on the English MMLU-Pro benchmark.
Why This Cleaned Version?
The original English subset of MMLU-ProX contained spacing issues where words were concatenated without
proper… See the full description on the dataset page: https://huggingface.co/datasets/ZQ-Dev/MMLU-ProX_EN_Cleaned.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
