datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ARCHIVE-TEXT-URLS
Internet Archive English Text URLs Dataset
Dataset Description
This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials.
Dataset Summary
Total Rows: 11,151,637
Language: English
Source: Internet Archive
Format: CSV with metadata and direct text file URLs
Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.nasa-science-repos-sme-benchmark
NASA Science Repos SME Benchmark
A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments.
Dataset Structure
Files
├── corpus.jsonl # 5,264 repositories with full metadata
├── queries.jsonl # 219 expert queries
└── qrels/
├── earth.tsv # Earth Science relevance judgments (162)
├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.french_narrativeqa
Description
Dataframe containing 143 French books in txt format.More precisely :
the texte column contains the texts
the titre column contains the book title
the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day)
the question column contains a single question asked about the associated text
the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.romanian-name-days
Romanian Name Days and Holidays
Zile onomastice și sărbători românești — the Romanian name-day calendar as
structured data.
In Romania, ziua onomastică — the feast day of the saint whose name you bear —
is widely celebrated, often more than a birthday. Until now this information
existed online only as HTML pages built for human readers. This is the
machine-readable version.
Published by trends.ro.
Dataset summary
Names
86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.aurora-think-1.5dataset_description:
"Aurora Think 1.5 is a meticulously crafted dataset containing a vast collection of questions and answers spanning a wide range of domains, including world knowledge, history, science, technology, philosophy, and more. It is specifically designed to be used for fine-tuning large language models (LLMs) to enhance their ability to understand and respond to complex, knowledge-intensive queries.
Key Features:
Extensive Coverage: The dataset encompasses a broad spectrum of… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/aurora-think-1.5.napolab
🌎 Natural Portuguese Language Benchmark (Napolab)
The Napolab is your go-to collection of Portuguese datasets for the evaluation of Large Language Models.
📊 Napolab for Large Language Models (LLMs)
A format of Napolab specifically designed for researchers experimenting with Large Language Models (LLMs) is now available. This format includes two main fields:
Prompt: The input prompt to be fed into the LLM.
Answer: The expected classification output label from the LLM… See the full description on the dataset page: https://huggingface.co/datasets/ruanchaves/napolab.Re-Auto-30K
Re-Auto-30K: A Comprehensive AI Safety Evaluation Dataset for Code Generation
Dataset Overview
Re-Auto-30K is a meticulously curated dataset containing 30,886 security-focused prompts designed specifically for evaluating AI safety in code generation scenarios. This dataset serves as a comprehensive benchmark for assessing Large Language Models (LLMs) across multiple dimensions of security, reliability, and autonomous behavior in software engineering contexts.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/navneetsatyamkumar/Re-Auto-30K.WinoBias-UK-Natural
WinoBias-UK Natural
WinoBias-UK Natural is a Ukrainian gender-counterfactual coreference evaluation set derived from
WinoBias. It provides natural masculine, feminine, mixed,
and cross-reference variants while preserving the source event and participant roles.
Current release
This preview contains 279 validated WinoBias source pairs and 1,674 Ukrainian variants. The current
release covers the validation Type 1 stratum. Full validation and test coverage is in… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoBias-UK-Natural.naija-pidgin-health-qa-rivers-2026theogonos-mirror-test
Theogonos Mirror Test
A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position.
Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness.
Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.refugiados_qa
Filtered Spanish Instruction Question-Answering Legal Refugiados
Dataset Description
Filtered Spanish Instruction Question-Answering Legal Refugiados is a collection of instruction queries filtered from the dataset at edumunozsala/instruct-legal-refugiados-es and split into train and test.
Dataset Summary
Compuesto por unos 10.326 registros que contienen los campos:
instrucción: una instrucción o consulta.
input: un contexto para resolver la consulta.
salida:… See the full description on the dataset page: https://huggingface.co/datasets/narhim/refugiados_qa.pashto-mental-health-counseling-3k
🧠 Pashto Mental Health Counseling 3K
This dataset is a specialized collection of 3,000 conversational pairs focused on mental health counseling, translated and culturally adapted into Pashto. It is designed to train LLMs to provide empathetic, supportive, and culturally relevant responses in a therapeutic context.
🌟 Overview
Mental health resources in Pashto are scarce. This dataset aims to bridge that gap by providing high-quality counseling dialogues. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-counseling-3k.prompt-variations
Prompt Variations and LLM Responses
Prompt variants and model responses used to evaluate the
Stability-Generalization Score (SGS) across eleven LLMs (eight
open-source + three closed-source) on six QA / instruction benchmarks
under six families of stylistic perturbations.
Splits
split
rows
source dataset
truthful_qa
99,888
TruthfulQA
natural_questions
41,040
Natural Questions
alpaca
13,872
Alpaca
simpleqa_verified
13,872
SimpleQA Verified… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/prompt-variations.Aurora-Think-1.0pashto-eagle-1k-cot
Pashto-Eagle-1K-CoT Dataset
Overview
Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot.
This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.pashto-opus-5k-reasoning-max
🚀 Pashto OPUS 5K Reasoning Max
This dataset is a high-quality collection of 5,000 reasoning-focused pairs, derived from the OPUS corpus and enhanced for Pashto Language Models. It is specifically curated to push the boundaries of "Chain-of-Thought" (CoT) and logical deduction in the Pashto language.
🌟 Overview
While standard OPUS data is often used for simple translation, Pashto-OPUS-5K-Reasoning-Max takes it a step further by focusing on complex instructions and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-opus-5k-reasoning-max.FINEST
FINEST
This is the official repository of FINEST: Improving LLM Responses to Sensitive Topics\Through Fine-Grained Evaluation (EACL 2026 Findings).
Dataset
We release the FINEST dataset in two complementary configurations to support both reproducibility and further research on fine-grained evaluation of LLM responses to sensitive topics.
1. raw_responses
The raw_responses configuration contains the full set of questions and model-generated responses used as… See the full description on the dataset page: https://huggingface.co/datasets/nayeon212/FINEST.wasp-5k
Wasp-Lang
This is a synthetic dataset created by an amplify model trained on the Wasp programming language quick-start documentation. Better data coming soon.
aurora-think-tinypashto-mental-health-support
Pashto Mental Health Support Dataset
Overview
Pashto-Mental-Health-Support is a specialized conversational dataset consisting of 172 high-quality samples focused on mental health awareness, emotional support, and psychological well-being. This dataset is a localized and translated version of the heliosbrahma/mental_health_chatbot_dataset.
This repository marks a strategic expansion of the iPashto.ai ecosystem, moving from logical reasoning into the domain of Emotional… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-support.pashto-otter-cot
Pashto-Otter-CoT Dataset
Overview
Pashto-Otter-CoT is a first-of-its-kind dataset specifically designed to bring Chain-of-Thought (CoT) Reasoning capabilities to Pashto language models. This dataset is a translated and curated version of a subset of the brendan-gho/gemma4b_paraphrased_otter_cot.
This project is part of the iPashto.ai initiative, led by Nassim الله (nassimjp), aimed at creating high-quality linguistic resources for the Pashto language.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-otter-cot.dictionary-embeddingsEmbeddings generated from the model multi-qa-mpnet-base-dot-v1 being trained on MAKILINGDING/english_dictionary
pashto-dragon-1k-cot
Pashto-Dragon-1K-CoT Dataset
Overview
Pashto-Dragon-1K-CoT is a specialized reasoning dataset containing 1,000+ samples, meticulously translated into Pashto to facilitate the development of advanced Chain-of-Thought (CoT) capabilities in Pashto LLMs. This dataset is a high-quality derivative of the brendan-gho/qwen3b_paraphrased_dragon_cot.
This repository is a core component of the iPashto.ai mission to move beyond simple web-scraping and focus on "Verified Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dragon-1k-cot.pashto-qwen-1k-cot
Pashto-Qwen-1K-CoT Dataset
Overview
Pashto-Qwen-1K-CoT is a high-quality reasoning dataset consisting of 1,024 samples, specifically curated to enhance the Chain-of-Thought (CoT) capabilities of Pashto language models. This dataset is a translated version of a subset from brendan-gho/qwen3b_paraphrased_cat_cot.
By focusing on "Reasoning" rather than just "Information," this dataset helps models like Baran and Roshan develop logical thinking paths in the Pashto language.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-qwen-1k-cot.NextGenAI
