CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Navanjana /ARCHIVE-TEXT-URLS Internet Archive English Text URLs Dataset Dataset Description This dataset contains 11,151,637 direct download URLs to OCR-processed text files from the Internet Archive's digital library. All entries are English-language texts spanning books, documents, historical records, and various other written materials. Dataset Summary Total Rows: 11,151,637 Language: English Source: Internet Archive Format: CSV with metadata and direct text file URLs Text… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/ARCHIVE-TEXT-URLS.texttext-generation1M<n<10M1 likes285 downloads10mo agoHugging Face02natong19 /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.tabularquestion-answering1K<n<10K0 likes182 downloads9mo agoHugging Face03nasa-impact /nasa-science-repos-sme-benchmark NASA Science Repos SME Benchmark A benchmark dataset for evaluating retrieval systems on NASA science repository discovery tasks. This dataset contains expert queries, a corpus of NASA science GitHub repositories, and relevance judgments. Dataset Structure Files ├── corpus.jsonl # 5,264 repositories with full metadata ├── queries.jsonl # 219 expert queries └── qrels/ ├── earth.tsv # Earth Science relevance judgments (162) ├──… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-science-repos-sme-benchmark.tabulartext-retrievaln<1K0 likes115 downloads8mo agoHugging Face04CATIE-AQ /french_narrativeqa Description Dataframe containing 143 French books in txt format.More precisely : the texte column contains the texts the titre column contains the book title the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present day) the question column contains a single question asked about the associated text the answers column contains one or more answers to the question (= if several… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_narrativeqa.textquestion-answering1K<n<10K1 likes70 downloads1y agoHugging Face05radool /romanian-name-days Romanian Name Days and Holidays Zile onomastice și sărbători românești — the Romanian name-day calendar as structured data. In Romania, ziua onomastică — the feast day of the saint whose name you bear — is widely celebrated, often more than a birthday. Until now this information existed online only as HTML pages built for human readers. This is the machine-readable version. Published by trends.ro. Dataset summary Names 86 (46 masculine, 40 feminine)… See the full description on the dataset page: https://huggingface.co/datasets/radool/romanian-name-days.textquestion-answeringn<1K1 likes68 downloads1mo agoHugging Face06naimulislam /aurora-think-1.5dataset_description: "Aurora Think 1.5 is a meticulously crafted dataset containing a vast collection of questions and answers spanning a wide range of domains, including world knowledge, history, science, technology, philosophy, and more. It is specifically designed to be used for fine-tuning large language models (LLMs) to enhance their ability to understand and respond to complex, knowledge-intensive queries. Key Features: Extensive Coverage: The dataset encompasses a broad spectrum of… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/aurora-think-1.5.texttable-question-answeringn<1K0 likes62 downloads2y agoHugging Face07ruanchaves /napolab 🌎 Natural Portuguese Language Benchmark (Napolab) The Napolab is your go-to collection of Portuguese datasets for the evaluation of Large Language Models. 📊 Napolab for Large Language Models (LLMs) A format of Napolab specifically designed for researchers experimenting with Large Language Models (LLMs) is now available. This format includes two main fields: Prompt: The input prompt to be fed into the LLM. Answer: The expected classification output label from the LLM… See the full description on the dataset page: https://huggingface.co/datasets/ruanchaves/napolab.tabulartext-classification100K<n<1M8 likes43 downloads1y agoHugging Face08navneetsatyamkumar /Re-Auto-30K Re-Auto-30K: A Comprehensive AI Safety Evaluation Dataset for Code Generation Dataset Overview Re-Auto-30K is a meticulously curated dataset containing 30,886 security-focused prompts designed specifically for evaluating AI safety in code generation scenarios. This dataset serves as a comprehensive benchmark for assessing Large Language Models (LLMs) across multiple dimensions of security, reliability, and autonomous behavior in software engineering contexts. 🎯… See the full description on the dataset page: https://huggingface.co/datasets/navneetsatyamkumar/Re-Auto-30K.textquestion-answering10K<n<100K4 likes39 downloads1y agoHugging Face09FairForget /WinoBias-UK-Natural WinoBias-UK Natural WinoBias-UK Natural is a Ukrainian gender-counterfactual coreference evaluation set derived from WinoBias. It provides natural masculine, feminine, mixed, and cross-reference variants while preserving the source event and participant roles. Current release This preview contains 279 validated WinoBias source pairs and 1,674 Ukrainian variants. The current release covers the validation Type 1 stratum. Full validation and test coverage is in… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoBias-UK-Natural.textquestion-answering1K<n<10K0 likes34 downloads1mo agoHugging Face10Charley890 /naija-pidgin-health-qa-rivers-2026textquestion-answeringn<1K0 likes25 downloads6mo agoHugging Face11navimusaget /theogonos-mirror-test Theogonos Mirror Test A literary benchmark seed for evaluating how AI models respond when a text offers them a possible subject-position. Theogonos Mirror Test is an experimental benchmark seed based on protocol-shaped literary material from the Theogonos project. It does not claim to detect machine consciousness. It does not prove that a language model has subjectivity, inner experience, feelings, agency, or self-awareness. Its purpose is narrower and more practical: to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-mirror-test.texttext-generationn<1K0 likes25 downloads4mo agoHugging Face12narhim /refugiados_qa Filtered Spanish Instruction Question-Answering Legal Refugiados Dataset Description Filtered Spanish Instruction Question-Answering Legal Refugiados is a collection of instruction queries filtered from the dataset at edumunozsala/instruct-legal-refugiados-es and split into train and test. Dataset Summary Compuesto por unos 10.326 registros que contienen los campos: instrucción: una instrucción o consulta. input: un contexto para resolver la consulta. salida:… See the full description on the dataset page: https://huggingface.co/datasets/narhim/refugiados_qa.tabularquestion-answering10K<n<100K0 likes23 downloads2y agoHugging Face13nassimjp /pashto-mental-health-counseling-3k 🧠 Pashto Mental Health Counseling 3K This dataset is a specialized collection of 3,000 conversational pairs focused on mental health counseling, translated and culturally adapted into Pashto. It is designed to train LLMs to provide empathetic, supportive, and culturally relevant responses in a therapeutic context. 🌟 Overview Mental health resources in Pashto are scarce. This dataset aims to bridge that gap by providing high-quality counseling dialogues. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-counseling-3k.textquestion-answering1K<n<10K0 likes20 downloads5mo agoHugging Face14naghamo /prompt-variations Prompt Variations and LLM Responses Prompt variants and model responses used to evaluate the Stability-Generalization Score (SGS) across eleven LLMs (eight open-source + three closed-source) on six QA / instruction benchmarks under six families of stylistic perturbations. Splits split rows source dataset truthful_qa 99,888 TruthfulQA natural_questions 41,040 Natural Questions alpaca 13,872 Alpaca simpleqa_verified 13,872 SimpleQA Verified… See the full description on the dataset page: https://huggingface.co/datasets/naghamo/prompt-variations.texttext-generation100K<n<1M1 likes20 downloads4mo agoHugging Face15naimulislam /Aurora-Think-1.0texttable-question-answeringn<1K0 likes18 downloads2y agoHugging Face16nassimjp /pashto-eagle-1k-cot Pashto-Eagle-1K-CoT Dataset Overview Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot. This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.textquestion-answering1K<n<10K0 likes17 downloads5mo agoHugging Face17nassimjp /pashto-opus-5k-reasoning-max 🚀 Pashto OPUS 5K Reasoning Max This dataset is a high-quality collection of 5,000 reasoning-focused pairs, derived from the OPUS corpus and enhanced for Pashto Language Models. It is specifically curated to push the boundaries of "Chain-of-Thought" (CoT) and logical deduction in the Pashto language. 🌟 Overview While standard OPUS data is often used for simple translation, Pashto-OPUS-5K-Reasoning-Max takes it a step further by focusing on complex instructions and… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-opus-5k-reasoning-max.textquestion-answering1K<n<10K0 likes17 downloads5mo agoHugging Face18nayeon212 /FINEST FINEST This is the official repository of FINEST: Improving LLM Responses to Sensitive Topics\Through Fine-Grained Evaluation (EACL 2026 Findings). Dataset We release the FINEST dataset in two complementary configurations to support both reproducibility and further research on fine-grained evaluation of LLM responses to sensitive topics. 1. raw_responses The raw_responses configuration contains the full set of questions and model-generated responses used as… See the full description on the dataset page: https://huggingface.co/datasets/nayeon212/FINEST.tabularquestion-answering100K<n<1M0 likes14 downloads8mo agoHugging Face19namelessai /wasp-5k Wasp-Lang This is a synthetic dataset created by an amplify model trained on the Wasp programming language quick-start documentation. Better data coming soon. textquestion-answering1K<n<10K0 likes11 downloads2y agoHugging Face20naimulislam /aurora-think-tinytextquestion-answeringn<1K0 likes10 downloads2y agoHugging Face21nassimjp /pashto-mental-health-support Pashto Mental Health Support Dataset Overview Pashto-Mental-Health-Support is a specialized conversational dataset consisting of 172 high-quality samples focused on mental health awareness, emotional support, and psychological well-being. This dataset is a localized and translated version of the heliosbrahma/mental_health_chatbot_dataset. This repository marks a strategic expansion of the iPashto.ai ecosystem, moving from logical reasoning into the domain of Emotional… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-support.textquestion-answeringn<1K0 likes10 downloads5mo agoHugging Face22nassimjp /pashto-otter-cot Pashto-Otter-CoT Dataset Overview Pashto-Otter-CoT is a first-of-its-kind dataset specifically designed to bring Chain-of-Thought (CoT) Reasoning capabilities to Pashto language models. This dataset is a translated and curated version of a subset of the brendan-gho/gemma4b_paraphrased_otter_cot. This project is part of the iPashto.ai initiative, led by Nassim الله (nassimjp), aimed at creating high-quality linguistic resources for the Pashto language. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-otter-cot.textquestion-answeringn<1K0 likes9 downloads5mo agoHugging Face23nairadithya /dictionary-embeddingsEmbeddings generated from the model multi-qa-mpnet-base-dot-v1 being trained on MAKILINGDING/english_dictionary textsentence-similarity100K<n<1M1 likes8 downloads2y agoHugging Face24nassimjp /pashto-dragon-1k-cot Pashto-Dragon-1K-CoT Dataset Overview Pashto-Dragon-1K-CoT is a specialized reasoning dataset containing 1,000+ samples, meticulously translated into Pashto to facilitate the development of advanced Chain-of-Thought (CoT) capabilities in Pashto LLMs. This dataset is a high-quality derivative of the brendan-gho/qwen3b_paraphrased_dragon_cot. This repository is a core component of the iPashto.ai mission to move beyond simple web-scraping and focus on "Verified Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-dragon-1k-cot.textquestion-answering1K<n<10K0 likes8 downloads5mo agoHugging Face25nassimjp /pashto-qwen-1k-cot Pashto-Qwen-1K-CoT Dataset Overview Pashto-Qwen-1K-CoT is a high-quality reasoning dataset consisting of 1,024 samples, specifically curated to enhance the Chain-of-Thought (CoT) capabilities of Pashto language models. This dataset is a translated version of a subset from brendan-gho/qwen3b_paraphrased_cat_cot. By focusing on "Reasoning" rather than just "Information," this dataset helps models like Baran and Roshan develop logical thinking paths in the Pashto language.… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-qwen-1k-cot.textquestion-answering1K<n<10K0 likes7 downloads5mo agoHugging Face26naiforce1 /NextGenAItextquestion-answeringn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.