CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nthakur /swim-ir-cross-lingual Dataset Card for SWIM-IR (Cross-lingual) This is the cross-lingual subset of the SWIM-IR dataset, where the query generated is in the target language and the passage is in English. The SWIM-IR dataset is available as CC-BY-SA 4.0. 18 languages (including English) are available in the cross-lingual dataset. For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website. What is SWIM-IR? SWIM-IR dataset is a synthetic multilingual… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-cross-lingual.texttext-retrieval10M<n<100M9 likes900 downloads2y agoHugging Face02piercewetter3 /irs-990-parsed IRS 990 Parsed Nonprofit Database Public relational extract of IRS Form 990 / 990-EZ / 990-PF filings, plus the colocated public files we join for address research: CMS NPPES + T-MSIS Medicare spend, FMCSA DOT carriers, OFAC SDN, FEC committees, and the IRS EO BMF. Generated: 2026-08-17Tables: 34Rows (sum): 459,069,505License: CC0 / public domain — derived from U.S. government recordsHub: https://huggingface.co/datasets/piercewetter3/irs-990-parsed Layout Tables… See the full description on the dataset page: https://huggingface.co/datasets/piercewetter3/irs-990-parsed.tabulartabular-classification100M<n<1B0 likes879 downloads1mo agoHugging Face03PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes634 downloads1y agoHugging Face04nthakur /swim-ir-monolingual Dataset Card for SWIM-IR (Monolingual) This is the monolingual subset of the SWIM-IR dataset, where the query generated and the passage are both in the same language. A few remaining languages will be added in the upcoming v2 version of SWIM-IR. The dataset is available as CC-BY-SA 4.0. For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website. What is SWIM-IR? SWIM-IR dataset is a synthetic multilingual retrieval dataset… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/swim-ir-monolingual.texttext-retrieval1M<n<10M10 likes361 downloads2y agoHugging Face05nthakur /indic-swim-ir-cross-lingual Dataset Card for Indic SWIM-IR (Cross-lingual) This is the cross-lingual Indic subset of the SWIM-IR dataset, where the query generated is in the Indo-European language and the passage is in English. The SWIM-IR dataset is available as CC-BY-SA 4.0. 18 languages (including English) are available in the cross-lingual dataset. For full details of the dataset, please read our upcoming NAACL 2024 paper and check out our website. What is SWIM-IR? SWIM-IR dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/nthakur/indic-swim-ir-cross-lingual.texttext-retrieval10K<n<100K2 likes358 downloads2y agoHugging Face06rmoham05 /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.textquestion-answering100K<n<1M0 likes251 downloads2mo agoHugging Face07ai-earth /Earth-Iron (ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs     Updates/News 🆕 🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉. Abstract Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated… See the full description on the dataset page: https://huggingface.co/datasets/ai-earth/Earth-Iron.textquestion-answering1K<n<10K0 likes137 downloads7mo agoHugging Face08Ireliya /hierarchical-geospatial-reasoningimagequestion-answeringn<1K2 likes104 downloads6mo agoHugging Face09nawaralseelawi /mizan-iraqi-arabic-benchmark Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes. 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865 🏆 Live leaderboard: https://mizan-bench.onrender.com 💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.textquestion-answeringn<1K1 likes74 downloads13d agoHugging Face10Tevatron /docmatix-ir Docmatix-IR Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering. To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.textquestion-answering1M<n<10M15 likes73 downloads2y agoHugging Face11nasa-impact /nasa-sde-IR-benchmark-20251024-v5 NASA SDE IR Benchmark v5 A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation. Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST. Code: NASA-IMPACT/st-training-workflow Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.texttext-retrieval100K<n<1M1 likes65 downloads4mo agoHugging Face12McGill-NLP /AdvBench-IR Exploiting Instruction-Following Retrievers for Malicious Information Retrieval This dataset includes malicious documents in response to AdvBench (Zou et al., 2023) queries. We have generated these documents using the Mistral-7B-Instruct-v0.2 language model. from datasets import load_dataset import transformers ds = load_dataset("McGill-NLP/AdvBench-IR", split="train") # Loads LlaMAGuard model to check the safety of the samples model_name = "meta-llama/Llama-Guard-3-1B" model =… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/AdvBench-IR.textquestion-answeringn<1K4 likes61 downloads2y agoHugging Face13AdaptKey /ustax-irc-qa-89k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.textquestion-answering10K<n<100K0 likes44 downloads6mo agoHugging Face14sosa123454321 /iran-turkiye-startup-landing-kb Iran → Türkiye Startup Landing — legal/migration knowledge base One dataset for the whole platform (rule: one dataset, one space, one vector index — never several). Nightly snapshots produced by rag/scrape.py: news, academic, official (göç idaresi / ministry), legislation and directory sources about Iranian founders landing startups in Türkiye. All PII is scrubbed at fetch time (rag/fetch.scrub_pii). Chunks are stored as JSONL per snapshot day under data/kb/<date>/. This… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/iran-turkiye-startup-landing-kb.textquestion-answeringn<1K0 likes43 downloads11d agoHugging Face15irioder /littleHermione-benchmark Dataset card: O.W.L. & N.E.W.T. Bench development set v0.4 Summary Version 0.4 is a reviewed development benchmark of 75 short-answer factual questions about the seven English-language Harry Potter novels. It contains two separately scored examinations: 30 challenging O.W.L. questions covering recurring book canon beyond famous entrance-level facts; 45 frontier N.E.W.T. questions covering chapter-level prose details, minor names, precise objects, prices, and… See the full description on the dataset page: https://huggingface.co/datasets/irioder/littleHermione-benchmark.textquestion-answeringn<1K0 likes42 downloads17d agoHugging Face16PrismaX /Earth-Iron Dataset Card for Earth-Iron Dataset Details Dataset Description Earth-Iron is a comprehensive question answering (QA) benchmark designed to evaluate the fundamental scientific exploration abilities of large language models (LLMs) within the Earth sciences. It features a substantial number of questions covering a wide range of topics and tasks crucial for basic understanding in this domain. This dataset aims to assess the foundational knowledge that underpins… See the full description on the dataset page: https://huggingface.co/datasets/PrismaX/Earth-Iron.textquestion-answering1K<n<10K0 likes39 downloads1y agoHugging Face17AdaptKey /ustax-irc-qa-36k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.textquestion-answering10K<n<100K0 likes37 downloads6mo agoHugging Face18qdai /Accel-IR Accel-IR Benchmark: A Gold Standard for Particle Accelerator Physics This repository contains the Accel-IR Benchmark, a domain-specific Information Retrieval (IR) dataset for particle accelerator physics. It was developed as part of the Master's Thesis "From Dataset to Optimization: A Benchmarking Framework for Information Retrieval in the Particle Accelerator Domain" by Qing Dai (University of Zurich, 2025), in collaboration with the Paul Scherrer Institute (PSI).… See the full description on the dataset page: https://huggingface.co/datasets/qdai/Accel-IR.tabulartext-retrieval1K<n<10K0 likes36 downloads10mo agoHugging Face19Amir7440 /IRAN-MADANI-LAWtextquestion-answeringn<1K1 likes34 downloads1y agoHugging Face20Irza /Arxiv_ph_indonesiatextquestion-answering1K<n<10K3 likes32 downloads3y agoHugging Face21Irfanuruchi /dsp-fft-sampling-aliasing Synthetic DSP Dataset: FFT + Sampling / Aliasing This repository contains synthetic instruction-style DSP samples designed for numerical reasoning and conceptual understanding of Digital Signal Processing (DSP) fundamentals. The dataset focuses on: FFT bin reasoning and frequency-domain interpretation Sampling theory Aliasing effects Dataset Origin & Verification This dataset was generated as part of the project: Fine-Tuning Lightweight Large Language Models for a… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/dsp-fft-sampling-aliasing.texttext-generation1K<n<10K0 likes30 downloads8mo agoHugging Face22liamirali /iranian-elderly-psychospiritual-interviews Iranian Elderly Psycho-Spiritual Interviews A culturally grounded, fully synthetic conversational interview dataset for assessing the mental and spiritual health of Iranian older adults, generated using Large Language Models. Dataset Summary This dataset introduces a culturally grounded, fully synthetic conversational interview corpus designed for the assessment and analysis of mental and spiritual health among Iranian older adults. All interviews are conducted in… See the full description on the dataset page: https://huggingface.co/datasets/liamirali/iranian-elderly-psychospiritual-interviews.tabulartext-generation1K<n<10K0 likes30 downloads5mo agoHugging Face23Pangeanic /Iraqi-Arabic-multidomain-QA-text Iraqi Arabic Multidomain QA Dataset The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models. This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.textquestion-answeringn<1K1 likes27 downloads4mo agoHugging Face24Irina-Na /AutenticHadithestextquestion-answering1K<n<10K2 likes26 downloads1y agoHugging Face25Adel-Elwan /Artificial-intelligence-dataset-for-IR-systems Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards information-retrieval semantic-search Languages English Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.tabularquestion-answering100K<n<1M0 likes25 downloads3y agoHugging Face26iradukunda-dev /offences_and_penalties_in_general_2018_datasettexttext-classificationn<1K0 likes25 downloads9mo agoHugging Face27SalahALHaismawi /uae-laws-irac UAE Laws Q&A Dataset (IRAC Format) A high-quality dataset of 9,477 question-answer pairs about UAE laws, formatted in IRAC (Issue, Rule, Application, Conclusion) legal reasoning structure. Dataset Creation Source Documents The dataset was built from a comprehensive collection of UAE legal documents, including: Federal Decrees and Laws Cabinet Resolutions Ministerial Decisions Civil and Commercial Codes Labor Law Traffic Law And more Creation Process… See the full description on the dataset page: https://huggingface.co/datasets/SalahALHaismawi/uae-laws-irac.textquestion-answering1K<n<10K1 likes24 downloads8mo agoHugging Face28JasonChen91 /Earth-Iron (ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs &nbsp; &nbsp; Updates/News 🆕 🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉. Abstract Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or… See the full description on the dataset page: https://huggingface.co/datasets/JasonChen91/Earth-Iron.textquestion-answering1K<n<10K0 likes21 downloads2mo agoHugging Face29sinamahallati /Iranian-Court-Rulings-25Kgated Iranian Court Rulings Dataset A structured Persian-language corpus containing 25,617 Iranian judicial rulings collected from publicly accessible pages of the Iranian National Judicial Opinions database. The dataset is intended for research and development in Persian Legal NLP, Information Retrieval, Retrieval-Augmented Generation (RAG), semantic search, legal document understanding, and related areas. Dataset Overview Number of records: 25,617 Language: Persian… See the full description on the dataset page: https://huggingface.co/datasets/sinamahallati/Iranian-Court-Rulings-25K.texttext-retrieval10K<n<100K0 likes21 downloads15d agoHugging Face30irahulpandey /NvidiaDocumentationQandApairs-llama2textquestion-answering1K<n<10K2 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.