CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes634 downloads1y agoHugging Face02rmoham05 /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.textquestion-answering100K<n<1M0 likes251 downloads2mo agoHugging Face03Ireliya /hierarchical-geospatial-reasoningimagequestion-answeringn<1K2 likes104 downloads6mo agoHugging Face04nawaralseelawi /mizan-iraqi-arabic-benchmark Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes. 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865 🏆 Live leaderboard: https://mizan-bench.onrender.com 💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.textquestion-answeringn<1K1 likes74 downloads13d agoHugging Face05Tevatron /docmatix-ir Docmatix-IR Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering. To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.textquestion-answering1M<n<10M15 likes73 downloads2y agoHugging Face06nasa-impact /nasa-sde-IR-benchmark-20251024-v5 NASA SDE IR Benchmark v5 A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation. Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST. Code: NASA-IMPACT/st-training-workflow Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.texttext-retrieval100K<n<1M1 likes65 downloads4mo agoHugging Face07McGill-NLP /AdvBench-IR Exploiting Instruction-Following Retrievers for Malicious Information Retrieval This dataset includes malicious documents in response to AdvBench (Zou et al., 2023) queries. We have generated these documents using the Mistral-7B-Instruct-v0.2 language model. from datasets import load_dataset import transformers ds = load_dataset("McGill-NLP/AdvBench-IR", split="train") # Loads LlaMAGuard model to check the safety of the samples model_name = "meta-llama/Llama-Guard-3-1B" model =… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/AdvBench-IR.textquestion-answeringn<1K4 likes61 downloads2y agoHugging Face08AdaptKey /ustax-irc-qa-89k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.textquestion-answering10K<n<100K0 likes44 downloads6mo agoHugging Face09sosa123454321 /iran-turkiye-startup-landing-kb Iran → Türkiye Startup Landing — legal/migration knowledge base One dataset for the whole platform (rule: one dataset, one space, one vector index — never several). Nightly snapshots produced by rag/scrape.py: news, academic, official (göç idaresi / ministry), legislation and directory sources about Iranian founders landing startups in Türkiye. All PII is scrubbed at fetch time (rag/fetch.scrub_pii). Chunks are stored as JSONL per snapshot day under data/kb/<date>/. This… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/iran-turkiye-startup-landing-kb.textquestion-answeringn<1K0 likes43 downloads12d agoHugging Face10irioder /littleHermione-benchmark Dataset card: O.W.L. & N.E.W.T. Bench development set v0.4 Summary Version 0.4 is a reviewed development benchmark of 75 short-answer factual questions about the seven English-language Harry Potter novels. It contains two separately scored examinations: 30 challenging O.W.L. questions covering recurring book canon beyond famous entrance-level facts; 45 frontier N.E.W.T. questions covering chapter-level prose details, minor names, precise objects, prices, and… See the full description on the dataset page: https://huggingface.co/datasets/irioder/littleHermione-benchmark.textquestion-answeringn<1K0 likes42 downloads17d agoHugging Face11AdaptKey /ustax-irc-qa-36k US Federal Tax Law QA Dataset (IRC — 36K pairs) Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC), used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1. Generation Pipeline IRC full text stored in a Qdrant vector store (chunked at ~512 tokens) An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk Generated pairs are deduplicated and split into train/validation Statistics Split Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.textquestion-answering10K<n<100K0 likes37 downloads6mo agoHugging Face12Amir7440 /IRAN-MADANI-LAWtextquestion-answeringn<1K1 likes34 downloads1y agoHugging Face13Irza /Arxiv_ph_indonesiatextquestion-answering1K<n<10K3 likes32 downloads3y agoHugging Face14Irfanuruchi /dsp-fft-sampling-aliasing Synthetic DSP Dataset: FFT + Sampling / Aliasing This repository contains synthetic instruction-style DSP samples designed for numerical reasoning and conceptual understanding of Digital Signal Processing (DSP) fundamentals. The dataset focuses on: FFT bin reasoning and frequency-domain interpretation Sampling theory Aliasing effects Dataset Origin & Verification This dataset was generated as part of the project: Fine-Tuning Lightweight Large Language Models for a… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/dsp-fft-sampling-aliasing.texttext-generation1K<n<10K0 likes30 downloads8mo agoHugging Face15Pangeanic /Iraqi-Arabic-multidomain-QA-text Iraqi Arabic Multidomain QA Dataset The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models. This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.textquestion-answeringn<1K1 likes27 downloads4mo agoHugging Face16Irina-Na /AutenticHadithestextquestion-answering1K<n<10K2 likes26 downloads1y agoHugging Face17SalahALHaismawi /uae-laws-irac UAE Laws Q&A Dataset (IRAC Format) A high-quality dataset of 9,477 question-answer pairs about UAE laws, formatted in IRAC (Issue, Rule, Application, Conclusion) legal reasoning structure. Dataset Creation Source Documents The dataset was built from a comprehensive collection of UAE legal documents, including: Federal Decrees and Laws Cabinet Resolutions Ministerial Decisions Civil and Commercial Codes Labor Law Traffic Law And more Creation Process… See the full description on the dataset page: https://huggingface.co/datasets/SalahALHaismawi/uae-laws-irac.textquestion-answering1K<n<10K1 likes24 downloads8mo agoHugging Face18ReliableAI /irish_belebelegatedIrish version of https://huggingface.co/datasets/facebook/belebele. Translated using facebook/nllb-200-3.3B, and the translations are verified by native Irish speakers. tabularquestion-answeringn<1K0 likes6 downloads2y agoHugging Face19IRUCAAI /doubao_Quanzhou_V3gatedtextquestion-answering100K<n<1M0 likes6 downloads2y agoHugging Face20IRUCAAI /doubao_Quanzhou_V2gatedtextquestion-answering100K<n<1M0 likes5 downloads2y agoHugging Face21IRUCAAI /doubao_Quanzhou_V1gatedtextquestion-answering100K<n<1M0 likes4 downloads2y agoHugging Face22IRUCAAI /doubao_HK_V2gatedtextquestion-answering1M<n<10M1 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.