datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.hierarchical-geospatial-reasoningmizan-iraqi-arabic-benchmark
Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set
Mizan is the first comprehensive, originally-authored evaluation benchmark
for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2
public development set: 340 originally-authored, dually-reviewed items across
two tracks (MSA baseline / Iraqi) and six axes.
📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865
🏆 Live leaderboard: https://mizan-bench.onrender.com
💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.docmatix-ir
Docmatix-IR
Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering.
To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.nasa-sde-IR-benchmark-20251024-v5
NASA SDE IR Benchmark v5
A comprehensive Information Retrieval benchmark dataset for the NASA Science Discovery Engine (SDE), containing synthetically generated query-document pairs for scientific content retrieval evaluation.
Paper: INDUS-SDE: A Language Model for Scientific Content Curation and Discovery — KDD 2026, AI for Sciences Track. This is the in-domain NASA SDE IR benchmark used to evaluate INDUS-SDE-ST.
Code: NASA-IMPACT/st-training-workflow
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-sde-IR-benchmark-20251024-v5.AdvBench-IR
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
This dataset includes malicious documents in response to AdvBench (Zou et al., 2023) queries. We have generated these documents using the Mistral-7B-Instruct-v0.2 language model.
from datasets import load_dataset
import transformers
ds = load_dataset("McGill-NLP/AdvBench-IR", split="train")
# Loads LlaMAGuard model to check the safety of the samples
model_name = "meta-llama/Llama-Guard-3-1B"
model =… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/AdvBench-IR.ustax-irc-qa-89k
US Federal Tax Law QA Dataset (IRC — 36K pairs)
Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v2.
Generation Pipeline
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Statistics
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-89k.iran-turkiye-startup-landing-kb
Iran → Türkiye Startup Landing — legal/migration knowledge base
One dataset for the whole platform (rule: one dataset, one space, one vector
index — never several). Nightly snapshots produced by rag/scrape.py:
news, academic, official (göç idaresi / ministry), legislation and directory
sources about Iranian founders landing startups in Türkiye.
All PII is scrubbed at fetch time (rag/fetch.scrub_pii).
Chunks are stored as JSONL per snapshot day under data/kb/<date>/.
This… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/iran-turkiye-startup-landing-kb.littleHermione-benchmark
Dataset card: O.W.L. & N.E.W.T. Bench development set v0.4
Summary
Version 0.4 is a reviewed development benchmark of 75 short-answer factual
questions about the seven English-language Harry Potter novels. It contains
two separately scored examinations:
30 challenging O.W.L. questions covering recurring book canon beyond famous
entrance-level facts;
45 frontier N.E.W.T. questions covering chapter-level prose details, minor
names, precise objects, prices, and… See the full description on the dataset page: https://huggingface.co/datasets/irioder/littleHermione-benchmark.ustax-irc-qa-36k
US Federal Tax Law QA Dataset (IRC — 36K pairs)
Synthetic question-answer pairs generated from the US Internal Revenue Code (IRC),
used to fine-tune AdaptKey/nemotron-30b-ustax-lora-v1.
Generation Pipeline
IRC full text stored in a Qdrant vector store (chunked at ~512 tokens)
An LLM-based Argo workflow (qdrant-qa-generator) generates QA pairs from each chunk
Generated pairs are deduplicated and split into train/validation
Statistics
Split
Records… See the full description on the dataset page: https://huggingface.co/datasets/AdaptKey/ustax-irc-qa-36k.IRAN-MADANI-LAWArxiv_ph_indonesiadsp-fft-sampling-aliasing
Synthetic DSP Dataset: FFT + Sampling / Aliasing
This repository contains synthetic instruction-style DSP samples
designed for numerical reasoning and conceptual understanding of
Digital Signal Processing (DSP) fundamentals.
The dataset focuses on:
FFT bin reasoning and frequency-domain interpretation
Sampling theory
Aliasing effects
Dataset Origin & Verification
This dataset was generated as part of the project:
Fine-Tuning Lightweight Large Language Models for a… See the full description on the dataset page: https://huggingface.co/datasets/Irfanuruchi/dsp-fft-sampling-aliasing.Iraqi-Arabic-multidomain-QA-text
Iraqi Arabic Multidomain QA Dataset
The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models.
This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.AutenticHadithesuae-laws-irac
UAE Laws Q&A Dataset (IRAC Format)
A high-quality dataset of 9,477 question-answer pairs about UAE laws, formatted in IRAC (Issue, Rule, Application, Conclusion) legal reasoning structure.
Dataset Creation
Source Documents
The dataset was built from a comprehensive collection of UAE legal documents, including:
Federal Decrees and Laws
Cabinet Resolutions
Ministerial Decisions
Civil and Commercial Codes
Labor Law
Traffic Law
And more
Creation Process… See the full description on the dataset page: https://huggingface.co/datasets/SalahALHaismawi/uae-laws-irac.irish_belebeleIrish version of https://huggingface.co/datasets/facebook/belebele.
Translated using facebook/nllb-200-3.3B, and the translations are verified by native Irish speakers.
doubao_Quanzhou_V3doubao_Quanzhou_V2doubao_Quanzhou_V1doubao_HK_V2
