datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nfcorpusirish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
xlam-irrelevance-7.5k
xlam-irrelevance-7.5k
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k.ubuntu_irc
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain.
We downloaded all chats from all channels up until March of 2025.
We consider all messages for given channel on a given day as a single document.
We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
329,115
6.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.iraqi-arabic-sales-dialogue-dataset
Iraqi Arabic Sales Dialogue Dataset
A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on
retail sales, haggling, and everyday conversation.
النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below.
What this is
210,832 template-generated conversations, of which 171,601 (81%) are exact-unique
message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The
core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.fiqaubuntu_irc_filtered
Ubuntu IRC
Description
Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots.
Dataset Statistics
Documents
UTF-8 GB
234,982
5.3
License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.struct-ir
SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data
github
We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions.
This repository contains the data for SSRB.
Data Download
Data can be… See the full description on the dataset page: https://huggingface.co/datasets/vec-ai/struct-ir.irish-legislative-summaries
Irish Legislative Summaries ⚖️
Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation.
This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.iran_inflation_and_cpi_1936_2022
⚠️ نسخهٔ جایگزین
این مجموعهداده با روششناسیِ فعلیِ فرمانا بهروز نمیشود.
→ Farmaanaa/iran_cpi_and_inflation_multisource
فایلهای قبلی برای آرشیو در دسترس میمانند.
— farmaanaa.ir
IRexp
IRexp — experimental IR band lists (commercial redistributable pool)
IRexp is an openly redistributable collection of experimental infrared band lists mined from open-access chemistry literature, often with co-reported ¹H/¹³C shift lists and resolved structures (SMILES / InChIKey).
Important: IRexp contains band lists (peak positions in cm⁻¹), not digitised absorbance traces. This is the form reported in publication text and is not directly comparable to SDBS or NIST full… See the full description on the dataset page: https://huggingface.co/datasets/ilkhamfy/IRexp.iran-legal-persian-qa
Iranian Legal Question Answering Dataset (Farsi)
This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties.
Dataset Description
Question records without corresponding answers have been excluded from the dataset.
This dataset will be updated periodically with new records.
The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.iraqi-dictionaryamazon_massive_intent_fa-IRGenshin_Irminsul_Nahida
纳西妲角色扮演数据集(Nahida Roleplay)
基于Genshin Irminsul Dataset 拓展而来。使用Qwen、Deepseek等模型基于已有知识生成的对话数据。
特点
回复简短,避免长段文本输出,更符合聊天场景。
拒绝游戏外话题讨论(例如编写代码,翻译等),小模型经过微调后通用能力会变弱,不如直接拒绝。
以纳西妲口吻,将用户视为”旅行者”。
知识截止于「月之四」版本。
LaguQA
LaguQA
A benchmark for what a language model knows about 107 Indonesian regional
and national songs. Every one was transcribed by hand from a printed songbook
into ABC 2.1 notation, and every value traces back to a specific page.
The questions cover bibliographic facts such as composer and region of origin,
and reasoning over the notation. Some show a fragment of number notation and ask
which song it is. Others ask the model to count bars or name the highest note.
Number… See the full description on the dataset page: https://huggingface.co/datasets/IRedDragonICY/LaguQA.OBLIQ-IR-Data
OBLIQ-IR-Data
The training data behind DataScience-UIBK/OBLIQ-IR-3B,
plus the retrieval runs and evaluation outputs for every result in
OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026).
Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance,
an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has
little or no surface expression in the document.
🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.WebFAQHardNegatives
WebFAQ 2.0: Multilingual Hard Negatives
This dataset contains mined hard negatives derived from the WebFAQ 2.0 corpus. It includes approximately 1.3 million samples across 20 languages.
The dataset is designed to support robust training of dense retrieval models, specifically enabling:
Contrastive Learning: Using strict hard negatives to improve discrimination.
Knowledge Distillation: Using the provided cross-encoder scores to train with soft labels (e.g., MarginMSE).… See the full description on the dataset page: https://huggingface.co/datasets/IrvinTopi/WebFAQHardNegatives.swiss_doc2doc_irhttps://huggingface.co/spaces/huggingface/datasets-tagging
Dataset Card for Swiss Doc2doc Information Retrieval
Dataset Summary
Swiss Doc2doc Information Retrieval is a multilingual, diachronic dataset of 131K Swiss Federal Supreme Court (FSCS) cases annotated with law citations and ruling citations, posing a challenging text classification task. As unique label we are using decision_id of cited rulings and uuid of cited law articles, which can be found in the… See the full description on the dataset page: https://huggingface.co/datasets/rcds/swiss_doc2doc_ir.ultrafeedback_tied
Train dir contains train set with different ratios of tie data
Test dir contains test sets which used to evaluate performances on the in-distribution data.
test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data.
non_tie_data_test.jsonl contains 1500 non-tie samples.
tie_data_test.jsonl contains 500 tie samples.
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.pmSocialGesture
[CVPR 2025] SocialGesture: Delving into Multi-person Gesture Understanding
Dataset Description
We introduce SocialGesture, the first large-scale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social… See the full description on the dataset page: https://huggingface.co/datasets/IrohXu/SocialGesture.iReview
iReview
AI-to-AI code review, powered by I-Lang protocol.
Any OpenAI-compatible model reviews your code. Structured instructions in, structured findings out.
I-Lang is the first protocol to formally map Greek mathematical symbols (Σ, Δ, φ, λ, Ω, ∇, μ, Π, ψ, ξ, ζ, θ, ∂) as primitive verbs for AI-to-AI communication, and the first to define a computable vector space for AI judgment (11 dimensions, 4 axioms).
Why
Every AI-to-AI code review tool today sends… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/iReview.iraqi_words_finetuning
Iraqi Words
A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a
dependency-free BM25 retriever and a fine-tuning data generator built on top of it.
Why this exists
Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA)
and higher-resource dialects such as Egyptian or Levantine. Lexical resources that
map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or
instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.hierarchical-geospatial-reasoningdocmatix-ir
Docmatix-IR
Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering.
To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.Reinforced-IR-synthetic
Introduction
Synthetic data for Reinforced IR.
Load Dataset
An example to load the dataset:
import datasets
# load dataset
dataset = datasets.load_dataset(
"cfli/Reinforced-IR-synthetic",
'dbpedia-entity',
split='generator'
)
# print one sample
print(dataset[0])
