CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01if-ir /nfcorpustexttext-retrieval1K<n<10K0 likes6.7k downloads1y agoHugging Face02ReliableAI /irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B. tabular100K<n<1M1 likes6.7k downloads2y agoHugging Face03MadeAgents /xlam-irrelevance-7.5k xlam-irrelevance-7.5k Overview The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs). Source and Construction This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k.text1K<n<10K21 likes1.9k downloads2y agoHugging Face04common-pile /ubuntu_irc Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 329,115 6.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.texttext-generation100K<n<1M0 likes755 downloads1y agoHugging Face05ameer4wisam /iraqi-arabic-sales-dialogue-dataset Iraqi Arabic Sales Dialogue Dataset A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on retail sales, haggling, and everyday conversation. النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below. What this is 210,832 template-generated conversations, of which 171,601 (81%) are exact-unique message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.texttext-generation100K<n<1M0 likes706 downloads2mo agoHugging Face06PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes652 downloads1y agoHugging Face07sander-wood /irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation. Dataset Summary The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.texttext-generation100K<n<1M28 likes629 downloads3y agoHugging Face08common-pile /ubuntu_irc_filtered Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 234,982 5.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.texttext-generation100K<n<1M2 likes560 downloads1y agoHugging Face09if-ir /fiqatexttext-retrieval10K<n<100K0 likes524 downloads1y agoHugging Face10vec-ai /struct-ir SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data github We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions. This repository contains the data for SSRB. Data Download Data can be… See the full description on the dataset page: https://huggingface.co/datasets/vec-ai/struct-ir.tabulartext-retrieval10M<n<100M3 likes463 downloads1y agoHugging Face11isaacus /irish-legislative-summaries Irish Legislative Summaries ⚖️ Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation. This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them. This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.texttext-retrieval1K<n<10K2 likes422 downloads11mo agoHugging Face12Farmaanaa /iran_inflation_and_cpi_1936_2022 ⚠️ نسخهٔ جایگزین این مجموعه‌داده با روش‌شناسیِ فعلیِ فرمانا به‌روز نمی‌شود. → Farmaanaa/iran_cpi_and_inflation_multisource فایل‌های قبلی برای آرشیو در دسترس می‌مانند. — farmaanaa.ir textn<1K0 likes261 downloads2mo agoHugging Face13rmoham05 /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.textquestion-answering100K<n<1M0 likes235 downloads2mo agoHugging Face14SaifSilverHand /iraqi-dictionarytextn<1K5 likes222 downloads2d agoHugging Face15SetFit /amazon_massive_intent_fa-IRtext10K<n<100K0 likes185 downloads4y agoHugging Face16ilkhamfy /IRexp IRexp — experimental IR band lists (commercial redistributable pool) IRexp is an openly redistributable collection of experimental infrared band lists mined from open-access chemistry literature, often with co-reported ¹H/¹³C shift lists and resolved structures (SMILES / InChIKey). Important: IRexp contains band lists (peak positions in cm⁻¹), not digitised absorbance traces. This is the form reported in publication text and is not directly comparable to SDBS or NIST full… See the full description on the dataset page: https://huggingface.co/datasets/ilkhamfy/IRexp.text100K<n<1M1 likes180 downloads11d agoHugging Face17IRedDragonICY /LaguQA LaguQA A benchmark for what a language model knows about 107 Indonesian regional and national songs. Every one was transcribed by hand from a printed songbook into ABC 2.1 notation, and every value traces back to a specific page. The questions cover bibliographic facts such as composer and region of origin, and reasoning over the notation. Some show a fragment of number notation and ask which song it is. Others ask the model to count bars or name the highest note. Number… See the full description on the dataset page: https://huggingface.co/datasets/IRedDragonICY/LaguQA.text10K<n<100K1 likes161 downloads21d agoHugging Face18rcds /swiss_doc2doc_irhttps://huggingface.co/spaces/huggingface/datasets-tagging Dataset Card for Swiss Doc2doc Information Retrieval Dataset Summary Swiss Doc2doc Information Retrieval is a multilingual, diachronic dataset of 131K Swiss Federal Supreme Court (FSCS) cases annotated with law citations and ruling citations, posing a challenging text classification task. As unique label we are using decision_id of cited rulings and uuid of cited law articles, which can be found in the… See the full description on the dataset page: https://huggingface.co/datasets/rcds/swiss_doc2doc_ir.tabulartext-classification100K<n<1M1 likes137 downloads3y agoHugging Face19irfankabir02 /OpenGameArt-GPL-3.0 Dataset Card for OpenGameArt-GPL-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.audioimage-classificationn<1K0 likes132 downloads10mo agoHugging Face20IrohXu /SocialGesture [CVPR 2025] SocialGesture: Delving into Multi-person Gesture Understanding Dataset Description We introduce SocialGesture, the first large-scale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social… See the full description on the dataset page: https://huggingface.co/datasets/IrohXu/SocialGesture.text10K<n<100K5 likes128 downloads1y agoHugging Face21irisxx /ultrafeedback_tied Train dir contains train set with different ratios of tie data Test dir contains test sets which used to evaluate performances on the in-distribution data. test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data. non_tie_data_test.jsonl contains 1500 non-tie samples. tie_data_test.jsonl contains 500 tie samples. Citation Please cite our paper if you find the dataset helpful in your work: @inproceedings{ guo2025todo, title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.tabular10K<n<100K0 likes122 downloads1y agoHugging Face22ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes111 downloads2mo agoHugging Face23Ireliya /hierarchical-geospatial-reasoningimagequestion-answeringn<1K2 likes107 downloads6mo agoHugging Face24i-Lang /iReview iReview AI-to-AI code review, powered by I-Lang protocol. Any OpenAI-compatible model reviews your code. Structured instructions in, structured findings out. I-Lang is the first protocol to formally map Greek mathematical symbols (Σ, Δ, φ, λ, Ω, ∇, μ, Π, ψ, ξ, ζ, θ, ∂) as primitive verbs for AI-to-AI communication, and the first to define a computable vector space for AI judgment (11 dimensions, 4 axioms). Why Every AI-to-AI code review tool today sends… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/iReview.textn<1K0 likes102 downloads4d agoHugging Face25DataScience-UIBK /OBLIQ-IR-Data OBLIQ-IR-Data The training data behind DataScience-UIBK/OBLIQ-IR-3B, plus the retrieval runs and evaluation outputs for every result in OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026). Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance, an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has little or no surface expression in the document. 🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.texttext-retrieval100K<n<1M4 likes100 downloads1mo agoHugging Face26IrvinTopi /WebFAQHardNegatives WebFAQ 2.0: Multilingual Hard Negatives This dataset contains mined hard negatives derived from the WebFAQ 2.0 corpus. It includes approximately 1.3 million samples across 20 languages. The dataset is designed to support robust training of dense retrieval models, specifically enabling: Contrastive Learning: Using strict hard negatives to improve discrimination. Knowledge Distillation: Using the provided cross-encoder scores to train with soft labels (e.g., MarginMSE).… See the full description on the dataset page: https://huggingface.co/datasets/IrvinTopi/WebFAQHardNegatives.textsentence-similarity100K<n<1M2 likes99 downloads7mo agoHugging Face27cfli /Reinforced-IR-synthetic Introduction Synthetic data for Reinforced IR. Load Dataset An example to load the dataset: import datasets # load dataset dataset = datasets.load_dataset( "cfli/Reinforced-IR-synthetic", 'dbpedia-entity', split='generator' ) # print one sample print(dataset[0]) text100K<n<1M1 likes94 downloads1y agoHugging Face28Benitoow /OfficeSmith-PPTX-IR OfficeSmith PPTX IR Synthetic bilingual business briefs paired with editable PPTX intermediate representations. Dataset summary This dataset is part of the OfficeSmith collection for training models to plan, build, clarify, critique, and repair editable business presentations. It contains observable outputs only: no hidden chain of thought, secret benchmark prompt, personal data, or API credential is included. Train rows: 862 Validation rows: 48 Test rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/Benitoow/OfficeSmith-PPTX-IR.texttext-generationn<1K0 likes91 downloads1mo agoHugging Face29Rev3auth /iris-mixtext1K<n<10K1 likes91 downloads5d agoHugging Face30Shadow0482 /iris_02text10K<n<100K0 likes90 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.