CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01if-ir /nfcorpustexttext-retrieval1K<n<10K0 likes7.1k downloads1y agoHugging Face02ReliableAI /irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B. tabular100K<n<1M1 likes6.9k downloads2y agoHugging Face03MadeAgents /xlam-irrelevance-7.5k xlam-irrelevance-7.5k Overview The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs). Source and Construction This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k.text1K<n<10K21 likes1.8k downloads2y agoHugging Face04common-pile /ubuntu_irc Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 329,115 6.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc.texttext-generation100K<n<1M0 likes740 downloads1y agoHugging Face05ameer4wisam /iraqi-arabic-sales-dialogue-dataset Iraqi Arabic Sales Dialogue Dataset A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on retail sales, haggling, and everyday conversation. النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below. What this is 210,832 template-generated conversations, of which 171,601 (81%) are exact-unique message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.texttext-generation100K<n<1M0 likes703 downloads2mo agoHugging Face06PerSets /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is dadrah.ir… See the full description on the dataset page: https://huggingface.co/datasets/PerSets/iran-legal-persian-qa.textquestion-answering100K<n<1M7 likes657 downloads1y agoHugging Face07sander-wood /irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation. Dataset Summary The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.texttext-generation100K<n<1M28 likes655 downloads3y agoHugging Face08if-ir /fiqatexttext-retrieval10K<n<100K0 likes523 downloads1y agoHugging Face09common-pile /ubuntu_irc_filtered Ubuntu IRC Description Logs of all discussions on the Ubuntu-hosted Internet Relay Chat (IRC) since 2004 have been archived and released into the Public Domain. We downloaded all chats from all channels up until March of 2025. We consider all messages for a given channel on a given day as a single document. We removed system messages as well as those from known bots. Dataset Statistics Documents UTF-8 GB 234,982 5.3 License Issues… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/ubuntu_irc_filtered.texttext-generation100K<n<1M2 likes516 downloads1y agoHugging Face10vec-ai /struct-ir SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data github We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions. This repository contains the data for SSRB. Data Download Data can be… See the full description on the dataset page: https://huggingface.co/datasets/vec-ai/struct-ir.tabulartext-retrieval10M<n<100M3 likes478 downloads1y agoHugging Face11isaacus /irish-legislative-summaries Irish Legislative Summaries ⚖️ Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation. This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them. This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.texttext-retrieval1K<n<10K2 likes440 downloads11mo agoHugging Face12Farmaanaa /iran_inflation_and_cpi_1936_2022 ⚠️ نسخهٔ جایگزین این مجموعه‌داده با روش‌شناسیِ فعلیِ فرمانا به‌روز نمی‌شود. → Farmaanaa/iran_cpi_and_inflation_multisource فایل‌های قبلی برای آرشیو در دسترس می‌مانند. — farmaanaa.ir textn<1K0 likes345 downloads2mo agoHugging Face13ilkhamfy /IRexp IRexp — experimental IR band lists (commercial redistributable pool) IRexp is an openly redistributable collection of experimental infrared band lists mined from open-access chemistry literature, often with co-reported ¹H/¹³C shift lists and resolved structures (SMILES / InChIKey). Important: IRexp contains band lists (peak positions in cm⁻¹), not digitised absorbance traces. This is the form reported in publication text and is not directly comparable to SDBS or NIST full… See the full description on the dataset page: https://huggingface.co/datasets/ilkhamfy/IRexp.text100K<n<1M1 likes236 downloads9d agoHugging Face14rmoham05 /iran-legal-persian-qa Iranian Legal Question Answering Dataset (Farsi) This dataset includes over 600K questions and 2M answers, all in written form. The questions were posed by ordinary Persian speakers (Iranians), while the responses were provided by attorneys from various specialties. Dataset Description Question records without corresponding answers have been excluded from the dataset. This dataset will be updated periodically with new records. The reference for this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/rmoham05/iran-legal-persian-qa.textquestion-answering100K<n<1M0 likes219 downloads2mo agoHugging Face15SaifSilverHand /iraqi-dictionarytextn<1K5 likes211 downloads10h agoHugging Face16SetFit /amazon_massive_intent_fa-IRtext10K<n<100K0 likes197 downloads4y agoHugging Face17charlin55 /Genshin_Irminsul_Nahida 纳西妲角色扮演数据集(Nahida Roleplay) 基于Genshin Irminsul Dataset 拓展而来。使用Qwen、Deepseek等模型基于已有知识生成的对话数据。 特点 回复简短,避免长段文本输出,更符合聊天场景。 拒绝游戏外话题讨论(例如编写代码,翻译等),小模型经过微调后通用能力会变弱,不如直接拒绝。 以纳西妲口吻,将用户视为”旅行者”。 知识截止于「月之四」版本。 texttext-generation10K<n<100K1 likes194 downloads3mo agoHugging Face18IRedDragonICY /LaguQA LaguQA A benchmark for what a language model knows about 107 Indonesian regional and national songs. Every one was transcribed by hand from a printed songbook into ABC 2.1 notation, and every value traces back to a specific page. The questions cover bibliographic facts such as composer and region of origin, and reasoning over the notation. Some show a fragment of number notation and ask which song it is. Others ask the model to count bars or name the highest note. Number… See the full description on the dataset page: https://huggingface.co/datasets/IRedDragonICY/LaguQA.text10K<n<100K1 likes163 downloads19d agoHugging Face19DataScience-UIBK /OBLIQ-IR-Data OBLIQ-IR-Data The training data behind DataScience-UIBK/OBLIQ-IR-3B, plus the retrieval runs and evaluation outputs for every result in OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026). Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance, an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has little or no surface expression in the document. 🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.texttext-retrieval100K<n<1M4 likes131 downloads1mo agoHugging Face20irfankabir02 /OpenGameArt-GPL-3.0 Dataset Card for OpenGameArt-GPL-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.audioimage-classificationn<1K0 likes130 downloads10mo agoHugging Face21IrvinTopi /WebFAQHardNegatives WebFAQ 2.0: Multilingual Hard Negatives This dataset contains mined hard negatives derived from the WebFAQ 2.0 corpus. It includes approximately 1.3 million samples across 20 languages. The dataset is designed to support robust training of dense retrieval models, specifically enabling: Contrastive Learning: Using strict hard negatives to improve discrimination. Knowledge Distillation: Using the provided cross-encoder scores to train with soft labels (e.g., MarginMSE).… See the full description on the dataset page: https://huggingface.co/datasets/IrvinTopi/WebFAQHardNegatives.textsentence-similarity100K<n<1M2 likes128 downloads7mo agoHugging Face22rcds /swiss_doc2doc_irhttps://huggingface.co/spaces/huggingface/datasets-tagging Dataset Card for Swiss Doc2doc Information Retrieval Dataset Summary Swiss Doc2doc Information Retrieval is a multilingual, diachronic dataset of 131K Swiss Federal Supreme Court (FSCS) cases annotated with law citations and ruling citations, posing a challenging text classification task. As unique label we are using decision_id of cited rulings and uuid of cited law articles, which can be found in the… See the full description on the dataset page: https://huggingface.co/datasets/rcds/swiss_doc2doc_ir.tabulartext-classification100K<n<1M1 likes125 downloads3y agoHugging Face23irisxx /ultrafeedback_tied Train dir contains train set with different ratios of tie data Test dir contains test sets which used to evaluate performances on the in-distribution data. test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data. non_tie_data_test.jsonl contains 1500 non-tie samples. tie_data_test.jsonl contains 500 tie samples. Citation Please cite our paper if you find the dataset helpful in your work: @inproceedings{ guo2025todo, title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.tabular10K<n<100K0 likes119 downloads1y agoHugging Face24if-ir /pmtexttext-retrieval100K<n<1M0 likes115 downloads1y agoHugging Face25IrohXu /SocialGesture [CVPR 2025] SocialGesture: Delving into Multi-person Gesture Understanding Dataset Description We introduce SocialGesture, the first large-scale dataset specifically designed for multi-person gesture analysis. SocialGesture features a diverse range of natural scenarios and supports multiple gesture analysis tasks, including video-based recognition and temporal localization, providing a valuable resource for advancing the study of gesture during complex social… See the full description on the dataset page: https://huggingface.co/datasets/IrohXu/SocialGesture.text10K<n<100K5 likes114 downloads1y agoHugging Face26i-Lang /iReview iReview AI-to-AI code review, powered by I-Lang protocol. Any OpenAI-compatible model reviews your code. Structured instructions in, structured findings out. I-Lang is the first protocol to formally map Greek mathematical symbols (Σ, Δ, φ, λ, Ω, ∇, μ, Π, ψ, ξ, ζ, θ, ∂) as primitive verbs for AI-to-AI communication, and the first to define a computable vector space for AI judgment (11 dimensions, 4 axioms). Why Every AI-to-AI code review tool today sends… See the full description on the dataset page: https://huggingface.co/datasets/i-Lang/iReview.textn<1K0 likes107 downloads2d agoHugging Face27ameer4wisam /iraqi_words_finetuning Iraqi Words A manually compiled Iraqi Arabic dialect lexicon (930 terms, 50 categories) with a dependency-free BM25 retriever and a fine-tuning data generator built on top of it. Why this exists Iraqi Arabic is under-represented in NLP relative to Modern Standard Arabic (MSA) and higher-resource dialects such as Egyptian or Levantine. Lexical resources that map Iraqi terms to their MSA meanings — the kind needed to ground retrieval or instruction-tuning for… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi_words_finetuning.texttranslationn<1K0 likes107 downloads2mo agoHugging Face28Ireliya /hierarchical-geospatial-reasoningimagequestion-answeringn<1K2 likes106 downloads6mo agoHugging Face29Tevatron /docmatix-ir Docmatix-IR Docmatix is originally a large dataset designed for fine-tuning large vision-language models on Visual Question Answering tasks. It contains a substantial collection of PDF images (2.4M) and a vast set of questions (9.5M) related to these images. However, many of the questions in the Docmatix dataset are not suitable for open-domain question answering. To address this, we have converted Docmatix into Docmatix-IR, a training set suitable for training document visual… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/docmatix-ir.textquestion-answering1M<n<10M15 likes93 downloads2y agoHugging Face30cfli /Reinforced-IR-synthetic Introduction Synthetic data for Reinforced IR. Load Dataset An example to load the dataset: import datasets # load dataset dataset = datasets.load_dataset( "cfli/Reinforced-IR-synthetic", 'dbpedia-entity', split='generator' ) # print one sample print(dataset[0]) text100K<n<1M1 likes91 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.