CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01THU-IAR /MIntRec Dataset details In real-world conversational interactions, we usually combine information from multiple modalities (e.g., text, video, audio) to help analyze human intentions. Though intent analysis has been widely explored in the Natural Language Processing community, there is a scarcity of data for multimodal intent analysis. Thus, we provide a novel multimodal intent benchmark dataset, MIntRec, to boom the research. To the best of our knowledge, it is the first multimodal intent… See the full description on the dataset page: https://huggingface.co/datasets/THU-IAR/MIntRec.text1K<n<10K1 likes3.6k downloads2y agoHugging Face02Yiyang-Ian-Li /LongDA LongDA Dataset Card Dataset Description LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis. Dataset Summary 505 queries extracted from 30 expert-written publications 17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.documentquestion-answeringn<1K1 likes1.4k downloads3mo agoHugging Face03autoiac-project /iac-eval IaC-Eval dataset (v1.1) IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities. This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now). | Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper | 2. Usage instructions Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.texttext-generationn<1K7 likes323 downloads2y agoHugging Face04NinjasOfSEO /ia-overviews-monitortext1K<n<10K0 likes303 downloads9h agoHugging Face05iarfmoose /qa_evaluatorThis is the same dataset as the question_generator dataset but with the context removed and the question and answer in separate fields. This is intended to be used with the question_generator repo to train the qa_evaluator model which predicts whether a question and answer pair makes sense. text100K<n<1M4 likes219 downloads5y agoHugging Face06iawen /java_unit_testtextn<1K2 likes217 downloads3y agoHugging Face07AmazonScience /Multi-IaC-Eval Multi-IaC-Eval We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation. Cloudformation: 263 Terraform: 446 CDK (Python): 64 CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.texttext-generationn<1K0 likes102 downloads1y agoHugging Face08Lucasdegeorge /ImageNet_TA_IA Library: https://github.com/lucasdegeorge/T2I-ImageNet How far can we go with ImageNet for Text-to-Image generation? Lucas Degeorge, Arijit Ghosh, Nicolas Dufour, David Picard, Vicky Kalogeiton This dataset has the captions used during the training of the models from the paper "How far can we go with ImageNet for Text-to-Image generation?" The core idea is that text-to-image generation models typically rely on vast datasets, prioritizing quantity over quality. The usual… See the full description on the dataset page: https://huggingface.co/datasets/Lucasdegeorge/ImageNet_TA_IA.text1M<n<10M5 likes89 downloads10mo agoHugging Face09apisdom /vulnerabilidades-ia-espanol Vulnerabilidades CVE en sistemas de IA (espanol) Corpus de advisories CVE/GHSA que afectan a paquetes y SDKs de IA (langchain, openai, anthropic, llamaindex, etc.) traducido al espanol por LaAutopsIA (ApisDom Intelligence Group). Fuente principal: GitHub Advisory Database (CC-BY-4.0). Cifras del snapshot actual 14 vulnerabilidades publicadas en este snapshot. Mes archivado: 2026-08. Ultima edicion: 2026-09-01T21:35:37.468Z. Frecuencia: sincronizacion mensual.… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/vulnerabilidades-ia-espanol.tabulartext-classificationn<1K0 likes80 downloads21d agoHugging Face10apisdom /indice-fallos-ia-espanol Indice de Fallos IA en espanol Snapshots mensuales del Indice de Fallos IA producido por el observatorio La AutopsIA (ApisDom Intelligence Group). Mide la fiabilidad de modelos LLM con benchmarks oficiales independientes, en formato citable y trazable. Cifras del snapshot actual 765 mediciones en este snapshot. Mes archivado: 2026-09. Recomputado: 2026-09-01T21:34:26.058Z. Frecuencia: sincronizacion mensual. Para que sirve este dataset Datos… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/indice-fallos-ia-espanol.tabulartabular-classification1K<n<10K0 likes76 downloads21d agoHugging Face11iapp /rag_thai_laws Thai Laws Dataset This dataset contains Thai law texts from the Office of the Council of State, Thailand. The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included: Converting system IDs to integer format Removing leading/trailing whitespace from titles and text Normalizing newlines to maintain consistent formatting Removing excessive blank lines The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/iapp/rag_thai_laws.texttext-generation10K<n<100K3 likes68 downloads2y agoHugging Face12hassan-IA /dah Dataset Card for DAH (DAtaset Hassaniya) DAH is a bilingual dataset created to support translation between Hassaniya dialect and English, with both Arabic script and Arabizi (Latin-transliterated) Hassaniya included. Dataset Details Name: dah Languages: Hassaniya Arabic (in both Arabic and Latin forms), English Creators: Ahlam Abdelkader Emani Babe Oumoukelthoum Sidenna Adapted from: Tatoeba Project parallel sentences Reviewed by: Founders of the… See the full description on the dataset page: https://huggingface.co/datasets/hassan-IA/dah.text1K<n<10K6 likes68 downloads6mo agoHugging Face13iamplus /CoTCoT Datasets from Google's FLan Dataset text10K<n<100K2 likes65 downloads3y agoHugging Face14IAmKarthik /protein-compound-affinity-esm2-molformertext100K<n<1M0 likes59 downloads3mo agoHugging Face15iarfmoose /question_generatorThis dataset is made up of data taken from SQuAD v2.0, RACE, CoQA, and MSMARCO. Some examples have been filtered out of the original datasets and others have been modified. There are two fields; question and text. The question field contains the question, and the text field contains both the answer and the context in the following format: "<answer> (answer text) <context> (context text)" The and are included as special tokens in the question generator's tokenizer. This dataset is intended to… See the full description on the dataset page: https://huggingface.co/datasets/iarfmoose/question_generator.text100K<n<1M12 likes56 downloads5y agoHugging Face16iac-eval-v2 /iac-eval-v2 IaC-Eval v2 Modernised Terraform code-generation benchmark — 186 tasks, Terraform 1.15 + OPA 1.16 (Rego v1). An updated and extended version of the IaC-Eval NeurIPS 2024 benchmark. Scoring is deterministic: the generated HCL either passes terraform plan + opa eval, or it doesn't — no LLM-as-judge. Dataset summary Field Value Tasks 186 (AWS only) Difficulty 1–6 (distribution: 1→35, 2→40, 3→51, 4→22, 5→9, 6→13) AWS services 34 distinct Terraform… See the full description on the dataset page: https://huggingface.co/datasets/iac-eval-v2/iac-eval-v2.texttext-generationn<1K0 likes52 downloads4mo agoHugging Face17MihaiIonascu /Azure_IaC_testtextn<1K0 likes51 downloads3y agoHugging Face18iamramzan /Global-Population-Data List of Countries and Dependencies by Population This dataset contains population-related information for countries and dependencies, scraped from Wikipedia. The dataset includes the following columns: Location: The country or dependency name. Population: Total population count. % of World: The percentage of the world's population this country or dependency represents. Date: The date of the population estimate. Source: Whether the source is official or derived from the United… See the full description on the dataset page: https://huggingface.co/datasets/iamramzan/Global-Population-Data.texttext-classificationn<1K1 likes49 downloads2y agoHugging Face19shishir-dwi /News-Article-Categorization_IAB Article and Category Dataset Overview This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more. Dataset Information Number of Samples: 871,909 Number of Categories: 26 Column Information text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.texttext-classification100K<n<1M4 likes45 downloads3y agoHugging Face20conversoriaecnae /iae-cnae-2025 IAE & CNAE 2025 Open Dataset (Spain) Three machine-readable Spanish tax-classification catalogs. Maintained by conversoriaecnae.es. License: CC-BY-4.0 — please attribute conversoriaecnae.es when reusing. Files File Rows Source data/v1/iae.{csv,jsonl} 1,189 AEAT — Tarifas IAE data/v1/cnae_2025.{csv,jsonl} 1,060 INE / Real Decreto 10/2025 Art. 6 data/v1/correspondencia_cnae2009_cnae2025.{csv,jsonl} 1,010 INE / RD 10/2025 Art. 7(c) Schema… See the full description on the dataset page: https://huggingface.co/datasets/conversoriaecnae/iae-cnae-2025.texttext-classification1K<n<10K0 likes43 downloads4mo agoHugging Face21iamjayeshc /reddit-self-medication-claim-dataset Reddit Self-Medication Claim Dataset Dataset Summary The Reddit Self-Medication Claim Dataset is an annotated NLP dataset designed to study self-medication claims expressed in informal online health discussions. The dataset focuses on identifying whether Reddit posts contain self-medication related claims, and further distinguishing between explicit and implicit expressions of self-medication behavior. This dataset was created as part of an independent research… See the full description on the dataset page: https://huggingface.co/datasets/iamjayeshc/reddit-self-medication-claim-dataset.tabulartext-classification1K<n<10K1 likes43 downloads2mo agoHugging Face22iai-group /clef2024_checkthat_task1_en Bibtex @inproceedings{Hasanain:CLEF:24, author = {Maram Hasanain and Reem Suwaileh and Sanne Weering and Chengkai Li and Tommaso Caselli and Wajdi Zaghouani and Alberto Barr{\'{o}}n{-}Cede{\~{n}}o and Preslav Nakov and Firoj Alam}, editor = {Guglielmo Faggioli and Nicola Ferro and Petra Galusc{\'{a}}kov{\'{a}} and… See the full description on the dataset page: https://huggingface.co/datasets/iai-group/clef2024_checkthat_task1_en.texttext-classification10K<n<100K0 likes42 downloads2y agoHugging Face23OSS-forge /Extended_Shellcode_IA32 Shellcode_IA32 Shellcode_IA32 is a dataset containing more than 20 years of shellcodes from a variety of sources and is the largest collection of shellcodes in assembly available to date. We are currently extending the dataset. Up to now, we released three versions of the dataset. Shellcode_IA32 was presented for the first time in the paper Shellcode_IA32: A Dataset for Automatic Shellcode Generation, accepted to the 1st Workshop on Natural Language Processing for Programming… See the full description on the dataset page: https://huggingface.co/datasets/OSS-forge/Extended_Shellcode_IA32.text1K<n<10K10 likes42 downloads10mo agoHugging Face24iakis /1_exploitstabularn<1K0 likes41 downloads2y agoHugging Face25iai-group /clef2024_checkthat_task1_es Bibtex @inproceedings{Hasanain:CLEF:24, author = {Maram Hasanain and Reem Suwaileh and Sanne Weering and Chengkai Li and Tommaso Caselli and Wajdi Zaghouani and Alberto Barr{\'{o}}n{-}Cede{\~{n}}o and Preslav Nakov and Firoj Alam}, editor = {Guglielmo Faggioli and Nicola Ferro and Petra Galusc{\'{a}}kov{\'{a}} and… See the full description on the dataset page: https://huggingface.co/datasets/iai-group/clef2024_checkthat_task1_es.texttext-classification10K<n<100K0 likes39 downloads2y agoHugging Face26DebasishDhal99 /IAST-corpus Dataset Details Dataset created by transliterating existing datasets to IAST by means of IAST transliteration library Languages include Sanskrit, Hindi, Odia, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Marathi, Bhojpuri, Nepali. Pre-existing dataset source(s) - Wikipedia. Across all subsets, 'source', 'target', 'source_lang', 'target_lang', 'source' columns are common. This is just a hobby dataset, but it should abide by the licenses of the input dataset(s). texttranslation100K<n<1M0 likes38 downloads1y agoHugging Face27IAMJB /RadEvalExpertDatasettextn<1K0 likes38 downloads1y agoHugging Face28iamramzan /sovereign-states-dataset Sovereign States Dataset This dataset provides a comprehensive list of sovereign states, along with their common and formal names, membership within the UN system, and details on sovereignty disputes and recognition status. The data was originally scraped from Wikipedia's List of Sovereign States and processed for clarity and usability. Dataset Features Common Name: The commonly used name of the country or state. Formal Name: The official/formal name of the country… See the full description on the dataset page: https://huggingface.co/datasets/iamramzan/sovereign-states-dataset.texttext-classificationn<1K1 likes36 downloads2y agoHugging Face29MihaiIonascu /Azure_IaC_reducedtextn<1K0 likes32 downloads3y agoHugging Face30iamshnoo /alpaca-cleaned-albaniantext10K<n<100K2 likes32 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.