datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MIntRec
Dataset details
In real-world conversational interactions, we usually combine information from multiple modalities (e.g., text, video, audio) to help analyze human intentions. Though intent analysis has been widely explored in the Natural Language Processing community, there is a scarcity of data for multimodal intent analysis. Thus, we provide a novel multimodal intent benchmark dataset, MIntRec, to boom the research. To the best of our knowledge, it is the first multimodal intent… See the full description on the dataset page: https://huggingface.co/datasets/THU-IAR/MIntRec.LongDA
LongDA Dataset Card
Dataset Description
LongDA is a data analysis benchmark for evaluating LLM-based agents under documentation-intensive analytical workflows. It features authentic U.S. government survey data with complete, long documentation, testing LLMs' ability to navigate complex real-world datasets before performing analysis.
Dataset Summary
505 queries extracted from 30 expert-written publications
17 U.S. national surveys covering health… See the full description on the dataset page: https://huggingface.co/datasets/Yiyang-Ian-Li/LongDA.iac-eval
IaC-Eval dataset (v1.1)
IaC-Eval dataset is the first human-curated and challenging Cloud Infrastructure-as-Code (IaC) dataset tailored to more rigorously benchmark large language models' IaC code generation capabilities.
This dataset contains 458 questions ranging from simple to difficult across various cloud services (targeting AWS for now).
| Github | 🏆 Leaderboard TBD | 📖 NeurIPS 2024 Paper |
2. Usage instructions
Option 1: Running the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/autoiac-project/iac-eval.ia-overviews-monitorqa_evaluatorThis is the same dataset as the question_generator dataset but with the context removed and the question and answer in separate fields. This is intended to be used with the question_generator repo to train the qa_evaluator model which predicts whether a question and answer pair makes sense.
java_unit_testMulti-IaC-Eval
Multi-IaC-Eval
We present Multi-IaC-Eval is a novel benchmark dataset for evaluating LLM-based IaC generation and mutation across AWS CloudFormation, Terraform, and Cloud Development Kit (CDK) formats. The dataset consists of triplets containing initial IaC templates, natural language modification requests, and corresponding updated templates, created through a synthetic data generation pipeline with rigorous validation.
Cloudformation: 263
Terraform: 446
CDK (Python): 64
CDK… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/Multi-IaC-Eval.ImageNet_TA_IA
Library: https://github.com/lucasdegeorge/T2I-ImageNet
How far can we go with ImageNet for Text-to-Image generation?
Lucas Degeorge, Arijit Ghosh, Nicolas Dufour, David Picard, Vicky Kalogeiton
This dataset has the captions used during the training of the models from the paper "How far can we go with ImageNet for Text-to-Image generation?"
The core idea is that text-to-image generation models typically rely on vast datasets, prioritizing quantity over quality. The usual… See the full description on the dataset page: https://huggingface.co/datasets/Lucasdegeorge/ImageNet_TA_IA.vulnerabilidades-ia-espanol
Vulnerabilidades CVE en sistemas de IA (espanol)
Corpus de advisories CVE/GHSA que afectan a paquetes y SDKs de IA (langchain, openai, anthropic, llamaindex, etc.) traducido al espanol por LaAutopsIA (ApisDom Intelligence Group). Fuente principal: GitHub Advisory Database (CC-BY-4.0).
Cifras del snapshot actual
14 vulnerabilidades publicadas en este snapshot.
Mes archivado: 2026-08.
Ultima edicion: 2026-09-01T21:35:37.468Z.
Frecuencia: sincronizacion mensual.… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/vulnerabilidades-ia-espanol.indice-fallos-ia-espanol
Indice de Fallos IA en espanol
Snapshots mensuales del Indice de Fallos IA producido por el observatorio La AutopsIA (ApisDom Intelligence Group). Mide la fiabilidad de modelos LLM con benchmarks oficiales independientes, en formato citable y trazable.
Cifras del snapshot actual
765 mediciones en este snapshot.
Mes archivado: 2026-09.
Recomputado: 2026-09-01T21:34:26.058Z.
Frecuencia: sincronizacion mensual.
Para que sirve este dataset
Datos… See the full description on the dataset page: https://huggingface.co/datasets/apisdom/indice-fallos-ia-espanol.rag_thai_laws
Thai Laws Dataset
This dataset contains Thai law texts from the Office of the Council of State, Thailand.
The dataset has been cleaned and processed by the iApp Team to improve data quality and accessibility. The cleaning process included:
Converting system IDs to integer format
Removing leading/trailing whitespace from titles and text
Normalizing newlines to maintain consistent formatting
Removing excessive blank lines
The cleaned dataset is now available on Hugging Face for easy… See the full description on the dataset page: https://huggingface.co/datasets/iapp/rag_thai_laws.dah
Dataset Card for DAH (DAtaset Hassaniya)
DAH is a bilingual dataset created to support translation between Hassaniya dialect and English, with both Arabic script and Arabizi (Latin-transliterated) Hassaniya included.
Dataset Details
Name: dah
Languages: Hassaniya Arabic (in both Arabic and Latin forms), English
Creators:
Ahlam Abdelkader
Emani Babe
Oumoukelthoum Sidenna
Adapted from: Tatoeba Project parallel sentences
Reviewed by: Founders of the… See the full description on the dataset page: https://huggingface.co/datasets/hassan-IA/dah.CoTCoT Datasets from Google's FLan Dataset
protein-compound-affinity-esm2-molformerquestion_generatorThis dataset is made up of data taken from SQuAD v2.0, RACE, CoQA, and MSMARCO. Some examples have been filtered out of the original datasets and others have been modified.
There are two fields; question and text. The question field contains the question, and the text field contains both the answer and the context in the following format:
"<answer> (answer text) <context> (context text)"
The and are included as special tokens in the question generator's tokenizer.
This dataset is intended to… See the full description on the dataset page: https://huggingface.co/datasets/iarfmoose/question_generator.iac-eval-v2
IaC-Eval v2
Modernised Terraform code-generation benchmark — 186 tasks, Terraform 1.15 + OPA 1.16 (Rego v1).
An updated and extended version of the IaC-Eval NeurIPS 2024 benchmark.
Scoring is deterministic: the generated HCL either passes terraform plan + opa eval, or it doesn't — no LLM-as-judge.
Dataset summary
Field
Value
Tasks
186 (AWS only)
Difficulty
1–6 (distribution: 1→35, 2→40, 3→51, 4→22, 5→9, 6→13)
AWS services
34 distinct
Terraform… See the full description on the dataset page: https://huggingface.co/datasets/iac-eval-v2/iac-eval-v2.Azure_IaC_testGlobal-Population-Data
List of Countries and Dependencies by Population
This dataset contains population-related information for countries and dependencies, scraped from Wikipedia. The dataset includes the following columns:
Location: The country or dependency name.
Population: Total population count.
% of World: The percentage of the world's population this country or dependency represents.
Date: The date of the population estimate.
Source: Whether the source is official or derived from the United… See the full description on the dataset page: https://huggingface.co/datasets/iamramzan/Global-Population-Data.News-Article-Categorization_IAB
Article and Category Dataset
Overview
This dataset contains a collection of articles, primarily news articles, along with their respective IAB (Interactive Advertising Bureau) categories. It can be a valuable resource for various natural language processing (NLP) tasks, including text classification, text generation, and more.
Dataset Information
Number of Samples: 871,909
Number of Categories: 26
Column Information
text: The text of the article.… See the full description on the dataset page: https://huggingface.co/datasets/shishir-dwi/News-Article-Categorization_IAB.iae-cnae-2025
IAE & CNAE 2025 Open Dataset (Spain)
Three machine-readable Spanish tax-classification catalogs. Maintained by conversoriaecnae.es. License: CC-BY-4.0 — please attribute conversoriaecnae.es when reusing.
Files
File
Rows
Source
data/v1/iae.{csv,jsonl}
1,189
AEAT — Tarifas IAE
data/v1/cnae_2025.{csv,jsonl}
1,060
INE / Real Decreto 10/2025 Art. 6
data/v1/correspondencia_cnae2009_cnae2025.{csv,jsonl}
1,010
INE / RD 10/2025 Art. 7(c)
Schema… See the full description on the dataset page: https://huggingface.co/datasets/conversoriaecnae/iae-cnae-2025.reddit-self-medication-claim-dataset
Reddit Self-Medication Claim Dataset
Dataset Summary
The Reddit Self-Medication Claim Dataset is an annotated NLP dataset designed to study self-medication claims expressed in informal online health discussions.
The dataset focuses on identifying whether Reddit posts contain self-medication related claims, and further distinguishing between explicit and implicit expressions of self-medication behavior.
This dataset was created as part of an independent research… See the full description on the dataset page: https://huggingface.co/datasets/iamjayeshc/reddit-self-medication-claim-dataset.clef2024_checkthat_task1_en
Bibtex
@inproceedings{Hasanain:CLEF:24,
author = {Maram Hasanain and
Reem Suwaileh and
Sanne Weering and
Chengkai Li and
Tommaso Caselli and
Wajdi Zaghouani and
Alberto Barr{\'{o}}n{-}Cede{\~{n}}o and
Preslav Nakov and
Firoj Alam},
editor = {Guglielmo Faggioli and
Nicola Ferro and
Petra Galusc{\'{a}}kov{\'{a}} and… See the full description on the dataset page: https://huggingface.co/datasets/iai-group/clef2024_checkthat_task1_en.Extended_Shellcode_IA32
Shellcode_IA32
Shellcode_IA32 is a dataset containing more than 20 years of shellcodes from a variety of sources and is the largest collection of shellcodes in assembly available to date. We are currently extending the dataset. Up to now, we released three versions of the dataset.
Shellcode_IA32 was presented for the first time in the paper Shellcode_IA32: A Dataset for Automatic Shellcode Generation, accepted to the 1st Workshop on Natural Language Processing for Programming… See the full description on the dataset page: https://huggingface.co/datasets/OSS-forge/Extended_Shellcode_IA32.1_exploitsclef2024_checkthat_task1_es
Bibtex
@inproceedings{Hasanain:CLEF:24,
author = {Maram Hasanain and
Reem Suwaileh and
Sanne Weering and
Chengkai Li and
Tommaso Caselli and
Wajdi Zaghouani and
Alberto Barr{\'{o}}n{-}Cede{\~{n}}o and
Preslav Nakov and
Firoj Alam},
editor = {Guglielmo Faggioli and
Nicola Ferro and
Petra Galusc{\'{a}}kov{\'{a}} and… See the full description on the dataset page: https://huggingface.co/datasets/iai-group/clef2024_checkthat_task1_es.IAST-corpus
Dataset Details
Dataset created by transliterating existing datasets to IAST by means of IAST transliteration library
Languages include Sanskrit, Hindi, Odia, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Marathi, Bhojpuri, Nepali.
Pre-existing dataset source(s) - Wikipedia.
Across all subsets, 'source', 'target', 'source_lang', 'target_lang', 'source' columns are common.
This is just a hobby dataset, but it should abide by the licenses of the input dataset(s).
RadEvalExpertDatasetsovereign-states-dataset
Sovereign States Dataset
This dataset provides a comprehensive list of sovereign states, along with their common and formal names, membership within the UN system, and details on sovereignty disputes and recognition status. The data was originally scraped from Wikipedia's List of Sovereign States and processed for clarity and usability.
Dataset Features
Common Name: The commonly used name of the country or state.
Formal Name: The official/formal name of the country… See the full description on the dataset page: https://huggingface.co/datasets/iamramzan/sovereign-states-dataset.Azure_IaC_reducedalpaca-cleaned-albanian
