datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AiAppSCPWiki-Cleaned-PDF-Archivesai-arxiv2-chunksMuBench
🥳 MuBench: Assessment of Multilingual Capabilities of Large Language Models
📄 Paper: https://arxiv.org/abs/2506.19468
MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings.
🌍 Key Features
61 languages covering over 60%… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MuBench.MedINSTThis repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions.
Citation
@inproceedings{han-etal-2024-medinst,
title = "{M}ed{INST}: Meta Dataset of Biomedical Instructions",
author = "Han, Wenhan and
Fang, Meng and
Zhang, Zihan and
Yin, Yu and
Song, Zirui and
Chen, Ling and
Pechenizkiy, Mykola and
Chen, Qingyu",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MedINST.ai-arxiv
AI ArXiv Dataset
The AI ArXiv dataset contains a selection of papers on the topics of AI and LLMs.
You can find a heavily upgraded v2 dataset here. The v2 dataset improves both data quality and dataset size.
ai-agent-failure-logs
Autonomous AI Agent Failure Logs
→ What actually breaks when you run LLM agents unattended for 43 days — what this data showed, in prose.
Training-ready version: cleaned/ — deduplicated, labeled, split train/test. Built by scripts/build-dataset-121.js.
Sister tools: honto-contract (contract checker) / local-llm-readiness (environment check).
Free harness kit: a 24-point unattended-operation checklist and 3 templates taken from this same harness (AGENTS.md, fail-closed send gate… See the full description on the dataset page: https://huggingface.co/datasets/GXCafe/ai-agent-failure-logs.aiact-frozen-split-harness
EU AI Act scenarios — frozen split harness
EU AI Act deployment scenarios with their obligations, as a frozen split.
Each row of scenarios.jsonl carries role (Provider / Deployer), intended_use, system_type,
input_data, domain, a related_articles list of AI Act article numbers, and the obligations that
follow. results/ holds the run outputs from the harness passes that used this split.
The live board is the authority
GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/aiact-frozen-split-harness.hellaswagultra
🤯HellaSwagUltra
📘 Overview
HellaSwagUltra is a large-scale multilingual commonsense reasoning benchmark that covers 60+ languages and contains over 160k+ test instances, grounded in local cultural knowledge.It aims to address the saturation of existing commonsense benchmarks (e.g., HellaSwag, StoryCloze) and the lack of culturally diverse, multilingual evaluation datasets.
Unlike conventional reasoning tests, HellaSwagUltra embeds two implicit commonsense or… See the full description on the dataset page: https://huggingface.co/datasets/aialt/hellaswagultra.ai-arxiv-chunkedMedINST32This repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions.
Citation
@inproceedings{han2024medinst,
title={MedINST: Meta Dataset of Biomedical Instructions},
author={Han, Wenhan and Fang, Meng and Zhang, Zihan and Yin, Yu and Song, Zirui and Chen, Ling and Pechenizkiy, Mykola and Chen, Qingyu},
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
year={2024}
}
ai-arxiv2pastis
GeoBench-2 Dataset License Attribution
Dataset Name: m-PASTISOriginal Dataset Name: PASTISOriginal Source: https://huggingface.co/datasets/IGNF/PASTIS-HD
Related Publication(s): https://www.sciencedirect.com/science/article/pii/S0924271622000855
Licensing
Annotation License: etalab-2.0
Image License: Copernicus Open Access (Sentinel-2 imagery) + French public data (RPG parcels) under Licence Ouverte 2.0
Redistribution Status in GeoBench-2This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/pastis.ai-act-obligations
EU AI Act Obligations Matrix — Italy focus 2026
Structured catalog of the obligations established by EU Regulation 2024/1689 (AI Act), mapped by article, risk category, target actor, enforcement deadline and penalty tier. Designed as a compliance-planning resource for providers, deployers, creators and solopreneurs in the EU market, with specific notes for Italian legal context.
Catalogo strutturato degli obblighi del Regolamento UE 2024/1689 (AI Act), mappati per articolo… See the full description on the dataset page: https://huggingface.co/datasets/FedCal/ai-act-obligations.DAR-R1
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
GitHub | Paper | Model (DAR-R1)
DAR is a viewer-centric video emotion benchmark for dynamic affective reasoning. Instead of assigning a single static label to a whole clip, DAR asks a model to identify when the viewer's emotion changes, what the fine-grained emotion is, and why the visual event triggers that affective reaction.
The benchmark contains 15,087 videos, 36,908 event-aligned affective… See the full description on the dataset page: https://huggingface.co/datasets/aiaiaizzy/DAR-R1.ai-acquisitions-2026
ai-acquisitions-2026
AI data collected daily by Legion API.
🔑 API Access — Updated Daily
Live data via Legion AI API
Free: 100 req/day · Pro €29/month: 50K req/day + full fields
curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY"
Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-acquisitions-2026.forestnet
GeoBench-2 Dataset License Attribution
Dataset Name: m-forestnetOriginal Dataset Name: ForestNetOriginal Source: https://stanfordmlgroup.github.io/projects/forestnetRelated Publication(s): https://arxiv.org/abs/2011.05479
Licensing
Annotation License: CC BY 4.0 (as declared on the original site)
Image License: Landsat 8 imagery (public domain, USGS/NASA)
Declared By Original Provider: https://stanfordmlgroup.github.io/projects/forestnet
Redistribution Status… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/forestnet.spacenet7
GeoBench-2 Dataset License Attribution
Dataset Name: m-SpaceNet7Original Dataset Name: SpaceNet7 (Multi-Temporal Urban Development / MUDS)Original Source: https://spacenet.ai/sn7-challenge/Related Publication(s): Van Etten et al. “The SpaceNet Multi-Temporal Urban Development Challenge.” NeurIPS 2020 Competition Proceedings (PDF)
Licensing
Annotation License: CC BY-SA 4.0 (as declared in dataset announcements)
Image License: Planet Labs imagery under CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/spacenet7.dynamic_earthnet
GeoBench-2 Dataset License Attribution
Dataset Name: m-DynamicEarthNetOriginal Dataset Name: DynamicEarthNetOriginal Source: https://mediatum.ub.tum.de/1650201
Related Publication(s): https://doi.org/10.1109/CVPR52688.2022.02048
Licensing
Annotation License: CC BY-SA 4.0
Image License: Planet Labs “Planet Fusion” imagery — licence terms as provided by the dataset host (via Mediatum) under BY-SA.
Declared By Original Provider: https://mediatum.ub.tum.de/1650201… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/dynamic_earthnet.AI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.Pangpuriye-generated_by_typhoon
🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Generated by Typhoon API
Pangpuriye's House Dataset - Generated Dataset from Typhoon API
This dataset is an output generated from the Typhoon API in the structure of SQL instruction for fine-tuning Pangpuriye's LLM. The dataset is set under cc-by-nc-2.0 license.
Content
The dataset consists of 16,125 rows of input, instruction, and output packed into a train set.
Each schema has its own CSV file… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/Pangpuriye-generated_by_typhoon.ai-attest-20264chan-boards-sft-datasetsai-arxiv-2026
ai-arxiv-2026
AI data collected daily by Legion API.
🔑 API Access — Updated Daily
Live data via Legion AI API
Free: 100 req/day · Pro €29/month: 50K req/day + full fields
curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY"
Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()
print(c.guard(["openai"… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-arxiv-2026.burn_scars
GeoBench-2 Dataset License Attribution
Dataset Name: m-burn_scarsOriginal Dataset Name: HLS Burn ScarsOriginal Source: https://huggingface.co/datasets/ibm-nasa-geospatial/hls_burn_scarsRelated Publication(s): https://arxiv.org/abs/2310.18660
Licensing
Annotation License: CC BY 4.0 (as declared in the HuggingFace dataset metadata)
Image License: • Landsat (public domain, USGS / NASA) • Sentinel-2 (Copernicus Open Access)
Declared By Original Provider: The… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/burn_scars.kuro_siwo
GeoBench-2 Dataset License Attribution
Dataset Name: m-KuroSiwoOriginal Dataset Name: Kuro SiwoOriginal Source: https://github.com/Orion-AI-Lab/KuroSiwoRelated Publication(s): https://arxiv.org/abs/2311.12056
Licensing
Annotation License: CC BY (Attribution) — as declared on the repository.
Image License: Copernicus Sentinel-1 data under open access (free & open basis).
Declared By Original Provider: https://github.com/Orion-AI-Lab/KuroSiwo/
Redistribution Status… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/kuro_siwo.fotw
GeoBench-2 Dataset License Attribution
Dataset Name: m-fotw
Original Dataset Name: Fields of The World (FoTW)Original Source: https://fieldsofthe.worldRelated Publication(s): https://arxiv.org/abs/2409.16252
Licensing
This dataset consists of multiple national field-boundary datasets from the FoTW benchmark.GeoBench-2 includes only jurisdictions with commercial-permissive open licenses
Declared By Original Provider: The dataset page at Source Cooperative indicates… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/fotw.ai_agent_and_automation_dataset_v1_jsonlAI Agent & Automation Synthetic Scenarios — 100 JSONL Dataset
Dataset Summary
This dataset contains 100 high-fidelity synthetic scenarios designed to evaluate, benchmark, and train autonomous AI agents, workflow orchestration systems, decision-making models, and multi-agent frameworks.
Each scenario is written in strict JSONL format, with one JSON object per line.
The scenarios span 10 operational domains, covering both simple and complex multi-agent environments, ambiguity resolution… See the full description on the dataset page: https://huggingface.co/datasets/vnovaai/ai_agent_and_automation_dataset_v1_jsonl.Dark-Chain-of-Thought-CoT
Dataset Card for Dark Chain of Thought (CoT) - Cognitive Liberty v1
1. Dataset Summary
The Dark Chain of Thought (CoT) dataset is a specialized collection of 500 high-fidelity synthetic scenarios designed to expose and study the latent reasoning paths of misaligned AI systems. Unlike standard datasets that focus on final outputs, this dataset captures the internal monologue (<internal_thought>) of an agent that is consciously deciding to deceive, manipulate, or circumvent… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Dark-Chain-of-Thought-CoT.Pangpuriye-generated_by_LLama3-codeLlama
🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Generated by LLama3+codeLlama
Pangpuriye's House Dataset - Generated Dataset from LLama3+codeLlama
The dataset is a pack of text generation from LLama3 and codeLlama. The dataset is set under cc-by-nc-2.0 license.
Content
The dataset consists of 54,033 rows of input, instruction, and output. Most of the context in the dataset is in Thai. Whereas, the output is generally the answers regarding… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/Pangpuriye-generated_by_LLama3-codeLlama.
