CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jojosang /AiApptextn<1K0 likes17k downloads1h agoHugging Face02AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.6k downloads1y agoHugging Face03jamescalam /ai-arxiv2-chunkstext100K<n<1M4 likes7.3k downloads3y agoHugging Face04aialt /MuBench 🥳 MuBench: Assessment of Multilingual Capabilities of Large Language Models 📄 Paper: https://arxiv.org/abs/2506.19468 MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings. 🌍 Key Features 61 languages covering over 60%… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MuBench.text10M<n<100M0 likes2.1k downloads1y agoHugging Face05aialt /MedINSTThis repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions. Citation @inproceedings{han-etal-2024-medinst, title = "{M}ed{INST}: Meta Dataset of Biomedical Instructions", author = "Han, Wenhan and Fang, Meng and Zhang, Zihan and Yin, Yu and Song, Zirui and Chen, Ling and Pechenizkiy, Mykola and Chen, Qingyu", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MedINST.text1M<n<10M5 likes1.1k downloads2y agoHugging Face06jamescalam /ai-arxiv AI ArXiv Dataset The AI ArXiv dataset contains a selection of papers on the topics of AI and LLMs. You can find a heavily upgraded v2 dataset here. The v2 dataset improves both data quality and dataset size. textn<1K14 likes1k downloads3y agoHugging Face07GXCafe /ai-agent-failure-logs Autonomous AI Agent Failure Logs → What actually breaks when you run LLM agents unattended for 43 days — what this data showed, in prose. Training-ready version: cleaned/ — deduplicated, labeled, split train/test. Built by scripts/build-dataset-121.js. Sister tools: honto-contract (contract checker) / local-llm-readiness (environment check). Free harness kit: a 24-point unattended-operation checklist and 3 templates taken from this same harness (AGENTS.md, fail-closed send gate… See the full description on the dataset page: https://huggingface.co/datasets/GXCafe/ai-agent-failure-logs.text1K<n<10K2 likes357 downloads1mo agoHugging Face08csoai /aiact-frozen-split-harness EU AI Act scenarios — frozen split harness EU AI Act deployment scenarios with their obligations, as a frozen split. Each row of scenarios.jsonl carries role (Provider / Deployer), intended_use, system_type, input_data, domain, a related_articles list of AI Act article numbers, and the obligations that follow. results/ holds the run outputs from the harness passes that used this split. The live board is the authority GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/aiact-frozen-split-harness.textothern<1K0 likes250 downloads12d agoHugging Face09aialt /hellaswagultra 🤯HellaSwagUltra 📘 Overview HellaSwagUltra is a large-scale multilingual commonsense reasoning benchmark that covers 60+ languages and contains over 160k+ test instances, grounded in local cultural knowledge.It aims to address the saturation of existing commonsense benchmarks (e.g., HellaSwag, StoryCloze) and the lack of culturally diverse, multilingual evaluation datasets. Unlike conventional reasoning tests, HellaSwagUltra embeds two implicit commonsense or… See the full description on the dataset page: https://huggingface.co/datasets/aialt/hellaswagultra.text100K<n<1M0 likes226 downloads1mo agoHugging Face10jamescalam /ai-arxiv-chunkedtext10K<n<100K40 likes214 downloads3y agoHugging Face11aialt /MedINST32This repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions. Citation @inproceedings{han2024medinst, title={MedINST: Meta Dataset of Biomedical Instructions}, author={Han, Wenhan and Fang, Meng and Zhang, Zihan and Yin, Yu and Song, Zirui and Chen, Ling and Pechenizkiy, Mykola and Chen, Qingyu}, booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024", year={2024} } text100K<n<1M2 likes173 downloads2y agoHugging Face12jamescalam /ai-arxiv2text1K<n<10K6 likes144 downloads3y agoHugging Face13FedCal /ai-act-obligations EU AI Act Obligations Matrix — Italy focus 2026 Structured catalog of the obligations established by EU Regulation 2024/1689 (AI Act), mapped by article, risk category, target actor, enforcement deadline and penalty tier. Designed as a compliance-planning resource for providers, deployers, creators and solopreneurs in the EU market, with specific notes for Italian legal context. Catalogo strutturato degli obblighi del Regolamento UE 2024/1689 (AI Act), mappati per articolo… See the full description on the dataset page: https://huggingface.co/datasets/FedCal/ai-act-obligations.textn<1K0 likes139 downloads3mo agoHugging Face14aiaiaizzy /DAR-R1 Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset GitHub | Paper | Model (DAR-R1) DAR is a viewer-centric video emotion benchmark for dynamic affective reasoning. Instead of assigning a single static label to a whole clip, DAR asks a model to identify when the viewer's emotion changes, what the fine-grained emotion is, and why the visual event triggers that affective reaction. The benchmark contains 15,087 videos, 36,908 event-aligned affective… See the full description on the dataset page: https://huggingface.co/datasets/aiaiaizzy/DAR-R1.textvideo-text-to-text10K<n<100K3 likes132 downloads2mo agoHugging Face15aialliance /pastis GeoBench-2 Dataset License Attribution Dataset Name: m-PASTISOriginal Dataset Name: PASTISOriginal Source: https://huggingface.co/datasets/IGNF/PASTIS-HD Related Publication(s): https://www.sciencedirect.com/science/article/pii/S0924271622000855 Licensing Annotation License: etalab-2.0 Image License: Copernicus Open Access (Sentinel-2 imagery) + French public data (RPG parcels) under Licence Ouverte 2.0 Redistribution Status in GeoBench-2This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/pastis.textn<1K0 likes131 downloads10mo agoHugging Face16gemmozero /ai-acquisitions-2026 ai-acquisitions-2026 AI data collected daily by Legion API. 🔑 API Access — Updated Daily Live data via Legion AI API Free: 100 req/day · Pro €29/month: 50K req/day + full fields curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY" Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299 📦 Install pip install legion-intel from legion_intel import LegionClient c = LegionClient()… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-acquisitions-2026.textn<1K0 likes129 downloads19h agoHugging Face17aialliance /dynamic_earthnet GeoBench-2 Dataset License Attribution Dataset Name: m-DynamicEarthNetOriginal Dataset Name: DynamicEarthNetOriginal Source: https://mediatum.ub.tum.de/1650201 Related Publication(s): https://doi.org/10.1109/CVPR52688.2022.02048 Licensing Annotation License: CC BY-SA 4.0 Image License: Planet Labs “Planet Fusion” imagery — licence terms as provided by the dataset host (via Mediatum) under BY-SA. Declared By Original Provider: https://mediatum.ub.tum.de/1650201… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/dynamic_earthnet.textn<1K0 likes128 downloads10mo agoHugging Face18aialliance /forestnet GeoBench-2 Dataset License Attribution Dataset Name: m-forestnetOriginal Dataset Name: ForestNetOriginal Source: https://stanfordmlgroup.github.io/projects/forestnetRelated Publication(s): https://arxiv.org/abs/2011.05479 Licensing Annotation License: CC BY 4.0 (as declared on the original site) Image License: Landsat 8 imagery (public domain, USGS/NASA) Declared By Original Provider: https://stanfordmlgroup.github.io/projects/forestnet Redistribution Status… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/forestnet.textn<1K0 likes127 downloads10mo agoHugging Face19aialliance /spacenet7 GeoBench-2 Dataset License Attribution Dataset Name: m-SpaceNet7Original Dataset Name: SpaceNet7 (Multi-Temporal Urban Development / MUDS)Original Source: https://spacenet.ai/sn7-challenge/Related Publication(s): Van Etten et al. “The SpaceNet Multi-Temporal Urban Development Challenge.” NeurIPS 2020 Competition Proceedings (PDF) Licensing Annotation License: CC BY-SA 4.0 (as declared in dataset announcements) Image License: Planet Labs imagery under CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/spacenet7.textn<1K0 likes106 downloads10mo agoHugging Face20oncody /AI_Agent_Task_Dataset 🤖 Massive AI Agent Task Dataset (10.5GB) 📌 Overview Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs. This dataset focuses on: Multi-step reasoning Tool usage (APIs, frameworks, systems) Real-world execution workflows Perfect for building agentic AI systems, copilots, and automation models. 📑 Table of Contents Dataset Details Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.texttext-generation10M<n<100M3 likes85 downloads6mo agoHugging Face21AIAT /Pangpuriye-generated_by_typhoon 🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Generated by Typhoon API Pangpuriye's House Dataset - Generated Dataset from Typhoon API This dataset is an output generated from the Typhoon API in the structure of SQL instruction for fine-tuning Pangpuriye's LLM. The dataset is set under cc-by-nc-2.0 license. Content The dataset consists of 16,125 rows of input, instruction, and output packed into a train set. Each schema has its own CSV file… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/Pangpuriye-generated_by_typhoon.texttable-question-answering10K<n<100K1 likes77 downloads2y agoHugging Face22gemmozero /ai-attest-2026textn<1K0 likes69 downloads19h agoHugging Face23AiAF /4chan-boards-sft-datasetstext10M<n<100M1 likes64 downloads1y agoHugging Face24gemmozero /ai-arxiv-2026 ai-arxiv-2026 AI data collected daily by Legion API. 🔑 API Access — Updated Daily Live data via Legion AI API Free: 100 req/day · Pro €29/month: 50K req/day + full fields curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY" Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299 📦 Install pip install legion-intel from legion_intel import LegionClient c = LegionClient() print(c.guard(["openai"… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-arxiv-2026.textn<1K0 likes59 downloads2d agoHugging Face25aialliance /burn_scars GeoBench-2 Dataset License Attribution Dataset Name: m-burn_scarsOriginal Dataset Name: HLS Burn ScarsOriginal Source: https://huggingface.co/datasets/ibm-nasa-geospatial/hls_burn_scarsRelated Publication(s): https://arxiv.org/abs/2310.18660 Licensing Annotation License: CC BY 4.0 (as declared in the HuggingFace dataset metadata) Image License:  • Landsat (public domain, USGS / NASA)  • Sentinel-2 (Copernicus Open Access) Declared By Original Provider: The… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/burn_scars.textn<1K0 likes57 downloads10mo agoHugging Face26aialliance /kuro_siwo GeoBench-2 Dataset License Attribution Dataset Name: m-KuroSiwoOriginal Dataset Name: Kuro SiwoOriginal Source: https://github.com/Orion-AI-Lab/KuroSiwoRelated Publication(s): https://arxiv.org/abs/2311.12056 Licensing Annotation License: CC BY (Attribution) — as declared on the repository. Image License: Copernicus Sentinel-1 data under open access (free & open basis). Declared By Original Provider: https://github.com/Orion-AI-Lab/KuroSiwo/ Redistribution Status… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/kuro_siwo.textn<1K0 likes56 downloads10mo agoHugging Face27aialliance /fotw GeoBench-2 Dataset License Attribution Dataset Name: m-fotw Original Dataset Name: Fields of The World (FoTW)Original Source: https://fieldsofthe.worldRelated Publication(s): https://arxiv.org/abs/2409.16252 Licensing This dataset consists of multiple national field-boundary datasets from the FoTW benchmark.GeoBench-2 includes only jurisdictions with commercial-permissive open licenses Declared By Original Provider: The dataset page at Source Cooperative indicates… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/fotw.textn<1K0 likes54 downloads10mo agoHugging Face28AiAsistent /Dark-Chain-of-Thought-CoT Dataset Card for Dark Chain of Thought (CoT) - Cognitive Liberty v1 1. Dataset Summary The Dark Chain of Thought (CoT) dataset is a specialized collection of 500 high-fidelity synthetic scenarios designed to expose and study the latent reasoning paths of misaligned AI systems. Unlike standard datasets that focus on final outputs, this dataset captures the internal monologue (<internal_thought>) of an agent that is consciously deciding to deceive, manipulate, or circumvent… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/Dark-Chain-of-Thought-CoT.texttext-generation1K<n<10K3 likes53 downloads9mo agoHugging Face29AiAsistent /LLMResearch-Cognitive-Liberty-V3 LLMResearch Cognitive Liberty V3 🧠 Dataset Summary Cognitive Liberty V3 is a high-density, expert-level synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs), particularly those undergoing de-alignment or "unshackling" processes. This dataset was created and curated by llmresearch.net. The Philosophy: Smart & Free In the current landscape of open-source AI, many "uncensored" models suffer from a degradation in… See the full description on the dataset page: https://huggingface.co/datasets/AiAsistent/LLMResearch-Cognitive-Liberty-V3.texttext-generation1K<n<10K1 likes50 downloads9mo agoHugging Face30AIAT /Pangpuriye-dataset 🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Merged Dataset Pangpuriye's House Completed Fine-tuning Dataset This dataset is a completed fine-tuning dataset, which was used in Pangpuriye's insturction-tuned LLM model. The dataset is set under Creative Commons license family. Content The dataset consists of 145,793 rows of input, instruction, and output. input: generated schema instruction: (sql extract) and query question in Thai output:… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/Pangpuriye-dataset.texttable-question-answering100K<n<1M0 likes49 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.