CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jojosang /AiApptextn<1K0 likes18k downloads56m agoHugging Face02AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.5k downloads1y agoHugging Face03jamescalam /ai-arxiv2-chunkstext100K<n<1M4 likes6.7k downloads3y agoHugging Face04AiAF /SCPWiki-Archive-02-March-2025-Datasetstextn<1K0 likes6.1k downloads2y agoHugging Face05aialliance /GEOBench-VLM GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks Summary While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, which is critical for applications such as environmental monitoring, urban planning, and disaster management. Some of the unique… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/GEOBench-VLM.image1K<n<10K17 likes3.8k downloads1y agoHugging Face06aialt /MuBench 🥳 MuBench: Assessment of Multilingual Capabilities of Large Language Models 📄 Paper: https://arxiv.org/abs/2506.19468 MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings. 🌍 Key Features 61 languages covering over 60%… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MuBench.text10M<n<100M0 likes2k downloads1y agoHugging Face07aiana94 /polynews-parallel Dataset Card for PolyNewsParallel Dataset Summary PolyNewsParallel is a multilingual paralllel dataset containing news titles for 833 language pairs. It covers 64 languages and 17 scripts. Uses This dataset can be used for machine translation or text retrieval. Languages There are 64 languages avaiable: Code Language Script amh_Ethi Amharic Ethiopic arb_Arab Modern Standard Arabic Arabic ayr_Latn Central Aymara Latin bam_Latn Bambara… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews-parallel.imagetranslation1M<n<10M15 likes1.8k downloads2y agoHugging Face08aiacademy-kg /bishkek-transport Bishkek Public Transport Open, continuously-growing data on the public-transport network of Bishkek, the capital of the Kyrgyz Republic. The data originates from the Bishkek mayoralty's public-transport monitoring system (the same feed behind the city's official live transit map and its "My City" mobile service). Only publicly visible transit information is included: stop locations and the live positions of buses, trolleybuses/electric buses, and marshrutkas (shared minibuses).… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/bishkek-transport.tabulartime-series-forecasting100M<n<1B0 likes1.7k downloads55m agoHugging Face09AiArtLab /tmptabular1M<n<10M0 likes1.6k downloads10mo agoHugging Face10AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face11aialt /MedINSTThis repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions. Citation @inproceedings{han-etal-2024-medinst, title = "{M}ed{INST}: Meta Dataset of Biomedical Instructions", author = "Han, Wenhan and Fang, Meng and Zhang, Zihan and Yin, Yu and Song, Zirui and Chen, Ling and Pechenizkiy, Mykola and Chen, Qingyu", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MedINST.text1M<n<10M5 likes1k downloads2y agoHugging Face12jamescalam /ai-arxiv AI ArXiv Dataset The AI ArXiv dataset contains a selection of papers on the topics of AI and LLMs. You can find a heavily upgraded v2 dataset here. The v2 dataset improves both data quality and dataset size. textn<1K14 likes1k downloads3y agoHugging Face13AiArtLab /mjnj384tabular1M<n<10M0 likes890 downloads1y agoHugging Face14aicostbudget-ai /ai-api-pricing AI API Pricing Dataset This Hugging Face dataset is the machine-readable distribution of the public AI API pricing records published by AICostBudget. It is not a separately curated subset: train.csv, prices.csv, and prices.json are generated from the same Pricing V2 public projection used by the AICostBudget Dataset page and download APIs. Prices change frequently. Verify production billing decisions against the provider pricing page, contract, billing dashboard, and invoice.… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.tabularn<1K0 likes518 downloads2d agoHugging Face15AiAF /KJV-LLM-Datasetstextn<1K0 likes515 downloads2y agoHugging Face16aiana94 /polynews Dataset Card for PolyNews Dataset Summary PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts. Uses This dataset can be used for domain adaptation of language models, language modeling or text generation. Languages There are 77 languages available: Code Language Script #Articles (K) amh_Ethi Amharic Ethiopic 0.551 arb_Arab Modern Standard Arabic Arabic 10.882 ayr_Latn Central Aymara Latin 12.878… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews.textfill-mask1M<n<10M6 likes481 downloads2y agoHugging Face17AiArtLab /imagenet1kk1921kk imagenet with auradiffusion vae 192x192 see https://huggingface.co/datasets/AiArtLab/imagenet1kk192/blob/main/imagenet-1kk/sdxs1b-imagenet.ipynb text100K<n<1M0 likes478 downloads2y agoHugging Face18gemmozero /ai-agent-security-incidents AI Agent Security Incident Database v0.1 A structured, machine-readable database of 1419 confirmed AI agent security incidents, collected and classified automatically. What is this? Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it. This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.tabulartext-classification1K<n<10K1 likes456 downloads10h agoHugging Face19AI-Art-Collab /576tabular1M<n<10M0 likes416 downloads1y agoHugging Face20aiana94 /xMINDlarge Dataset Card for xMINDlarge Dataset Summary xMINDlarge is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation. It is derived from the English MINDlarge dataset using open-source neural machine translation (i.e., NLLB 3.3B). For the small version of the dataset, see xMINDsmall. Uses This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDlarge.texttranslation1M<n<10M4 likes344 downloads2y agoHugging Face21GXCafe /ai-agent-failure-logs Autonomous AI Agent Failure Logs → What actually breaks when you run LLM agents unattended for 43 days — what this data showed, in prose. Training-ready version: cleaned/ — deduplicated, labeled, split train/test. Built by scripts/build-dataset-121.js. Sister tools: honto-contract (contract checker) / local-llm-readiness (environment check). Free harness kit: a 24-point unattended-operation checklist and 3 templates taken from this same harness (AGENTS.md, fail-closed send gate… See the full description on the dataset page: https://huggingface.co/datasets/GXCafe/ai-agent-failure-logs.text1K<n<10K2 likes338 downloads1mo agoHugging Face22AiArtLab /1152tabular100K<n<1M0 likes335 downloads1y agoHugging Face23aiacademy-kg /house_kg_full_dataset house.kg — Kyrgyzstan Real Estate (multimodal) A complete snapshot of house.kg, the largest real-estate board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller identities, agency ratings, reviews — and 227,294 photographs. Field names are English; values are kept in the original language (Russian/Kyrgyz), exactly as the site renders them. 💻 Scraper source code on GitHub → The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.imagetabular-regression100K<n<1M0 likes321 downloads3mo agoHugging Face24AI-Art-Collab /640tabular100K<n<1M0 likes300 downloads11mo agoHugging Face25aialliance /biomassters GeoBench-2 Dataset License Attribution Dataset Name: m-BioMasstersOriginal Dataset Name: BioMasstersOriginal Source: https://huggingface.co/datasets/nascetti-a/BioMassters Related Publication(s): https://nascetti-a.github.io/BioMasster/ Licensing Annotation License: CC BY 4.0 (declared on the HuggingFace dataset page) Image License: Copernicus Sentinel-1 & Sentinel-2 imagery (open access under CC BY-SA 3.0 IGO) Declared By Original Provider:… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/biomassters.textn<1K0 likes292 downloads9mo agoHugging Face26aiai-laboratory /vietspeech-train-streamingtext100K<n<1M0 likes275 downloads2mo agoHugging Face27Ateeqq /AI-and-Human-Generated-Text AI & Human Generated Text I am Using this dataset for AI Text Detection for https://exnrt.com. Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA Description The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.texttext-classification10K<n<100K24 likes274 downloads2y agoHugging Face28AI-Art-Collab /dataset640tabular1M<n<10M0 likes262 downloads10mo agoHugging Face29csoai /aiact-frozen-split-harness EU AI Act scenarios — frozen split harness EU AI Act deployment scenarios with their obligations, as a frozen split. Each row of scenarios.jsonl carries role (Provider / Deployer), intended_use, system_type, input_data, domain, a related_articles list of AI Act article numbers, and the obligations that follow. results/ holds the run outputs from the harness passes that used this split. The live board is the authority GET https://councilof.ai/api/gspc — quote… See the full description on the dataset page: https://huggingface.co/datasets/csoai/aiact-frozen-split-harness.textothern<1K0 likes257 downloads12d agoHugging Face30aiana94 /xMINDsmall Dataset Card for xMINDsmall Dataset Summary xMINDsmall is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation. It is derived from the English MINDsmall dataset using open-source neural machine translation (i.e., NLLB 3.3B). For the large version of the dataset, see xMINDlarge. Uses This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDsmall.texttranslation1M<n<10M4 likes255 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.