CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jojosang /AiApptextn<1K0 likes18k downloads2h agoHugging Face02AiAF /SCPWiki-Cleaned-PDF-Archivesdocumenttext-generationn<1K1 likes7.5k downloads1y agoHugging Face03jamescalam /ai-arxiv2-chunkstext100K<n<1M4 likes6.7k downloads3y agoHugging Face04AiAF /SCPWiki-Archive-02-March-2025-Datasetstextn<1K0 likes6.1k downloads2y agoHugging Face05hrrsmjd /AIA_12hour_512x5122 likes4k downloads2y agoHugging Face06aialliance /GEOBench-VLM GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks Summary While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, which is critical for applications such as environmental monitoring, urban planning, and disaster management. Some of the unique… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/GEOBench-VLM.image1K<n<10K17 likes3.8k downloads1y agoHugging Face07AiAF /JFK-Assassination-Records-2025-Documents-Release0 likes3.3k downloads1y agoHugging Face08aialt /MuBench 🥳 MuBench: Assessment of Multilingual Capabilities of Large Language Models 📄 Paper: https://arxiv.org/abs/2506.19468 MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings. 🌍 Key Features 61 languages covering over 60%… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MuBench.text10M<n<100M0 likes2k downloads1y agoHugging Face09aiana94 /polynews-parallel Dataset Card for PolyNewsParallel Dataset Summary PolyNewsParallel is a multilingual paralllel dataset containing news titles for 833 language pairs. It covers 64 languages and 17 scripts. Uses This dataset can be used for machine translation or text retrieval. Languages There are 64 languages avaiable: Code Language Script amh_Ethi Amharic Ethiopic arb_Arab Modern Standard Arabic Arabic ayr_Latn Central Aymara Latin bam_Latn Bambara… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews-parallel.imagetranslation1M<n<10M15 likes1.8k downloads2y agoHugging Face10aiacademy-kg /bishkek-transport Bishkek Public Transport Open, continuously-growing data on the public-transport network of Bishkek, the capital of the Kyrgyz Republic. The data originates from the Bishkek mayoralty's public-transport monitoring system (the same feed behind the city's official live transit map and its "My City" mobile service). Only publicly visible transit information is included: stop locations and the live positions of buses, trolleybuses/electric buses, and marshrutkas (shared minibuses).… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/bishkek-transport.tabulartime-series-forecasting100M<n<1B0 likes1.7k downloads3h agoHugging Face11AiArtLab /tmptabular1M<n<10M0 likes1.6k downloads10mo agoHugging Face12AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face13aialt /MedINSTThis repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions. Citation @inproceedings{han-etal-2024-medinst, title = "{M}ed{INST}: Meta Dataset of Biomedical Instructions", author = "Han, Wenhan and Fang, Meng and Zhang, Zihan and Yin, Yu and Song, Zirui and Chen, Ling and Pechenizkiy, Mykola and Chen, Qingyu", editor = "Al-Onaizan, Yaser and Bansal, Mohit and Chen… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MedINST.text1M<n<10M5 likes1k downloads2y agoHugging Face14jamescalam /ai-arxiv AI ArXiv Dataset The AI ArXiv dataset contains a selection of papers on the topics of AI and LLMs. You can find a heavily upgraded v2 dataset here. The v2 dataset improves both data quality and dataset size. textn<1K14 likes1k downloads3y agoHugging Face15AiArtLab /mjnj384tabular1M<n<10M0 likes890 downloads1y agoHugging Face16aiagentvn4u /movies0 likes726 downloads3mo agoHugging Face17aicostbudget-ai /ai-api-pricing AI API Pricing Dataset This Hugging Face dataset is the machine-readable distribution of the public AI API pricing records published by AICostBudget. It is not a separately curated subset: train.csv, prices.csv, and prices.json are generated from the same Pricing V2 public projection used by the AICostBudget Dataset page and download APIs. Prices change frequently. Verify production billing decisions against the provider pricing page, contract, billing dashboard, and invoice.… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.tabularn<1K0 likes518 downloads2d agoHugging Face18AiAF /KJV-LLM-Datasetstextn<1K0 likes515 downloads2y agoHugging Face19aiana94 /polynews Dataset Card for PolyNews Dataset Summary PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts. Uses This dataset can be used for domain adaptation of language models, language modeling or text generation. Languages There are 77 languages available: Code Language Script #Articles (K) amh_Ethi Amharic Ethiopic 0.551 arb_Arab Modern Standard Arabic Arabic 10.882 ayr_Latn Central Aymara Latin 12.878… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews.textfill-mask1M<n<10M6 likes481 downloads2y agoHugging Face20AiArtLab /imagenet1kk1921kk imagenet with auradiffusion vae 192x192 see https://huggingface.co/datasets/AiArtLab/imagenet1kk192/blob/main/imagenet-1kk/sdxs1b-imagenet.ipynb text100K<n<1M0 likes478 downloads2y agoHugging Face21AiArtLab /3840 likes460 downloads1y agoHugging Face22gemmozero /ai-agent-security-incidents AI Agent Security Incident Database v0.1 A structured, machine-readable database of 1419 confirmed AI agent security incidents, collected and classified automatically. What is this? Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it. This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.tabulartext-classification1K<n<10K1 likes456 downloads8h agoHugging Face23ai-api-key-free-finder /quantum-video Dataset Card for Dataset Name quantum suite video quantum suite video dataset, used to train the quantum suite video model. Dataset Details will upload, collecting data Dataset Description This dataset is aimed to be curated for the quantum suite video dataset with video from wikimedia (Cc allowing commercial use), this dataset allows commercial use, and will become useful for you to use in your video models, the size is aimed to be 2.5tb, after we… See the full description on the dataset page: https://huggingface.co/datasets/ai-api-key-free-finder/quantum-video.text-to-video0 likes431 downloads27d agoHugging Face24AI-Art-Collab /576tabular1M<n<10M0 likes416 downloads1y agoHugging Face25aiana94 /xMINDlarge Dataset Card for xMINDlarge Dataset Summary xMINDlarge is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation. It is derived from the English MINDlarge dataset using open-source neural machine translation (i.e., NLLB 3.3B). For the small version of the dataset, see xMINDsmall. Uses This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDlarge.texttranslation1M<n<10M4 likes344 downloads2y agoHugging Face26GXCafe /ai-agent-failure-logs Autonomous AI Agent Failure Logs → What actually breaks when you run LLM agents unattended for 43 days — what this data showed, in prose. Training-ready version: cleaned/ — deduplicated, labeled, split train/test. Built by scripts/build-dataset-121.js. Sister tools: honto-contract (contract checker) / local-llm-readiness (environment check). Free harness kit: a 24-point unattended-operation checklist and 3 templates taken from this same harness (AGENTS.md, fail-closed send gate… See the full description on the dataset page: https://huggingface.co/datasets/GXCafe/ai-agent-failure-logs.text1K<n<10K2 likes338 downloads1mo agoHugging Face27AiArtLab /1152tabular100K<n<1M0 likes335 downloads1y agoHugging Face28aiacademy-kg /house_kg_full_dataset house.kg — Kyrgyzstan Real Estate (multimodal) A complete snapshot of house.kg, the largest real-estate board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller identities, agency ratings, reviews — and 227,294 photographs. Field names are English; values are kept in the original language (Russian/Kyrgyz), exactly as the site renders them. 💻 Scraper source code on GitHub → The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.imagetabular-regression100K<n<1M0 likes321 downloads3mo agoHugging Face29AI-Art-Collab /640tabular100K<n<1M0 likes300 downloads11mo agoHugging Face30aialliance /biomassters GeoBench-2 Dataset License Attribution Dataset Name: m-BioMasstersOriginal Dataset Name: BioMasstersOriginal Source: https://huggingface.co/datasets/nascetti-a/BioMassters Related Publication(s): https://nascetti-a.github.io/BioMasster/ Licensing Annotation License: CC BY 4.0 (declared on the HuggingFace dataset page) Image License: Copernicus Sentinel-1 & Sentinel-2 imagery (open access under CC BY-SA 3.0 IGO) Declared By Original Provider:… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/biomassters.textn<1K0 likes292 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.