datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AiAppSCPWiki-Cleaned-PDF-Archivesai-arxiv2-chunksSCPWiki-Archive-02-March-2025-DatasetsAIA_12hour_512x512GEOBench-VLM
GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks
Summary
While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they fall short in addressing the unique demands of geospatial applications. Generic VLM benchmarks are not designed to handle the complexities of geospatial data, which is critical for applications such as environmental monitoring, urban planning, and disaster management. Some of the unique… See the full description on the dataset page: https://huggingface.co/datasets/aialliance/GEOBench-VLM.JFK-Assassination-Records-2025-Documents-ReleaseMuBench
🥳 MuBench: Assessment of Multilingual Capabilities of Large Language Models
📄 Paper: https://arxiv.org/abs/2506.19468
MuBench is a meta-dataset for evaluating the multilingual capabilities of large language models (LLMs) across 61 languages and 3.9M aligned samples.It provides a unified framework to assess understanding, reasoning, factual knowledge, and truthfulness in both single-language and code-switched settings.
🌍 Key Features
61 languages covering over 60%… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MuBench.polynews-parallel
Dataset Card for PolyNewsParallel
Dataset Summary
PolyNewsParallel is a multilingual paralllel dataset containing news titles for 833 language pairs. It covers 64 languages and 17 scripts.
Uses
This dataset can be used for machine translation or text retrieval.
Languages
There are 64 languages avaiable:
Code
Language
Script
amh_Ethi
Amharic
Ethiopic
arb_Arab
Modern Standard Arabic
Arabic
ayr_Latn
Central Aymara
Latin
bam_Latn
Bambara… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews-parallel.bishkek-transport
Bishkek Public Transport
Open, continuously-growing data on the public-transport network of Bishkek,
the capital of the Kyrgyz Republic. The data originates from the Bishkek
mayoralty's public-transport monitoring system (the same feed behind the city's
official live transit map and its "My City" mobile service). Only publicly
visible transit information is included: stop locations and the live positions of
buses, trolleybuses/electric buses, and marshrutkas (shared minibuses).… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/bishkek-transport.tmpAll-Prompt-JailbreakMedINSTThis repository contains the data of the paper MedINST: Meta Dataset of Biomedical Instructions.
Citation
@inproceedings{han-etal-2024-medinst,
title = "{M}ed{INST}: Meta Dataset of Biomedical Instructions",
author = "Han, Wenhan and
Fang, Meng and
Zhang, Zihan and
Yin, Yu and
Song, Zirui and
Chen, Ling and
Pechenizkiy, Mykola and
Chen, Qingyu",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen… See the full description on the dataset page: https://huggingface.co/datasets/aialt/MedINST.ai-arxiv
AI ArXiv Dataset
The AI ArXiv dataset contains a selection of papers on the topics of AI and LLMs.
You can find a heavily upgraded v2 dataset here. The v2 dataset improves both data quality and dataset size.
mjnj384moviesKJV-LLM-Datasetsai-api-pricing
AI API Pricing Dataset
This Hugging Face dataset is the machine-readable distribution of the public AI API pricing records published by AICostBudget. It is not a separately curated subset: train.csv, prices.csv, and prices.json are generated from the same Pricing V2 public projection used by the AICostBudget Dataset page and download APIs.
Prices change frequently. Verify production billing decisions against the provider pricing page, contract, billing dashboard, and invoice.… See the full description on the dataset page: https://huggingface.co/datasets/aicostbudget-ai/ai-api-pricing.polynews
Dataset Card for PolyNews
Dataset Summary
PolyNews is a multilingual dataset containing news titles in 77 languages and 19 scripts.
Uses
This dataset can be used for domain adaptation of language models, language modeling or text generation.
Languages
There are 77 languages available:
Code
Language
Script
#Articles (K)
amh_Ethi
Amharic
Ethiopic
0.551
arb_Arab
Modern Standard Arabic
Arabic
10.882
ayr_Latn
Central Aymara
Latin
12.878… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/polynews.ai-agent-security-incidents
AI Agent Security Incident Database v0.1
A structured, machine-readable database of 1405 confirmed AI agent security incidents, collected and classified automatically.
What is this?
Every time an AI agent causes unintended harm — escaping a sandbox, exploiting an API, taking unauthorized actions, exfiltrating data — this database captures it.
This is not a list of theoretical risks. Every entry describes something that actually happened, with a verifiable source… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents.384imagenet1kk1921kk imagenet with auradiffusion vae 192x192
see https://huggingface.co/datasets/AiArtLab/imagenet1kk192/blob/main/imagenet-1kk/sdxs1b-imagenet.ipynb
7681152576quantum-video
Dataset Card for Dataset Name
quantum suite video
quantum suite video dataset, used to train the quantum suite video model.
Dataset Details
will upload, collecting data
Dataset Description
This dataset is aimed to be curated for the quantum suite video dataset with video from wikimedia (Cc allowing commercial use), this dataset allows commercial use, and will become useful for you to use in your video models,
the size is aimed to be 2.5tb, after we… See the full description on the dataset page: https://huggingface.co/datasets/ai-api-key-free-finder/quantum-video.ai-agent-failure-logs
Autonomous AI Agent Failure Logs
→ What actually breaks when you run LLM agents unattended for 43 days — what this data showed, in prose.
Training-ready version: cleaned/ — deduplicated, labeled, split train/test. Built by scripts/build-dataset-121.js.
Sister tools: honto-contract (contract checker) / local-llm-readiness (environment check).
Free harness kit: a 24-point unattended-operation checklist and 3 templates taken from this same harness (AGENTS.md, fail-closed send gate… See the full description on the dataset page: https://huggingface.co/datasets/GXCafe/ai-agent-failure-logs.xMINDlarge
Dataset Card for xMINDlarge
Dataset Summary
xMINDlarge is an open, large-scale multi-parallel news dataset for multi- and cross-lingual news recommendation.
It is derived from the English MINDlarge dataset using open-source neural machine translation (i.e., NLLB 3.3B).
For the small version of the dataset, see xMINDsmall.
Uses
This dataset can be used for machine translation, text retrieval, or as a benchmark dataset for news recommendation.… See the full description on the dataset page: https://huggingface.co/datasets/aiana94/xMINDlarge.house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.work-arena
