datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pile-uncopyrighted
Pile Uncopyrighted
In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA.
MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.wikiann
Dataset Card for WikiANN
Dataset Summary
WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, and test splits of Rahimi et al. (2019), which supports 176 of the 282 languages from the original WikiANN corpus.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/unimelb-nlp/wikiann.turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.unit4-students-scoresShareGPT_Vicuna_unfiltered
Dataset Card
This is a reupload of this dataset that was further cleaned by gozfarb.
un_pc
Dataset Card for United Nations Parallel Corpus
Dataset Summary
The United Nations Parallel Corpus is the first parallel corpus composed from United Nations documents published by the original data creator.
The parallel corpus consists of manually translated UN documents from the last 25 years (1990 to 2014)
for the six official UN languages, Arabic, Chinese, English, French, Russian, and Spanish.
The corpus is freely available for download under a liberal license.… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/un_pc.university-1652
University-1652: Drone-based Geo-localization Benchmark 🚁
University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 500+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation.
🔗 Official code & baseline: layumi/University1652-Baseline
· Leaderboard: State-of-the-art results
🔐 Access
This dataset is gated: click "Request… See the full description on the dataset page: https://huggingface.co/datasets/layumi/university-1652.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.Universe-Daily-Price
Universe Daily Price
The index of every symbol carried by the public daily price datasets, with the repository that holds it.
21,693 rows over 20,895 symbols, 4 columns. Updated by Papers With Backtest.
Why It Matters
This is the lookup table the other price datasets need:
Routing: A symbol on its own does not say which file holds it. repo_id answers that in one join, so a strategy that mixes equities, futures and rates loads from the right place without… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Universe-Daily-Price.cqadupstack-unix
CQADupstackUnixRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Written, Web, Programming
Reference
http://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackUnixRetrieval"])
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-unix.UnsolvedMath🌐 Browse UnsolvedMath online
✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark
UnsolvedMath Dataset
A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com.
Paper: "Open Mathematical Problems as an AI Reasoning Benchmark"
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.universal_dependencies
Dataset Card (v2.0) for Universal Dependencies Treebank
Version 2.0.0 introduces significant improvements and breaking changes:
Parquet Format: faster loading with HuggingFace datasets >=4.0.0
MWT Support: New mwt field provides structured multi-word token information
Enhanced Security: No more trust_remote_code=True required
Separate Versioning: Loader version (2.0.0) distinct from UD data version (2.18)
Breaking Changes:
Token sequences now exclude MWT surface forms… See the full description on the dataset page: https://huggingface.co/datasets/universal-dependencies/universal_dependencies.UnicEdit-10M
CVPR 2026 | UnicEdit-10M: Large-scale Image Editing Dataset
🔗 Quick Links
📄 Paper: UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
💻 Code: GitHub - WeChatCV/UnicBench
🌐 Project Page: UnicEdit-10M
🤗 Benchmark: UnicBench
🌟 Support Us: If you find this dataset or our work useful, please verify it by giving us a star on GitHub! Your support encourages us to keep open-sourcing… See the full description on the dataset page: https://huggingface.co/datasets/xiaotanhua/UnicEdit-10M.UniML3D
UniML3D
UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.LLaMmlein-DatasetThis dataset is a strict subset of the RedPajama V2 dataset and therefore retains all licenses from RedPajama V2.
More details in our preprint!
Data Take Down
ri-conicet
RI CONICET — metadatos completos + PDFs de acceso abierto
Snapshot del repositorio institucional CONICET Digital
(ri.conicet.gov.ar, DSpace, 266.824 publicaciones):
metadatos de toda la producción científica y tecnológica de CONICET,
junto con los PDFs de acceso abierto.
⚠️ Dataset en progreso — La recolección está en curso: el snapshot de
metadatos y el corpus de PDFs crecen con el tiempo y se actualizan
automáticamente. Si necesitás el repositorio completo de una vez… See the full description on the dataset page: https://huggingface.co/datasets/UNRN/ri-conicet.Unified-FeedbackCollections of pairwise feedback datasets.
openai/summarize_from_feedback
openai/webgpt_comparisons
Dahoas/instruct-synthetic-prompt-responses
Anthropic/hh-rlhf
lmsys/chatbot_arena_conversations
openbmb/UltraFeedback
argilla/ultrafeedback-binarized-preferences-cleaned
berkeley-nest/Nectar
Codes to reproduce the dataset: jdf-prog/UnifiedFeedback
Dataset formats
{
"id": "...",
"conv_A": [
{
"role": "user",
"content": "...",
},
{
"role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.browsecompMusic-AVQAgta-data-files-universalunsplash-lite
The Unsplash Lite Dataset (v1.2.1)
The Lite dataset contains all of the same fields as the Full dataset, but is limited to ~25,000 photos.
It can be used for both commercial and non-commercial usage, provided you abide by the terms.
The Unsplash Dataset is made available for research purposes.
It cannot be used to redistribute the images contained within.
To use the Unsplash library in a product, see the Unsplash API.
dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models
Training
Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data!
We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model,
the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.OmniScience
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
🚀 2026-05-17: This work was accepted by the KDD 2026 Dataset & Benchmark Track. 🚀 2026-05-01: The OmniScience dataset surpassed 20,000 average monthly downloads. 🚀 2026-01-21: The OmniScience dataset ranked Top 8 on Hugging Face Datasets Trending (Top 1 on Image Caption Filed). 🚀 2026-01-17: The OmniScience dataset surpassed 5,000 downloads within 5 days of its release. 🚀 2026-01-12:… See the full description on the dataset page: https://huggingface.co/datasets/UniParser/OmniScience.Sea-Undistort
Dataset Card for Sea-Undistort
Sea-Undistort is a synthetic dataset for through-water image restoration in high-resolution airborne bathymetry. It contains 1,200 scenes with four 512×512 RGB images per scene: (1) ground/no water, (2) undistorted/no waves, (3) no sunglint, (4) distorted (all effects). Each scene comes with structured per-image metadata describing camera, water, sky/illumination, and seafloor parameters. Images were procedurally rendered in Blender to emulate… See the full description on the dataset page: https://huggingface.co/datasets/maxkromer/Sea-Undistort.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.UniDoc-Bench
UNIDOC-BENCH Dataset
A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG).
Dataset Description
UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.UnlearnCanvas
Dataset Card for UnlearnCanvas
This dataset card introduces "UnlearnCanvas", a high-resolution stylized image dataset for benchmarking generative modeling tasks, in particular for machine unlearning in diffusion models. Developed to address the societal concerns arising from diffusion models, such as harmful content generation, copyright disputes, and the perpetuation of stereotypes and biases, UnlearnCanvas aims at facilitating the evaluation and improvement of machine unlearning… See the full description on the dataset page: https://huggingface.co/datasets/OPTML-Group/UnlearnCanvas.oas-unpaired
OAS Unpaired
The OAS unpaired dataset Observed Antibody Space (OAS), available as parquet with content-defined chunking on HuggingFace.
Configs and Splits
This dataset exposes 91 configs:
Config
Splits
Description
default
heavy, light
All sequences, split by chain
heavy
train
All heavy chain sequences
light
train
All light chain sequences
{Author et al., YYYY}
heavy, light, or both
One author's sequences
from datasets import load_dataset
# All heavy… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/oas-unpaired.
