CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01monology /pile-uncopyrighted Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA. MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.text100M<n<1B175 likes80k downloads3y agoHugging Face02unimelb-nlp /wikiann Dataset Card for WikiANN Dataset Summary WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, and test splits of Rahimi et al. (2019), which supports 176 of the 282 languages from the original WikiANN corpus. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/unimelb-nlp/wikiann.texttoken-classification1M<n<10M124 likes30k downloads3y agoHugging Face03h8st6ptv /turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format: All Universities in Turkey Dataset Description This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities. Fields 1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.imagen<1K2 likes22k downloads2y agoHugging Face04UnipatAI /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.imagetext-generationn<1K2 likes14k downloads5mo agoHugging Face05agents-course /unit4-students-scorestext10K<n<100K20 likes14k downloads2h agoHugging Face06Aeala /ShareGPT_Vicuna_unfiltered Dataset Card This is a reupload of this dataset that was further cleaned by gozfarb. text100K<n<1M52 likes9.7k downloads3y agoHugging Face07Helsinki-NLP /un_pc Dataset Card for United Nations Parallel Corpus Dataset Summary The United Nations Parallel Corpus is the first parallel corpus composed from United Nations documents published by the original data creator. The parallel corpus consists of manually translated UN documents from the last 25 years (1990 to 2014) for the six official UN languages, Arabic, Chinese, English, French, Russian, and Spanish. The corpus is freely available for download under a liberal license.… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/un_pc.texttranslation100M<n<1B28 likes9.6k downloads2y agoHugging Face08layumi /university-1652gated University-1652: Drone-based Geo-localization Benchmark 🚁 University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 500+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation. 🔗 Official code & baseline: layumi/University1652-Baseline · Leaderboard: State-of-the-art results 🔐 Access This dataset is gated: click "Request… See the full description on the dataset page: https://huggingface.co/datasets/layumi/university-1652.imageimage-feature-extraction100K<n<1M8 likes9.6k downloads3mo agoHugging Face09Samuelsantos777 /psg-audio-v3-unofficial-mirror PSG-Audio v3 — Unofficial Complete Mirror Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset. This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community. Overview PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.textaudio-classificationn<1K1 likes8.2k downloads3mo agoHugging Face10paperswithbacktest /Universe-Daily-Pricegated Universe Daily Price The index of every symbol carried by the public daily price datasets, with the repository that holds it. 21,693 rows over 20,895 symbols, 4 columns. Updated by Papers With Backtest. Why It Matters This is the lookup table the other price datasets need: Routing: A symbol on its own does not say which file holds it. repo_id answers that in one join, so a strategy that mixes equities, futures and rates loads from the right place without… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Universe-Daily-Price.tabulartime-series-forecasting10K<n<100K0 likes6.5k downloads24d agoHugging Face11mteb /cqadupstack-unix CQADupstackUnixRetrieval An MTEB dataset Massive Text Embedding Benchmark CQADupStack: A Benchmark Data Set for Community Question-Answering Research Task category t2t Domains Written, Web, Programming Reference http://nlp.cis.unimelb.edu.au/resources/cqadupstack/ How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["CQADupstackUnixRetrieval"]) evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-unix.texttext-retrieval10K<n<100K0 likes6.3k downloads1y agoHugging Face12ulamai /UnsolvedMath🌐 Browse UnsolvedMath online ✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark UnsolvedMath Dataset A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com. Paper: "Open Mathematical Problems as an AI Reasoning Benchmark" Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.documentquestion-answering10K<n<100K80 likes5.8k downloads9d agoHugging Face13universal-dependencies /universal_dependencies Dataset Card (v2.0) for Universal Dependencies Treebank Version 2.0.0 introduces significant improvements and breaking changes: Parquet Format: faster loading with HuggingFace datasets >=4.0.0 MWT Support: New mwt field provides structured multi-word token information Enhanced Security: No more trust_remote_code=True required Separate Versioning: Loader version (2.0.0) distinct from UD data version (2.18) Breaking Changes: Token sequences now exclude MWT surface forms… See the full description on the dataset page: https://huggingface.co/datasets/universal-dependencies/universal_dependencies.texttoken-classification1M<n<10M8 likes5.8k downloads7d agoHugging Face14xiaotanhua /UnicEdit-10M CVPR 2026 | UnicEdit-10M: Large-scale Image Editing Dataset 🔗 Quick Links 📄 Paper: UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits 💻 Code: GitHub - WeChatCV/UnicBench 🌐 Project Page: UnicEdit-10M 🤗 Benchmark: UnicBench 🌟 Support Us: If you find this dataset or our work useful, please verify it by giving us a star on GitHub! Your support encourages us to keep open-sourcing… See the full description on the dataset page: https://huggingface.co/datasets/xiaotanhua/UnicEdit-10M.imageimage-to-image1M<n<10M10 likes5.7k downloads6mo agoHugging Face15Linzhan /UniML3D UniML3D UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.imagetext-to-3d10K<n<100K8 likes4.9k downloads2d agoHugging Face16LSX-UniWue /LLaMmlein-DatasetThis dataset is a strict subset of the RedPajama V2 dataset and therefore retains all licenses from RedPajama V2. More details in our preprint! Data Take Down texttext-generation100M<n<1B5 likes4.6k downloads11mo agoHugging Face17UNRN /ri-conicet RI CONICET — metadatos completos + PDFs de acceso abierto Snapshot del repositorio institucional CONICET Digital (ri.conicet.gov.ar, DSpace, 266.824 publicaciones): metadatos de toda la producción científica y tecnológica de CONICET, junto con los PDFs de acceso abierto. ⚠️ Dataset en progreso — La recolección está en curso: el snapshot de metadatos y el corpus de PDFs crecen con el tiempo y se actualizan automáticamente. Si necesitás el repositorio completo de una vez… See the full description on the dataset page: https://huggingface.co/datasets/UNRN/ri-conicet.document100K<n<1M1 likes4.5k downloads13d agoHugging Face18llm-blender /Unified-FeedbackCollections of pairwise feedback datasets. openai/summarize_from_feedback openai/webgpt_comparisons Dahoas/instruct-synthetic-prompt-responses Anthropic/hh-rlhf lmsys/chatbot_arena_conversations openbmb/UltraFeedback argilla/ultrafeedback-binarized-preferences-cleaned berkeley-nest/Nectar Codes to reproduce the dataset: jdf-prog/UnifiedFeedback Dataset formats { "id": "...", "conv_A": [ { "role": "user", "content": "...", }, { "role": "assistant"… See the full description on the dataset page: https://huggingface.co/datasets/llm-blender/Unified-Feedback.tabular1M<n<10M18 likes4.5k downloads2y agoHugging Face19Chat-UniVi /browsecomptextn<1K0 likes4.2k downloads1y agoHugging Face20UnFaZeD07 /Music-AVQAtabular10K<n<100K0 likes4.2k downloads7mo agoHugging Face21dfdffdfdfdfd /gta-data-files-universalgeospatial3 likes4.1k downloads3mo agoHugging Face221aurent /unsplash-lite The Unsplash Lite Dataset (v1.2.1) The Lite dataset contains all of the same fields as the Full dataset, but is limited to ~25,000 photos. It can be used for both commercial and non-commercial usage, provided you abide by the terms. The Unsplash Dataset is made available for research purposes. It cannot be used to redistribute the images contained within. To use the Unsplash library in a product, see the Unsplash API. texttext-to-image10K<n<100K10 likes3.9k downloads3y agoHugging Face23lesc-unifi /dragon Dataset Card for DRAGON 🧾 ArXiv Preprint DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models. The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures. Dataset Details Dataset Description The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.imageimage-classification1M<n<10M9 likes3.9k downloads1y agoHugging Face24nomic-ai /nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models Training Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data! We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model, the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.text100M<n<1B21 likes3.8k downloads2y agoHugging Face25UniParser /OmniSciencegated OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding 🚀 2026-05-17: This work was accepted by the KDD 2026 Dataset & Benchmark Track. 🚀 2026-05-01: The OmniScience dataset surpassed 20,000 average monthly downloads. 🚀 2026-01-21: The OmniScience dataset ranked Top 8 on Hugging Face Datasets Trending (Top 1 on Image Caption Filed). 🚀 2026-01-17: The OmniScience dataset surpassed 5,000 downloads within 5 days of its release. 🚀 2026-01-12:… See the full description on the dataset page: https://huggingface.co/datasets/UniParser/OmniScience.imageimage-to-text1M<n<10M129 likes3.7k downloads26d agoHugging Face26maxkromer /Sea-Undistort Dataset Card for Sea-Undistort Sea-Undistort is a synthetic dataset for through-water image restoration in high-resolution airborne bathymetry. It contains 1,200 scenes with four 512×512 RGB images per scene: (1) ground/no water, (2) undistorted/no waves, (3) no sunglint, (4) distorted (all effects). Each scene comes with structured per-image metadata describing camera, water, sky/illumination, and seafloor parameters. Images were procedurally rendered in Blender to emulate… See the full description on the dataset page: https://huggingface.co/datasets/maxkromer/Sea-Undistort.imageimage-to-image1K<n<10K1 likes3.7k downloads11mo agoHugging Face27unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face28Salesforce /UniDoc-Bench UNIDOC-BENCH Dataset A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG). Dataset Description UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.imagequestion-answering1K<n<10K15 likes3.3k downloads10mo agoHugging Face29OPTML-Group /UnlearnCanvas Dataset Card for UnlearnCanvas This dataset card introduces "UnlearnCanvas", a high-resolution stylized image dataset for benchmarking generative modeling tasks, in particular for machine unlearning in diffusion models. Developed to address the societal concerns arising from diffusion models, such as harmful content generation, copyright disputes, and the perpetuation of stereotypes and biases, UnlearnCanvas aims at facilitating the evaluation and improvement of machine unlearning… See the full description on the dataset page: https://huggingface.co/datasets/OPTML-Group/UnlearnCanvas.image1K<n<10K2 likes3.3k downloads3y agoHugging Face30ConvergeBio /oas-unpaired OAS Unpaired The OAS unpaired dataset Observed Antibody Space (OAS), available as parquet with content-defined chunking on HuggingFace. Configs and Splits This dataset exposes 91 configs: Config Splits Description default heavy, light All sequences, split by chain heavy train All heavy chain sequences light train All light chain sequences {Author et al., YYYY} heavy, light, or both One author's sequences from datasets import load_dataset # All heavy… See the full description on the dataset page: https://huggingface.co/datasets/ConvergeBio/oas-unpaired.tabular1B<n<10B2 likes3.2k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.