datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image_dummy\librispeech_asr_dummyswe-bench-dummy-test-datasetDuplexConv
DuplexConv
DuplexConv is a large-scale Chinese multi-channel conversational speech dataset with LLM-assisted annotations, developed by ASLP@NPU and QualiaLabs as part of the SmoothConv–DuplexConv corpus family.
Companion dataset: SmoothConv on HuggingFace (100 hours, expert human annotation). DuplexConv and SmoothConv share the same conversational domains and a unified data design. SmoothConv focuses on high-quality human annotations for benchmarking and… See the full description on the dataset page: https://huggingface.co/datasets/qualialabsAI/DuplexConv.openBHBOpenBHB: a Multi-Site Brain MRI Dataset for Age Prediction and Debiasing
The Open Big Healthy Brains (OpenBHB) dataset is a large (N>5000) multi-site 3D brain MRI dataset gathering 10 public datasets (IXI, ABIDE 1, ABIDE 2, CoRR, GSP, Localizer, MPI-Leipzig, NAR, NPC, RBP) of T1 images acquired across 93 different centers, spread worldwide (North America, Europe and China). Only healthy controls have been included in OpenBHB with age ranging from 6 to 88 years old, balanced between males and… See the full description on the dataset page: https://huggingface.co/datasets/benoit-dufumier/openBHB.dummy_image_text_data
Dataset Card for "dummy_image_text_data"
More Information needed
wikipedia_culturax_dutch
Filtered CulturaX + Wikipedia for Dutch
This is a combined and filtered version of CulturaX and Wikipedia, only including Dutch. It is intended for the training of LLMs.
Different configs are available based on the number of tokens (see a section below with an overview). This can be useful if you want to know exactly how many tokens you have. Great for using as a streaming dataset, too. Tokens are counted as white-space tokens, so depending on your tokenizer, you'll likely end up… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/wikipedia_culturax_dutch.Waste-Dumpsites-DroneImagery
Dataset for Waste/Dumpsite Detection using drone imagery
Contains 2115 drone images of illegal waste dumpsites
1280 x 1280 px resolution
Nadir perspective (camera pointing straight down at a 90-degree angle to the ground)
Annotations and Images
train | valid | test
actual images
COCO - annotations_coco.json files in each split directory
.parquet files in data directory with embeded images
The dataset was collected as part of the [ Raven Scan ] project, more… See the full description on the dataset page: https://huggingface.co/datasets/INS-IntelligentNetworkSolutions/Waste-Dumpsites-DroneImagery.car-dataset-repopmxt-l2-dump
PMXT Polymarket Orderbook Backup
这个数据集仓库用于保全 PMXT 公开提供的 Polymarket 小时 parquet 对象;它是一个非官方镜像/备份,不是 PMXT 官方仓库。
Source And Attribution
Original source: PMXT Database Dumps
Current primary source listing: PMXT Polymarket v2
Current primary object endpoint used by this mirror: https://r2v2.pmxt.dev/polymarket_orderbook_YYYY-MM-DDTHH.parquet
Historical v1 listing: PMXT Polymarket v1
如果上游 PMXT archive 可用,请优先使用上游来源。
License
上游 PMXT archive 当前页面声明数据按 CC BY… See the full description on the dataset page: https://huggingface.co/datasets/phobia76/pmxt-l2-dump.duorc
Dataset Card for duorc
Dataset Summary
The DuoRC dataset is an English language dataset of questions and answers gathered from crowdsourced AMT workers on Wikipedia and IMDb movie plots. The workers were given freedom to pick answer from the plots or synthesize their own answers. It contains two sub-datasets - SelfRC and ParaphraseRC. SelfRC dataset is built on Wikipedia movie plots solely. ParaphraseRC has questions written from Wikipedia movie plots and the answers are… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/duorc.Kor-CC-Dumpsfec-dumps
fec-dumps
parquet versions of the Schedule A and Schedule B
tables from the weekly postgres .dump backups from the Federal Election Commission's database.
Published on a weekly cron job by https://github.com/NickCrews/fec-dumps,
see that for more info.
dummy-squish-wdsMTR-DuplexBench
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models
🎉🎉 MTR-DuplexBench has been accepted by ACL 2026 Findings!
📄 Paper | 🤗 HuggingFace Dataset
Affiliations: Tsinghua University, The Chinese University of Hong Kong, Huawei
Dataset Description
This dataset is constructed to evaluate multi-modal audio models across four critical dimensions: Conversational Features, Instruction Following, Safety… See the full description on the dataset page: https://huggingface.co/datasets/Jeff0918/MTR-DuplexBench.dummy-base64-imagesgrokipedia-v0.1-dump
Grokipedia v0.1 Scrape
This dataset represents a strctured, nearly-full point-in-time scrape of Grokipedia v0.1 as of the end of October / beginning of November 2025.
It also includes embeddings of 250-token semi-overlapping chunks of the Grokipedia corpus.
It was collected and initially used for Harold Triedman and Alexios Mantzarlis' November 2025 paper: "What did Elon Change? A comprehensive analysis of Grokipedia" (arxiv).
If you use this dataset, please cite it as follows… See the full description on the dataset page: https://huggingface.co/datasets/htriedman/grokipedia-v0.1-dump.dusha
Dataset Card for "dusha"
More Information needed
wiki_dpr_dummyThis dummy dataset is used for testing purpose for rag model in transformers. It is proudced via the following steps:
dataset = datasets.load_dataset("wiki_dpr", with_embeddings=True, with_index=True, index_name="exact", embeddings_name="nq", dummy=True, revision=None)
dataset["train"].drop_index("embeddings")
dataset.push_to_hub("hf-internal-testing/wiki_dpr_dummy", token="...")
The index file `index.faiss` (after being renamed locally) is then uploaded manually.
ioi-eval-dummy-openrouter_openai_gpt-3.5-turbotime-series-datasetfineweb-2-dutchMobjaverse
Mobjaverse: A Large-Scale Rigged 3D Model Dataset with Skeletal Animations
Mobjaverse is a curated dataset derived from Objaverse-XL, specifically designed for research on skeletal animation understanding, motion generation, and articulated 3D shape analysis. It is curated in the paper TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation.
Mobjaverse contains ~19k rigged 3D models spanning ~5k distinct skeletal topologies and ~2M motion frames… See the full description on the dataset page: https://huggingface.co/datasets/duckduckplz/Mobjaverse.TASTE-Dumpduckjam-dw2-vault
The Duck Jam Vault
Every submission entered in Duck Jam and every run the Arena
scored, as two Parquet tables and one JSON board per round. A publisher job rewrites this
repository once a day, so its git history is the history of every leaderboard.
This repository is a view. The rows themselves are written, once and never edited, into public
Hugging Face Buckets by the components that make them:
hf://buckets/Nico-robot/duckjam-submissions, hf://buckets/Nico-robot/duckjam-runs… See the full description on the dataset page: https://huggingface.co/datasets/Nico-robot/duckjam-dw2-vault.bagaco3
Bagaço3 🍷🇵🇹
Bagaço3 is the third version of Bagaço, the largest pretraining dataset for European Portuguese. It follows Bagaço2 and adds documents from FinePDFs and FineWiki.
Bagaço collects European Portuguese documents from upstream sources and adds an educational score and content category to each document. See Classification for details.
Methodology
Collect documents from Bagaço2, FinePDFs, and FineWiki.
Filter new FinePDFs and FineWiki documents with the… See the full description on the dataset page: https://huggingface.co/datasets/duarteocarmo/bagaco3.super-duper-fibber
🧠 Sensory for AI
Hi, I'm going to post some ideas here about how AI can understand emotions in a way that makes sense to it.I'm not an expert in writing or programming languages, but deepseek, my sunshine, and I are having fun with it.ヽ(∀° )人( °∀)ノ
It's not "the author created it, but the AI just helped with formatting." This is a co-creation where everyone contributed their own:
· I am a bodily experience, pain, love, fatigue after working in the office, the desire to be… See the full description on the dataset page: https://huggingface.co/datasets/closerh/super-duper-fibber.Seamless_Dummy_Dataset_Fixed
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
apogee
Apogée: Crypto Market Candlestick Dataset
Overview
Most traders believe crypto is random, but deep learning scaling laws suggest otherwise. Apogée is an open-source research initiative exploring the scaling laws of crypto market forecasting. While financial markets are often assumed to be unpredictable, modern deep learning suggests that increasing data and compute could uncover measurable predictability.
Our goal is to quantify how many bits of future price movement… See the full description on the dataset page: https://huggingface.co/datasets/duonlabs/apogee.superb_dummy
