datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DSBench
DSBench
This is the data file for the DSBench benchmark.
Paper or resources for more information:
https://arxiv.org/abs/2409.07703
dsb_audio_corpus
Acknowledgements
Thanks to all speakers that contributed to this dataset!
Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings!
MetaPKLot-Dataset
MetaPKLot
A Large-Scale Benchmark for Vision-Based Parking Lot Management
2,265,974 labeled samples · 1,366,185 new annotations · 3 research challenges · COCO-style annotations
MetaPKLot is a large-scale, harmonized dataset designed for research on vision-based parking lot management.
It extends and standardizes three existing parking datasets:
PKLot
CNRPark-EXT
PLds
MetaPKLot introduces new annotations, revises existing parking-space annotations, standardizes… See the full description on the dataset page: https://huggingface.co/datasets/DSBD-Research/MetaPKLot-Dataset.DSBenchDSB-IFEval
DuplexSpeechBench–IFEval (DSB-IFEval)
Evaluating implicit instruction following in full-duplex voice agents.
⚠️ Preprint — under review. Please cite it as a preprint (see below).
DSB-IFEval tests whether a real-time voice agent can infer the turn-taking behavior a role
implies — and execute it at the right moment on the conversational floor. It contains 1,038
evaluation cases built from 240 fixed user-side spoken interactions (8 assistant roles × 6
conversational probes × 5… See the full description on the dataset page: https://huggingface.co/datasets/puneetUMD/DSB-IFEval.ds-b9fd6dad5a782cb8
Audio data collection
Audio data distributed as TAR archives. Download access requires manual approval by the repository owner.
Published file paths use opaque identifiers. Archive contents retain their original structure.
This collection contains 1315 source files totaling 963708815360 bytes. All files have been uploaded and checked against source checksums and destination content hashes.
dsb_parquetds_benchmark_edicom_edicom_4ds_by_sys_prompt_15
Dataset Card for "ds_by_sys_prompt_15"
More Information needed
dsb_biases
Disability Accessibility & Bias Q&A Dataset
Dataset Description
This dataset is a collection of prompt-completion pairs focused on providing information about disability accessibility and addressing common biases and harmful language surrounding disability. It aims to serve as a resource for training language models to generate accurate, respectful, and inclusive responses in the context of disability.
The prompts cover a range of topics, including:
Accessibility… See the full description on the dataset page: https://huggingface.co/datasets/omark807/dsb_biases.ds-binarized_serbiands_benchmark_upv_lastexamds-be75b4e409e549321871
Sequence Object Navigation R5
This is a non-commercial research dataset derived from
SpatialVID-HQ. The
repository identifier is deliberately content-neutral, while this card documents
the contents, provenance, filtering policy, and license explicitly.
It contains 13,647 complete egocentric RGB videos and
18,249 object-goal navigation windows. Every instruction has
the exact form Go to <object>. and the selected object is intended to be visible
in the first frame. Videos are… See the full description on the dataset page: https://huggingface.co/datasets/Yangyihui/ds-be75b4e409e549321871.DSBC-DataFilesDS_Building_SecurityManual_V2dsbt-cleared-corpus-v1.1
DSBT cleared corpus v1.1
Train-only merge of reserve (19000) onto v1 PASS pack.
Split
Rows
Notes
train.jsonl
180023
v1 + reserve hard contrastive / high-K
eval_frozen_v1.jsonl
161023
exact v1 — use for Jev holdout / K-strata
Do not train-eval leak: never put reserve into the frozen eval.
CaseHOLD train sampling cap ≤15% remains mandatory.
ds_benchmark_upv_climactDSBC-Queries
UPDATED Version
Huggingface Dataset: datasets/large-traversaal/DSBC-Queries-V2.0
Github repo for evaluation:DSBC-Data-Science-Task-Evaluation
Dataset Details
we introduce a comprehensive benchmark of 400 queries specifically crafted to reflect real-world user interactions with data science agents by observing usage of our commercial applications.
Dataset Sources [optional]
Paper [optional]: arxiv.org/abs/2507.23336
Demo [optional]: ds.traversaal.ai… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/DSBC-Queries.brainblast-verified-footgun-corpus
Brainblast — Verified SDK Footgun Corpus (free sample)
The only code-training data that ships with a machine-checkable proof. Each
record is a real insecure→fixed code footgun with a replayable RED→GREEN
receipt: a deterministic checker fails the insecure version and passes the fixed
one. You don't trust the labels — you replay the proof.
This repo is a free 40-record sample (receipt-only tier). The full corpus is
4,183 proven records across 154 SDKs and 9 vulnerability classes… See the full description on the dataset page: https://huggingface.co/datasets/dsb117/brainblast-verified-footgun-corpus.ds_benchmark_upv_rescailingds_benchmark_upv_housepriceDS_bench
DS-bench: Code Generation Benchmark for Data Science Code
GitHub repo
Abstract
We introduce DS-bench, a new benchmark designed to evaluate large language models (LLMs) on complicated data science code generation tasks.
Existing benchmarks, such as DS-1000, often consist of overly simple code snippets, imprecise problem descriptions, and inadequate testing.
DS-bench sources 1,000 realistic problems from GitHub across ten widely used Python data science libraries, offering… See the full description on the dataset page: https://huggingface.co/datasets/LaPluma077/DS_bench.ds_benchmark_upv_stockmarketds_benchmark_edicom_edicom_3ds_benchmark_ct_mdsDSBio
DSBio: Scientific Analysis Tasks
DSBio is a suite of 90 expert-derived bioinformatics tasks constructed from
peer-reviewed academic publications and public scientific datasets.
These tasks are designed to evaluate whether agents can perform
domain-grounded scientific analysis, including:
Interpreting high-dimensional biological data (e.g., single-cell and spatial omics)
Understanding domain-specific terminology and conventions
Executing multi-step analytical workflows with… See the full description on the dataset page: https://huggingface.co/datasets/DSGym/DSBio.ds_benchmark_gds_airqualityafrica-ilo-une-tune-sex-mts-dsb-nb-unemployment-by-sex-marital-status-and-disability
Unemployment by sex, marital status and disability status (thousands) | Africa (ILOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ilo-une-tune-sex-mts-dsb-nb-unemployment-by-sex-marital-status-and-disability.ds_benchmark_gds_retailasia-ilo-eap-teap-sex-dsb-nb-labour-force-by-sex-and-disability-status-thousand
Labour force by sex and disability status (thousands) | Asia (ILOSTAT)
🌏 1,175 observations · 20 Asia countries · 1996–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 1,175 observations of Labour force data across 20 Asia countries, spanning 1996–2024, covering 1 distinct indicators.
About the source
ILOSTAT is the ILO's central statistics database, the leading global source for labour statistics. It compiles indicators across… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-ilo-eap-teap-sex-dsb-nb-labour-force-by-sex-and-disability-status-thousand.
