datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dsb_biases
Disability Accessibility & Bias Q&A Dataset
Dataset Description
This dataset is a collection of prompt-completion pairs focused on providing information about disability accessibility and addressing common biases and harmful language surrounding disability. It aims to serve as a resource for training language models to generate accurate, respectful, and inclusive responses in the context of disability.
The prompts cover a range of topics, including:
Accessibility… See the full description on the dataset page: https://huggingface.co/datasets/omark807/dsb_biases.DSBC-Queries
UPDATED Version
Huggingface Dataset: datasets/large-traversaal/DSBC-Queries-V2.0
Github repo for evaluation:DSBC-Data-Science-Task-Evaluation
Dataset Details
we introduce a comprehensive benchmark of 400 queries specifically crafted to reflect real-world user interactions with data science agents by observing usage of our commercial applications.
Dataset Sources [optional]
Paper [optional]: arxiv.org/abs/2507.23336
Demo [optional]: ds.traversaal.ai… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/DSBC-Queries.DSBio
DSBio: Scientific Analysis Tasks
DSBio is a suite of 90 expert-derived bioinformatics tasks constructed from
peer-reviewed academic publications and public scientific datasets.
These tasks are designed to evaluate whether agents can perform
domain-grounded scientific analysis, including:
Interpreting high-dimensional biological data (e.g., single-cell and spatial omics)
Understanding domain-specific terminology and conventions
Executing multi-step analytical workflows with… See the full description on the dataset page: https://huggingface.co/datasets/DSGym/DSBio.DSbench
