datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ServiceProjectFall2023
Deep Learning Service Project (Fall 2023)
Getting Started
Clone the repository with git lfs disabled or not installed.
ON WINDOWS
set GIT_LFS_SKIP_SMUDGE=1
git clone https://huggingface.co/datasets/DataScienceClubUVU/ServiceProjectFall2023
ON LINUX
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/DataScienceClubUVU/ServiceProjectFall2023
Download the pytorch file (.pth) from… See the full description on the dataset page: https://huggingface.co/datasets/DataScienceClubUVU/ServiceProjectFall2023.datapoints_round1_dpsk_data_science_shard1_daytona_n100k1swesmith-datascience-skorch-sandboxesData_Science-21voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.DiEm_HTR
Dataset Card for DiEm HTR
The DiEm HTR dataset is a ground truth dataset for historical danish handwriting in the 17th and 18th century, generated as part of the Digitalisering af Enesteministerialbøger project at the Danish National Archives.
Dataset Details
Dataset Description
The Digitalisering af Enesteministerialbøger project (DiEm) at the Danish National Archives aims to transcribe and make publically available all of the danish parish registers from… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/DiEm_HTR.data-science-en-id
Data Science EN-ID Parallel Corpus (Scientific Domain)
Dataset Description
This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language.
Primary Languages: English (EN) and Indonesian (ID)
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.nemotron-terminal-data_science
nemotron-terminal-data_science
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_science". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_science.terminal_bench_2_nemotron_terminal_data_science__Qwen3_8B_20260414_004932modern-danish-handwriting
Dataset Card for Modern Danish Handwriting
The Modern Danish Handwriting dataset is a Danish-language dataset containing more than 200 pages of transcribed and proofread handwritten text.
Dataset Details
Dataset Description
The Modern Danish Handwriting dataset currently consists of handwritten samples of text from the ePAROLE dataset. The samples were created by volunteers at the Danish National Archives and guests at the festival Historiske Dage in 2025.… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/modern-danish-handwriting.Data_Science-21medium_20000-data_science_n100k1easy_5000-data_science_n100k1medium_5000-data_science_n100k1DPO-datasciencedatascience-stackexchange-posts
Dataset Card for "datascience-stackexchange-posts"
More Information needed
Data_Science-21medium_5000_data_science_fixedscienceqa-cot-datamedium_5000_data_scienceDiEm_HTR-Numbers
Dataset Card for DiEm HTR Numbers
The DiEm HTR Numbers dataset is a ground truth dataset consisting of numbers written in historical danish handwriting from the 18th century, generated as part of the Digitalisering af Enesteministerialbøger project at the Danish National Archives.
Dataset Details
Dataset Description
The Digitalisering af Enesteministerialbøger project (DiEm) at the Danish National Archives aims to transcribe and make publically available… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/DiEm_HTR-Numbers.easy_5000_data_scienceworkplace-dynamics-survey
Workplace Dynamics Survey
Anonymized workplace survey data for organizational research.
Usage
from datasets import load_dataset
dataset = load_dataset("org-science-data/workplace-dynamics-survey")
df = dataset["train"].to_pandas()
Or use the provided loader:
from loader import load_data
df = load_data()
Schema
Metrics
Column
Type
Description
role_diversity
float
Normalized metric
goal_alignment
float
Normalized metric… See the full description on the dataset page: https://huggingface.co/datasets/org-science-data/workplace-dynamics-survey.router_SFT_larger_model_generated_data_mmlu_pro_science_Qwen3-4B_aimerouter_PEFT_data_mmlu_pro_science_5_shot_shuffle_Meta-Llama-3-8B-Instructmixed_1000_data_sciencerouter_PEFT_data_Science_v2_Qwen3-8B_70_30scienceqadatascience-instruct
Dataset Card for Dataset Name
The datascience instruct dataset is a collection of question answers based around various topics of datascience.
Dataset Description
The primary goal of this dataset is to fine tune base LLMs for responding to data science queries. According to our observation, most base LLMs (2B to 7B) are good in understanding data science concepts but they lack in responding step by step. This dataset contains well structured user agent interaction… See the full description on the dataset page: https://huggingface.co/datasets/hanzla/datascience-instruct.easy_5000_data_science_fixed
