datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.generalization-science-dataServiceProjectFall2023
Deep Learning Service Project (Fall 2023)
Getting Started
Clone the repository with git lfs disabled or not installed.
ON WINDOWS
set GIT_LFS_SKIP_SMUDGE=1
git clone https://huggingface.co/datasets/DataScienceClubUVU/ServiceProjectFall2023
ON LINUX
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/DataScienceClubUVU/ServiceProjectFall2023
Download the pytorch file (.pth) from… See the full description on the dataset page: https://huggingface.co/datasets/DataScienceClubUVU/ServiceProjectFall2023.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/GG-samrt/DataScience-Instruct-500K.data-science-en-id
Data Science EN-ID Parallel Corpus (Scientific Domain)
Dataset Description
This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language.
Primary Languages: English (EN) and Indonesian (ID)
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.PopMCQ
🎯 PopMCQ
Does your model pick the famous answer or the correct one?
📌 Overview
PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty.
The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.datapoints_round1_dpsk_data_science_shard1_daytona_n100k1swesmith-datascience-skorch-sandboxesData_Science-21current-dataYour mother
nobody is going to see this probably
I saw
DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/fantos/DataScience-Instruct-500K.OBLIQ-IR-Data
OBLIQ-IR-Data
The training data behind DataScience-UIBK/OBLIQ-IR-3B,
plus the retrieval runs and evaluation outputs for every result in
OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026).
Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance,
an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has
little or no surface expression in the document.
🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.data-science-job-salaries
Dataset Card for Data Science Job Salaries
Dataset Summary
Content
Column
Description
work_year
The year the salary was paid.
experience_level
The experience level in the job during the year with the following possible values: EN Entry-level / Junior MI Mid-level / Intermediate SE Senior-level / Expert EX Executive-level / Director
employment_type
The type of employement for the role: PT Part-time FT Full-time CT Contract FL Freelance
job_title… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/data-science-job-salaries.voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.DiEm_HTR
Dataset Card for DiEm HTR
The DiEm HTR dataset is a ground truth dataset for historical danish handwriting in the 17th and 18th century, generated as part of the Digitalisering af Enesteministerialbøger project at the Danish National Archives.
Dataset Details
Dataset Description
The Digitalisering af Enesteministerialbøger project (DiEm) at the Danish National Archives aims to transcribe and make publically available all of the danish parish registers from… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/DiEm_HTR.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/Fan0718/DataScience-Instruct-500K.nemotron-terminal-data_science
nemotron-terminal-data_science
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_science". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_science.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/binzhango/DataScience-Instruct-500K.terminal_bench_2_nemotron_terminal_data_science__Qwen3_8B_20260414_004932turkish-science-data
Turkish Science Data
Yüksek kaliteli Türkçe fen bilimleri sentetik veri seti.
İçerik
Alan
Konu Başlıkları
Fizik
Kinematik, Newton yasaları, Enerji/İş, Elektrik devreleri
Kimya
Mol hesabı, Asit-baz/pH, Kimyasal denklemler, Periyodik tablo
Biyoloji
Mendel genetiği, DNA, Mitoz/Mayoz, Organeller, Hormonlar
Format
Her örnek soru-cevap formatında Türkçe düz metin:
Soru: ...
Çözüm: ...
Cevap: ...
Kullanım
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/DevHunterAI/turkish-science-data.science-journal-for-kids-data
Science Journal for Kids Data
This repository contains a dataset of abstracts from the Science Journal for Kids website and the original academic papers. It includes metadata such as titles, URLs, reading levels, and links to the full academic papers. The dataset is designed to support research and analysis of educational content tailored for young learners.
Data
The dataset is a curated collection of 284 original scientific abstracts and their adapted abstracts for… See the full description on the dataset page: https://huggingface.co/datasets/loukritia/science-journal-for-kids-data.SO-Python_QA-Data_Science_and_Machine_Learning_classData_Science-21modern-danish-handwriting
Dataset Card for Modern Danish Handwriting
The Modern Danish Handwriting dataset is a Danish-language dataset containing more than 200 pages of transcribed and proofread handwritten text.
Dataset Details
Dataset Description
The Modern Danish Handwriting dataset currently consists of handwritten samples of text from the ePAROLE dataset. The samples were created by volunteers at the Danish National Archives and guests at the festival Historiske Dage in 2025.… See the full description on the dataset page: https://huggingface.co/datasets/RA-Data-Science/modern-danish-handwriting.medium_20000-data_science_n100k1data-science-job-descriptions
Data Science Job DescriptionsThese data encompass the title, company, and description of the outer-join job board between October 2021 and today.
license: wtfpl
task_categories:
- text-classification
- feature-extraction
language:
- en
tags:
- jobs
pretty_name: ds-jobs
size_categories:
- 1K<n<10K
easy_5000-data_science_n100k1data-science-workflows-sft-100k
Data Science Workflows SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment.
Motivation
Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.datascience-bowl2019https://www.kaggle.com/c/data-science-bowl-2019
medium_5000-data_science_n100k1
