datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
generalization-science-dataDataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/GG-samrt/DataScience-Instruct-500K.PopMCQ
🎯 PopMCQ
Does your model pick the famous answer or the correct one?
📌 Overview
PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty.
The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.modis-lake-powell-toy-dataset
MODIS Water Lake Powell Toy Dataset
Dataset Summary
Tabular dataset comprised of MODIS surface reflectance bands along with calculated indices and a label (water/not-water)
Dataset Structure
Data Fields
water: Label, water or not-water (binary)
sur_refl_b01_1: MODIS surface reflection band 1 (-100, 16000)
sur_refl_b02_1: MODIS surface reflection band 2 (-100, 16000)
sur_refl_b03_1: MODIS surface reflection band 3 (-100, 16000)
sur_refl_b04_1: MODIS… See the full description on the dataset page: https://huggingface.co/datasets/nasa-cisto-data-science-group/modis-lake-powell-toy-dataset.Data_Science-21DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/fantos/DataScience-Instruct-500K.data-science-job-salaries
Dataset Card for Data Science Job Salaries
Dataset Summary
Content
Column
Description
work_year
The year the salary was paid.
experience_level
The experience level in the job during the year with the following possible values: EN Entry-level / Junior MI Mid-level / Intermediate SE Senior-level / Expert EX Executive-level / Director
employment_type
The type of employement for the role: PT Part-time FT Full-time CT Contract FL Freelance
job_title… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/data-science-job-salaries.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/Fan0718/DataScience-Instruct-500K.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/binzhango/DataScience-Instruct-500K.SO-Python_QA-Data_Science_and_Machine_Learning_classData_Science-21datascience-bowl2019https://www.kaggle.com/c/data-science-bowl-2019
Data_Science-21workplace-dynamics-survey
Workplace Dynamics Survey
Anonymized workplace survey data for organizational research.
Usage
from datasets import load_dataset
dataset = load_dataset("org-science-data/workplace-dynamics-survey")
df = dataset["train"].to_pandas()
Or use the provided loader:
from loader import load_data
df = load_data()
Schema
Metrics
Column
Type
Description
role_diversity
float
Normalized metric
goal_alignment
float
Normalized metric… See the full description on the dataset page: https://huggingface.co/datasets/org-science-data/workplace-dynamics-survey.router_SFT_larger_model_generated_data_mmlu_pro_science_Qwen3-4B_aimerouter_PEFT_data_mmlu_pro_science_5_shot_shuffle_Meta-Llama-3-8B-Instructdata-science-job-salariesrouter_SFT_larger_model_generated_data_mmlu_pro_science_Qwen3-14BAI-and-Data-Science-Job-Market-Dataset
About Dataset
Edit
The AI & Data Science Job Market Dataset (2020–2026) is a synthetically generated dataset designed to simulate real-world hiring patterns across the artificial intelligence and data science job market.
The dataset contains structured information about job roles, company characteristics, required technical skills, education levels, experience requirements, and salary ranges. It reflects hiring data across multiple countries, industries, and company sizes.
This… See the full description on the dataset page: https://huggingface.co/datasets/shree0910/AI-and-Data-Science-Job-Market-Dataset.router_SFT_self_generated_data_mmlu_pro_science_OLMo-2-1124-13B-Instructrouter_PEFT_data_mmlu_pro_science_5_shot_shuffle_Qwen3-0.6B_rolloutrouter_SFT_larger_model_generated_data_mmlu_pro_science_OLMo-2-1124-13B-Instructrouter_SFT_self_generated_data_mmlu_pro_science_Meta-Llama-3-8B-Instructrouter_SFT_self_generated_data_mmlu_pro_science_Qwen3-4B_aimerouter_PEFT_data_mmlu_pro_science_5_shot_shuffle_Qwen3-1.7B_rolloutrouter_PEFT_data_Science_v2_Qwen3-14B_rolloutrouter_SFT_self_generated_data_mmlu_pro_science_Qwen3-0.6Brouter_SFT_self_generated_data_mmlu_pro_science_Qwen3-8Brouter_SFT_larger_model_generated_data_mmlu_pro_science_OLMo-2-1124-7B-Instruct
