datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/GG-samrt/DataScience-Instruct-500K.PopMCQ
🎯 PopMCQ
Does your model pick the famous answer or the correct one?
📌 Overview
PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty.
The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.OBLIQ-IR-Data
OBLIQ-IR-Data
The training data behind DataScience-UIBK/OBLIQ-IR-3B,
plus the retrieval runs and evaluation outputs for every result in
OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026).
Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance,
an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has
little or no surface expression in the document.
🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/Fan0718/DataScience-Instruct-500K.DataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/fantos/DataScience-Instruct-500K.SO-Python_QA-Data_Science_and_Machine_Learning_classDataScience-Instruct-500K
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du
DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting:
🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation.
🔍… See the full description on the dataset page: https://huggingface.co/datasets/binzhango/DataScience-Instruct-500K.turkish-science-data
Turkish Science Data
Yüksek kaliteli Türkçe fen bilimleri sentetik veri seti.
İçerik
Alan
Konu Başlıkları
Fizik
Kinematik, Newton yasaları, Enerji/İş, Elektrik devreleri
Kimya
Mol hesabı, Asit-baz/pH, Kimyasal denklemler, Periyodik tablo
Biyoloji
Mendel genetiği, DNA, Mitoz/Mayoz, Organeller, Hormonlar
Format
Her örnek soru-cevap formatında Türkçe düz metin:
Soru: ...
Çözüm: ...
Cevap: ...
Kullanım
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/DevHunterAI/turkish-science-data.data-science-job-descriptions
Data Science Job DescriptionsThese data encompass the title, company, and description of the outer-join job board between October 2021 and today.
license: wtfpl
task_categories:
- text-classification
- feature-extraction
language:
- en
tags:
- jobs
pretty_name: ds-jobs
size_categories:
- 1K<n<10K
data-science-workflows-sft-100k
Data Science Workflows SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment.
Motivation
Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.Connor-Data-Sciences_humainesConnor-Data-SciencesConnor-Data-AEX-sciencesenglish-data-science-basics-30Final-ALevel-Science-DataConnor-Data-DatascienceConnor-Data-AEX-datascienceConnor-Data-AEX-sciences_humainesDataScience-ML-DATASETS
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/eshmoideas/DataScience-ML-DATASETS.intrain_data_sciencechat01imageimage-text-jsonimage-text2
