datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
data-science-workflows-sft-100k
Data Science Workflows SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment.
Motivation
Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.DataScience-ML-DATASETS
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/eshmoideas/DataScience-ML-DATASETS.
