CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-DataLab /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/RUC-DataLab/DataScience-Instruct-500K.tabular10K<n<100K76 likes852 downloads11mo agoHugging Face02GG-samrt /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/GG-samrt/DataScience-Instruct-500K.tabular10K<n<100K0 likes287 downloads8mo agoHugging Face03DataScience-UIBK /PopMCQ 🎯 PopMCQ Does your model pick the famous answer or the correct one? 📌 Overview PopMCQ renders the same question six ways. The question and the correct answer never change — only how popular the three distractors are. That makes option popularity an independent variable, so an accuracy swing across S1–S6 is attributable to popularity rather than to question difficulty. The swings are large. Under the hardest setting (S2), models pick a popular-but-wrong… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/PopMCQ.tabularquestion-answering10K<n<100K0 likes158 downloads15d agoHugging Face04fantos /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/fantos/DataScience-Instruct-500K.tabular10K<n<100K0 likes97 downloads11mo agoHugging Face05DataScience-UIBK /OBLIQ-IR-Data OBLIQ-IR-Data The training data behind DataScience-UIBK/OBLIQ-IR-3B, plus the retrieval runs and evaluation outputs for every result in OBLIQ-IR: Training a Dense Retriever for Oblique Queries (EMNLP 2026). Oblique retrieval is the setting where relevance is decided by a latent attribute — an implicit stance, an abstract proof strategy, an authorial style, a lossy recollection of a rhetorical exchange — that has little or no surface expression in the document. 🤖 Model:… See the full description on the dataset page: https://huggingface.co/datasets/DataScience-UIBK/OBLIQ-IR-Data.texttext-retrieval100K<n<1M4 likes96 downloads1mo agoHugging Face06Fan0718 /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/Fan0718/DataScience-Instruct-500K.tabular10K<n<100K0 likes77 downloads8mo agoHugging Face07binzhango /DataScience-Instruct-500K DeepAnalyze: Agentic Large Language Models for Autonomous Data Science Authors: Shaolei Zhang, Ju Fan*, Meihao Fan, Guoliang Li, Xiaoyong Du DeepAnalyze is the first agentic LLM for autonomous data science. It can autonomously complete a wide range of data-centric tasks without human intervention, supporting: 🛠 Entire data science pipeline: Automatically perform any data science tasks such as data preparation, analysis, modeling, visualization, and report generation. 🔍… See the full description on the dataset page: https://huggingface.co/datasets/binzhango/DataScience-Instruct-500K.tabular10K<n<100K0 likes70 downloads10mo agoHugging Face08RazinAleks /SO-Python_QA-Data_Science_and_Machine_Learning_classtabular1K<n<10K6 likes55 downloads3y agoHugging Face09DevHunterAI /turkish-science-data Turkish Science Data Yüksek kaliteli Türkçe fen bilimleri sentetik veri seti. İçerik Alan Konu Başlıkları Fizik Kinematik, Newton yasaları, Enerji/İş, Elektrik devreleri Kimya Mol hesabı, Asit-baz/pH, Kimyasal denklemler, Periyodik tablo Biyoloji Mendel genetiği, DNA, Mitoz/Mayoz, Organeller, Hormonlar Format Her örnek soru-cevap formatında Türkçe düz metin: Soru: ... Çözüm: ... Cevap: ... Kullanım from datasets… See the full description on the dataset page: https://huggingface.co/datasets/DevHunterAI/turkish-science-data.text100K<n<1M0 likes53 downloads3mo agoHugging Face10nathansutton /data-science-job-descriptions Data Science Job DescriptionsThese data encompass the title, company, and description of the outer-join job board between October 2021 and today. license: wtfpl task_categories: - text-classification - feature-extraction language: - en tags: - jobs pretty_name: ds-jobs size_categories: - 1K<n<10K text1K<n<10K3 likes42 downloads3y agoHugging Face11stindardlogic /data-science-workflows-sft-100k Data Science Workflows SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment. Motivation Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.texttext-generation100K<n<1M0 likes39 downloads2mo agoHugging Face12Hamzasajjad38 /data-science-chatbot 📊 Data Science Chatbot Dataset (2000 Samples) 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts. This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a Data Science Tutor Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.texttext-generation1K<n<10K0 likes11 downloads5mo agoHugging Face13hadxs /Connor-Data-Sciences_humainestext10K<n<100K0 likes7 downloads3mo agoHugging Face14hadxs /Connor-Data-Sciencestext10K<n<100K0 likes6 downloads3mo agoHugging Face15hadxs /Connor-Data-AEX-sciencestext1K<n<10K0 likes5 downloads3mo agoHugging Face16drxzero /Final-ALevel-Science-Datatext1K<n<10K1 likes4 downloads3mo agoHugging Face17hadxs /Connor-Data-Datasciencetext10K<n<100K0 likes4 downloads3mo agoHugging Face18hadxs /Connor-Data-AEX-sciences_humainestext1K<n<10K0 likes4 downloads3mo agoHugging Face19eshmoideas /DataScience-ML-DATASETSgated MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/eshmoideas/DataScience-ML-DATASETS.texttext-generation1M<n<10M1 likes3 downloads3mo agoHugging Face20hadxs /Connor-Data-AEX-datasciencetext100K<n<1M0 likes3 downloads3mo agoHugging Face21beldua /english-data-science-basics-30textn<1K0 likes2 downloads9mo agoHugging Face22Pranav2640 /intrain_data_sciencetextn<1K0 likes1 downloads2y agoHugging Face23DataScienceGroup /chat01textn<1K0 likes1 downloads8mo agoHugging Face24DataScienceGroup /imageimagen<1K0 likes1 downloads9mo agoHugging Face25DataScienceGroup /image-text-jsontextn<1K0 likes1 downloads9mo agoHugging Face26DataScienceGroup /image-text2textn<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.