CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ik45 /data-science-en-id Data Science EN-ID Parallel Corpus (Scientific Domain) Dataset Description This dataset is a curated English-Indonesian (EN-ID) parallel corpus specifically designed for the Scientific and Data Science domains. It was developed to support the training of Machine Translation (NMT) models and Large Language Models (LLMs) to better handle technical terminology, academic structures, and formal scientific language. Primary Languages: English (EN) and Indonesian (ID) Domain:… See the full description on the dataset page: https://huggingface.co/datasets/Ik45/data-science-en-id.texttext-generation10M<n<100M0 likes193 downloads6mo agoHugging Face02KellanF89 /Cannabis_Science_Data Cannabis Science Literature QA Dataset This dataset contains 161,170 high-quality question-answer pairs derived from over 400 peer-reviewed cannabis science research papers and textbooks. Created to advance AI research in cannabis science and medical applications, it provides a comprehensive resource for training language models on cannabis-related scientific knowledge. Dataset Details Dataset Description This dataset was systematically generated from a curated… See the full description on the dataset page: https://huggingface.co/datasets/KellanF89/Cannabis_Science_Data.question-answering100K<n<1M5 likes114 downloads1y agoHugging Face03Emulated-Inc /data-science-code-training-pool Data science code training pool Public questions about writing Python with numpy, pandas, matplotlib, scikit-learn, scipy, pytorch and tensorflow, each with the code that answers it, gathered from the datasets named below at the pinned revisions and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 339575 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/data-science-code-training-pool.text-generation100K<n<1M0 likes51 downloads11d agoHugging Face04stindardlogic /data-science-workflows-sft-100k Data Science Workflows SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level data science practice across data cleaning, EDA, ML pipelines, feature engineering, SQL analytics, statistical analysis, model evaluation, visualization, and production deployment. Motivation Data science is one of the most in-demand technical skills — companies need models that can reason through real analytical problems with the rigor of a senior data scientist. Models… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/data-science-workflows-sft-100k.texttext-generation100K<n<1M0 likes38 downloads2mo agoHugging Face05tekwhisperer /Cannabis_Science_Data Cannabis Science Literature QA Dataset This dataset contains 161,170 high-quality question-answer pairs derived from over 400 peer-reviewed cannabis science research papers and textbooks. Created to advance AI research in cannabis science and medical applications, it provides a comprehensive resource for training language models on cannabis-related scientific knowledge. Dataset Details Dataset Description This dataset was systematically generated from a curated… See the full description on the dataset page: https://huggingface.co/datasets/tekwhisperer/Cannabis_Science_Data.question-answering100K<n<1M0 likes16 downloads6mo agoHugging Face06Hamzasajjad38 /data-science-chatbot 📊 Data Science Chatbot Dataset (2000 Samples) 🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts. This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way. 🎯 Objective The goal of this dataset is to: Train LLMs to act as a Data Science Tutor Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.texttext-generation1K<n<10K0 likes14 downloads5mo agoHugging Face07pymlex /datasciencejobs-tg Data Science Jobs Telegram Overview This dataset contains 1.8k descriptions of ML job offers from the t.me/datasciencejobs Telegram channel. The data was scraped using this script. Each row includes the post ID, publication date, number of views, and the raw text. Disclaimer Scraping violates Telegram's Terms of Service. This data is provided for educational purposes only. Benchmark We also provide a small benchmark with 243 QA pairs generated by… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/datasciencejobs-tg.text-generation1K<n<10K0 likes9 downloads5mo agoHugging Face08eshmoideas /DataScience-ML-DATASETSgated MachineLearningLM Pretraining Corpus This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/eshmoideas/DataScience-ML-DATASETS.texttext-generation1M<n<10M1 likes2 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.