datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Synthetic-AI-ML-Dataset
Synthetic-AI-ML-Dataset
Synthetic Q&A dataset on AI and Machine Learning
Dataset Details
Metric
Value
Topic
AI and Machine Learning
Total Q&A Pairs
14021
Valid Pairs
14021
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
14,941,957
Completion Tokens
17,159,263
Total Tokens
32,101,220
GPU Energy
12.9628 kWh
Sources
This dataset was generated from 474 scholarly papers:
#… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/Synthetic-AI-ML-Dataset.ai-ml-instruction-dataset
AI/ML Engineering Instruction Dataset
Comprehensive instruction dataset covering machine learning concepts, PyTorch implementations, NLP with transformers, model evaluation, and feature engineering.
Dataset Details
Dataset Description
This is a high-quality instruction-tuning dataset focused on Ai Ml topics. Each entry includes:
A clear instruction/question
Optional input context
A detailed response/solution
Chain-of-thought reasoning process
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/bernabepuente/ai-ml-instruction-dataset.ml-interview-sft-dataset
ML/AI Interview Coach — SFT Dataset
A curated dataset of 566 high-quality Q&A pairs covering ML, Deep Learning, NLP, LLMs, RAG, Vector Databases, LangChain, Agentic AI, MLOps, and more — designed for fine-tuning an ML Interview Coach model.
Dataset Summary
Stat
Value
Total Q&A pairs
566
Unique topics
75
Format
ChatML (system + user + assistant)
Language
English
Avg answer length
~800 tokens
Sources
15+ interview prep documents + hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/raghu298/ml-interview-sft-dataset.Amazon_ml_challenge_flitered_datasetarxiv-ml-qa-dataset
ArXiv ML Q&A Dataset
Dataset Description
951 high-quality Q&A pairs generated from ArXiv machine learning paper
abstracts. Built for fine-tuning language models on technical ML questions.
Repository: GitHub Repository - Source code for dataset generation, cleaning, and filtering.
How It Was Built
Scraped 2000+ ArXiv ML papers via ArXiv API (cat:cs.LG)
Cleaned — deduplication, length filter (50–400 words)
Generated — Llama-3-8B-Instruct via… See the full description on the dataset page: https://huggingface.co/datasets/vaadewoyin/arxiv-ml-qa-dataset.MLdataset
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Kartheesh/MLdataset.
