datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.Machine-Learning-Socratic-DatasetSO-Python_QA-Data_Science_and_Machine_Learning_classMachine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio.
It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.schemaforge-ai-and-machine-learning-11
bdtechtalks.com
Auto-refined by SchemaForge
Metadata
Topic: AI and Machine Learning
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
AI is writing code but who's reviewing it
Nvidia's ASPIRE framework accelerates robot programming
DeepSearch turbocharges product and market research
LLMs should stop thinking out loud
Codev 3.0 engineers the AI-powered dev team
The future of agentic AI is all about the harness
Machine learning in… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-and-machine-learning-11.schemaforge-machine-learning-research-13
www.nature.com
Auto-refined by SchemaForge
Metadata
Topic: Machine Learning Research
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
Machine learning is the ability of a machine to improve its performance based on previous results.
Machine learning methods enable computers to learn without being explicitly programmed.
Privacy risks from medical AI tools are not shared equally.
Learning from routine health system data builds… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-machine-learning-research-13.mmlu-machine-learningschemaforge-ai-and-machine-learning-10
bdtechtalks.com
Auto-refined by SchemaForge
Metadata
Topic: AI and Machine Learning
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
AI is writing code but who's reviewing it
Nvidia's ASPIRE framework accelerates robot programming
Goodfire's block-sparse featurizers are a breakthrough in AI interpretability
Machine learning in space is revolutionizing exploration
LLMs should stop thinking out loud
Codev 3.0 engineers the… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-and-machine-learning-10.Wikipedia_Articles_on_Machine_Learning_and_DS
Wikipedia Machine Learning Corpus (wiki-ml-corpus)
A curated dataset of over 100 Wikipedia articles related to Machine Learning, Statistics, Probability, Data Science, and Deep Learning.
This dataset is designed for use in:
NLP tasks like summarization, QA, and topic modeling
ML interview prep and curriculum design
Ontology-driven QA systems and SPARQL-based pipelines
Building structured knowledge graphs from unstructured text
Dataset Structure
Each example in the… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/Wikipedia_Articles_on_Machine_Learning_and_DS.schemaforge-machine-learning-research-12
www.nature.com
Auto-refined by SchemaForge
Metadata
Topic: Machine Learning Research
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
Machine learning is the ability of a machine to improve its performance based on previous results.
Machine learning methods enable computers to learn without being explicitly programmed.
Privacy risks from medical AI tools are not shared equally.
Learning from routine health system data builds… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-machine-learning-research-12.schemaforge-machine-learning-research-15
www.nature.com
Auto-refined by SchemaForge
Metadata
Topic: Machine Learning Research
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
Machine learning is the ability of a machine to improve its performance based on previous results.
Machine learning methods enable computers to learn without being explicitly programmed.
Privacy risks from medical AI tools are not shared equally.
Learning from routine health system data builds… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-machine-learning-research-15.MachineLearning_DDA
