datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MachineLearning
Machine Learning Tier
This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection.
The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.machine_learning_questions
Dataset Card for "machine_learning_questions"
More Information needed
mmlu-pro-cot-opus5bbh-cot-opus5musr-cot-opus5quantum-machine-learning-theory
Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data
A research-depth, proof-oriented vertical on the learning theory of quantum models and quantum data. Covers why parameterized quantum circuits train or don't (barren plateaus), what they can represent, when they generalize or provably beat classical models, and — for quantum data — how to predict properties of unknown states/channels with few measurements (classical… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-theory.quantum-machine-learninga continuous data scrape of arxiv and google scholar papers of quantum machine learning papers particularly regarding climate.
Machine-Learning-Socratic-DatasetSO-Python_QA-Data_Science_and_Machine_Learning_classquantum-machine-learning-models
Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures
A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.machine-failure-logsboolq-cot-opus5Machine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio.
It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.Solid-State-Battery-Properties-Enabled-by-Machine-learningMachine-Learning-Instruct
The dataset was messily gathered from various sources such as Unsloth Github, Kohya_SS Github, Transformers docs, PEFT docs and some more.
Then it was augmented and used as seed text to generate multi-turn, updated ML conversations. So, each conversation should be self-contained and ready for training (tho this is not always the case).
ai-auto-train-machine-learning-quantum-mindmap-visualization-builder-quantum-chip-setup-datasetmmlu-machine_learning
Dataset Card for "mmlu-machine_learning"
More Information needed
task718_mmmlu_answer_generation_machine_learning
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task718_mmmlu_answer_generation_machine_learning
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task718_mmmlu_answer_generation_machine_learning.Machine-Learning-QA-datasetschemaforge-ai-and-machine-learning-11
bdtechtalks.com
Auto-refined by SchemaForge
Metadata
Topic: AI and Machine Learning
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
AI is writing code but who's reviewing it
Nvidia's ASPIRE framework accelerates robot programming
DeepSearch turbocharges product and market research
LLMs should stop thinking out loud
Codev 3.0 engineers the AI-powered dev team
The future of agentic AI is all about the harness
Machine learning in… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-and-machine-learning-11.schemaforge-machine-learning-research-13
www.nature.com
Auto-refined by SchemaForge
Metadata
Topic: Machine Learning Research
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
Machine learning is the ability of a machine to improve its performance based on previous results.
Machine learning methods enable computers to learn without being explicitly programmed.
Privacy risks from medical AI tools are not shared equally.
Learning from routine health system data builds… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-machine-learning-research-13.mmlu-machine_learning-verbal-neg-prepend
Dataset Card for "mmlu-machine_learning-verbal-neg-prepend"
More Information needed
mmlu-machine-learningmmlu-machine_learning-neg-answer
Dataset Card for "mmlu-machine_learning-neg-answer"
More Information needed
mmlu-machine_learning-dev
Dataset Card for "mmlu-machine_learning-dev"
More Information needed
machine-learning-forecasting-dataschemaforge-ai-and-machine-learning-10
bdtechtalks.com
Auto-refined by SchemaForge
Metadata
Topic: AI and Machine Learning
Quality Score: 0.92
Source: Autonomous web scraper
Extracted Facts
AI is writing code but who's reviewing it
Nvidia's ASPIRE framework accelerates robot programming
Goodfire's block-sparse featurizers are a breakthrough in AI interpretability
Machine learning in space is revolutionizing exploration
LLMs should stop thinking out loud
Codev 3.0 engineers the… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-and-machine-learning-10.Wikipedia_Articles_on_Machine_Learning_and_DS
Wikipedia Machine Learning Corpus (wiki-ml-corpus)
A curated dataset of over 100 Wikipedia articles related to Machine Learning, Statistics, Probability, Data Science, and Deep Learning.
This dataset is designed for use in:
NLP tasks like summarization, QA, and topic modeling
ML interview prep and curriculum design
Ontology-driven QA systems and SPARQL-based pipelines
Building structured knowledge graphs from unstructured text
Dataset Structure
Each example in the… See the full description on the dataset page: https://huggingface.co/datasets/nabin2004/Wikipedia_Articles_on_Machine_Learning_and_DS.
