datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MachineLearning
Machine Learning Tier
This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection.
The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.tokenization-multiplicity-data
Dataset: Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service
This dataset contains the official experiment inference traces for the paper Tokenization Multiplicity Leads to Arbitrary Price Variation in LLM-as-a-service by Ivi Chatzi, Nina Corvelo Benz, Stratis Tsirtsis and Manuel Gomez-Rodriguez.
📂 Dataset Structure
The dataset is organized into folders as follows:
.\{model}\{task}\{lang}\{seed}_{10*temperature}.jsonl
where {model}… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/tokenization-multiplicity-data.machine_learning_questions
Dataset Card for "machine_learning_questions"
More Information needed
strategic-ttc-data
Dataset: Strategic Test-Time Compute (TTC)
This dataset contains the official experiment inference traces for the paper "Test-Time Compute Games" (arXiv:2601.21839).
It includes full model generations, token counts, and correctness verifications for various Large Language Models (LLMs) across three major reasoning benchmarks: GSM8K, AIME, and GPQA.
This data allows researchers to analyze the relationship between test-time compute and model performance without needing to re-run… See the full description on the dataset page: https://huggingface.co/datasets/Human-Centric-Machine-Learning/strategic-ttc-data.mmlu-pro-cot-opus5bbh-cot-opus5musr-cot-opus5IndustryCorpus2_artificial_intelligence_machine_learning
IndustryCorpus2: Artificial Intelligence
This repository contains the IndustryCorpus2: Artificial Intelligence domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_artificial_intelligence_machine_learning.machine_learning_coursesquantum-machine-learning-theory
Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data
A research-depth, proof-oriented vertical on the learning theory of quantum models and quantum data. Covers why parameterized quantum circuits train or don't (barren plateaus), what they can represent, when they generalize or provably beat classical models, and — for quantum data — how to predict properties of unknown states/channels with few measurements (classical… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-theory.quantum-machine-learninga continuous data scrape of arxiv and google scholar papers of quantum machine learning papers particularly regarding climate.
Machine-Learning-Socratic-DatasetSO-Python_QA-Data_Science_and_Machine_Learning_classquantum-machine-learning-models
Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures
A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.machine-failure-logsboolq-cot-opus5Machine-Learning-Credit-Card-Fraud-Detection-Projectmachine-learning-glossary-ai
📚 Machine Learning & AI Technical Glossary Dataset
Curated benchmark dataset covering core terminology, mathematical formulations, and engineering principles across Deep Learning, Transformers, and MLOps.
Maintained and documented by AheadMint.
📌 Dataset Overview
Category
Key Concepts
Reference Documentation
Neural Networks
Backpropagation, Attention, Loss Functions
AheadMint Deep Learning
Generative AI
RAG Architectures, Vector Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/aheadmint/machine-learning-glossary-ai.Machine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio.
It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.Adversarial_Machine_Learning_TextFooler_DatasetAdversarial Machine Learning TextFooler Dataset
Overview
This dataset, adversarial_machine_learning_textfooler_dataset.jsonl, is designed for research and evaluation in adversarial machine learning, specifically focusing on text-based adversarial attacks. It contains pairs of original and adversarial text samples, primarily generated using the TextFooler attack method, along with other adversarial techniques such as Homoglyph, Number Substitution, Typo, Emoji, Semantic Shift, Paraphrase, and… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Adversarial_Machine_Learning_TextFooler_Dataset.Machine-Learning-Project
Datasets of AI2612
3D models(.glb) and rendergraphs about architecture.
The source data are from sketchfab.
Antiviral-Prediction-using-Machine-Learning-Cheminformaticsmachine-learningImplementation_1-Machine_learningmachine-learning-wikipedia-dataset
Machine Learning Wikipedia Dataset
Dataset Description
This dataset contains 50 curated Wikipedia articles related to the Machine Learning niche, including core concepts, key algorithms, conferences, and tools within the field.
Each record includes the article title, source URL, a summary, full text content, Wikipedia categories, outbound references, and associated images.
Dataset Structure
Each line in the .jsonl file is a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/Ashfaq2000/machine-learning-wikipedia-dataset.Graph_Machine_Learning_Homework_DatasetsSolid-State-Battery-Properties-Enabled-by-Machine-learningMachine-Learning-Instruct
The dataset was messily gathered from various sources such as Unsloth Github, Kohya_SS Github, Transformers docs, PEFT docs and some more.
Then it was augmented and used as seed text to generate multi-turn, updated ML conversations. So, each conversation should be self-contained and ready for training (tho this is not always the case).
