datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MachineLearning
Machine Learning Tier
This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection.
The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.machine_learning_questions
Dataset Card for "machine_learning_questions"
More Information needed
machine_learning_coursesMachine-Learning-Socratic-DatasetMachine-Learning-Credit-Card-Fraud-Detection-Projectmachine-learning-glossary-ai
📚 Machine Learning & AI Technical Glossary Dataset
Curated benchmark dataset covering core terminology, mathematical formulations, and engineering principles across Deep Learning, Transformers, and MLOps.
Maintained and documented by AheadMint.
📌 Dataset Overview
Category
Key Concepts
Reference Documentation
Neural Networks
Backpropagation, Attention, Loss Functions
AheadMint Deep Learning
Generative AI
RAG Architectures, Vector Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/aheadmint/machine-learning-glossary-ai.Machine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio.
It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.Machine-Learning-Project
Datasets of AI2612
3D models(.glb) and rendergraphs about architecture.
The source data are from sketchfab.
machine-learningmachine-learning-wikipedia-dataset
Machine Learning Wikipedia Dataset
Dataset Description
This dataset contains 50 curated Wikipedia articles related to the Machine Learning niche, including core concepts, key algorithms, conferences, and tools within the field.
Each record includes the article title, source URL, a summary, full text content, Wikipedia categories, outbound references, and associated images.
Dataset Structure
Each line in the .jsonl file is a JSON object with the… See the full description on the dataset page: https://huggingface.co/datasets/Ashfaq2000/machine-learning-wikipedia-dataset.Machine-Learning-Instruct
The dataset was messily gathered from various sources such as Unsloth Github, Kohya_SS Github, Transformers docs, PEFT docs and some more.
Then it was augmented and used as seed text to generate multi-turn, updated ML conversations. So, each conversation should be self-contained and ready for training (tho this is not always the case).
machine_learning_projectThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"robot_type": "so-100",
"codebase_version": "v3.0",
"total_episodes": 27,
"total_frames": 18800,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:27"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4",
"features": {… See the full description on the dataset page: https://huggingface.co/datasets/Delcastillo8/machine_learning_project.machine-learningCattleSSFR
CattleSSFR (cattle single sample face recognition) dataset
The set of 311 cropped single cattle face images comprises the CattleSSFR dataset, which is publicly available. CattleSSFR is the first reported dataset for SSFR containing 311 classes of cattle faces single images. Indicative images from CattleSSFR dataset are illustrated in Fig.2, aiming to demonstrate the diversity of different classes for the problem under study and the challenges related to the SSFR problem, e.g.… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningVisionRG/CattleSSFR.Machine-Learning-QA-datasetorena-segment-annotations
ORena FOCUS 2026 — SEGMENT supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 SEGMENT
track. This is the openly releasable half; the LapChole-FOCUS half is withheld under that
dataset's usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
1,300
vqa/hernia_mesh.jsonl — hernia videos
140
total
1,440
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-segment-annotations.orena-frame-annotations
ORena FOCUS 2026 — FRAME supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 FRAME track.
This is the openly releasable half; the LapChole-FOCUS half is withheld under that dataset's
usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
7,608
vqa/hernia_mesh.jsonl — hernia videos
1,350
total
8,958
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-frame-annotations.machine-learning-and-simulationmachine-learning-forecasting-dataMachineLearning_EmojiDataset_Nov17machine_learningMachine-Learning-for-ECG-Based-Human-Func-tional-State-AssessmentCode files for the article "Machine Learning Framework for ECG-Based Human Func-tional State Assessment Using Interpretable Feature Analysis And Unsupervised Clustering"
Abstract This repository contains the code and materials developed by the authors for a hybrid machine learning pipeline designed to assess functional states from ECG data. The pipeline integrates feature-based supervised classification with interpretable analyses and unsupervised clustering techniques. The underlying dataset… See the full description on the dataset page: https://huggingface.co/datasets/insamhlaithe/Machine-Learning-for-ECG-Based-Human-Func-tional-State-Assessment.machine-learninghttps://github.com/ostad-ai/Machine-Learning?tab=readme-ov-file#machine-learing-and-data-science
machinelearningMachine-Learning-QA-Datasetmachine_learningmachine-learning-projectsMachine-Learning-Car-Price-Prediction-Project
