datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MachineLearning
Machine Learning Tier
This dataset is a collection of synthetic microlensing light curves from the Nancy Grace Roman Space Telescope Galactic Bulge Time Domain Survey. It is intended for the training and benchmarking of machine learning models for microlensing event classification, parameter estimation, and anomaly detection.
The raw distribution of event properties is not representative of what Roman will see, but should span a statistically larger set of events. More… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/MachineLearning.machinelearninglm-scm-synthetic-tabularml
MachineLearningLM Pretraining Corpus
This repository contains the pretraining corpus for MachineLearningLM, a framework designed to equip large language models (LLMs) with robust in-context machine learning (ML) capabilities. The dataset consists of ML tasks synthesized from millions of structural causal models (SCMs), spanning various shot counts up to 1,024. It is designed to enable LLMs to learn from many in-context examples on standard ML tasks purely via in-context learning… See the full description on the dataset page: https://huggingface.co/datasets/MachineLearningLM/machinelearninglm-scm-synthetic-tabularml.machine_learning_questions
Dataset Card for "machine_learning_questions"
More Information needed
Machine-Learning-Socratic-DatasetMachine_Learning_QA_Dataset_LlamaDataset created based on win-wang/Machine_Learning_QA_Collection
This Dataset was created for the finetuning test of Machine Learning Questions and Answers.
It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
The dataset is formatted for llama3 using the chat template
<|begin_of_text|><|start_header_id|>system<|end_header_id|>
Cutting Knowledge Date: December 2023
Today Date: 23 July 2024
You are a helpful… See the full description on the dataset page: https://huggingface.co/datasets/aamanlamba/Machine_Learning_QA_Dataset_Llama.Machine_Learning_QA_CollectionThis Dataset was created for the finetuning test of Machine Learning Questions and Answers. It combined 7 Machine Learning, Data Science, and AI Questions and Answers datasets.
This collection dataset only extracted the questions and answers from those datasets mentioned below. The original collection of all datasets contains about 12.4k records, which are split into train set, dev set, and test set in a 7:1:2 ratio.
It was used to test the Finetuning Gemma 2 model by MLX on Apple Silicon.… See the full description on the dataset page: https://huggingface.co/datasets/win-wang/Machine_Learning_QA_Collection.Machine-Learning-Instruct
The dataset was messily gathered from various sources such as Unsloth Github, Kohya_SS Github, Transformers docs, PEFT docs and some more.
Then it was augmented and used as seed text to generate multi-turn, updated ML conversations. So, each conversation should be self-contained and ready for training (tho this is not always the case).
Machine-Learning-QA-datasetmachine-learning-forecasting-dataMachineLearning_EmojiDataset_Nov17machine_learningmachine-learninghttps://github.com/ostad-ai/Machine-Learning?tab=readme-ov-file#machine-learing-and-data-science
machinelearningMachine-Learning-QA-Datasetmachine_learningMachine-Learning-Car-Price-Prediction-ProjectMachineLearningDataset
Dataset Card for "MachineLearningDataset"
More Information needed
MachineLearning_DDAMachine-Learningdatasetsmachine-learningMachine_LearningMachineLearning
