datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
booksummaries_cleanedMSR_data_cleaned
MSR Data Cleaned - C/C++ Code Vulnerability Dataset
📌 Dataset Description
A curated collection of C/C++ code vulnerabilities paired with:
CVE details (scores, classifications, exploit status)
Code changes (commit messages, added/deleted lines)
File-level and function-level diffs
🔍 Sample Data Structure from original file
+---------------+-----------------+----------------------+---------------------------+
| CVE ID | Attack Origin | Publish Date… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/MSR_data_cleaned.cleaned_prompt_safety_datasetgrab-safe-driver-telematics-cleaned-datasetcardiovascular-cleaned-datasetCT-RATE-Dataset-cleanedCA_Weather_Fire_Dataset_Cleaned📦 Dataset Card: CA_Weather_Fire_Dataset_Cleaned
Dataset Summary
This dataset contains cleaned and preprocessed weather and fire incident data for California (1984–2025).
The original dataset, California Weather and Fire Prediction Dataset (1984–2025) with Engineered Features, includes features such as temperature, humidity, wind speed, fire occurrence, and seasonal indicators.
From the Original Dataset, I changed the data types to floats, rearranged the columns, removed… See the full description on the dataset page: https://huggingface.co/datasets/MaxPrestige/CA_Weather_Fire_Dataset_Cleaned.databricks-dolly-15k-cleaned
Summary
databricks-dolly-15k-cleaned is a cleaned up version of the popular databricks-dolly-15k dataset after automatically removing low quality datapoints detected using Cleanlab.
alpaca_data_cleaned_bhojpuri
Dataset Card for Dataset Name
This repository contains a translated version of the Alpaca-Cleaned dataset, originally provided by Yahma on Hugging Face. The dataset has been translated into Bhojpuri, a language spoken in the northern-eastern part of India and the Terai region of Nepal.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The Alpaca-Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/SatyamDev/alpaca_data_cleaned_bhojpuri.OpenOrca_cleaned_kor_linkbricks_single_dataset_with_prompt_text_huggingfacecleaned_questions_answers_datasetg13-cleaned-datasetUsing AGBonnet/augmented-clinical-notes dataset and additionally processing such as chunking, creating vector embeddings and cleaning up columns on top of it.
cleaned_dataset_prvy_plcymerged_movies_books_cleanedcleaned_ta_test_datacleaneddatatoxicity_dataset_cleanedjamare-cleaned-text-datasetcleaned_qa_datasetFinal_cleaned_data_06_03cleaned-datacleaned_datasetdata_jobs_cleanedcleaned_scaled_fire_data.csvcleaned_orca_dataset.csvnew_combined_dataset_cleanedcleaned-datacicids-dataset-cleanedengine-sensor-data-cleanedalpaca_data_cleaned_standardized
