datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PlanarGSExperimentsHaSPeR
HᴀSPᴇR: An Image Repository for Hand Shadow Puppet Recognition
This repository contains the code, data, and models of the paper titled "HᴀSPᴇR: An Image Repository for Hand Shadow Puppet Recognition" accepted in the Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops (WCCA Oral).
License: Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International
Data Directory
Please navigate to the Hugging Face repository… See the full description on the dataset page: https://huggingface.co/datasets/Starscream-11813/HaSPeR.PrimeVul
Dataset Card: PrimeVul Dataset Splits
Overview
PrimeVul is a dataset crafted for vulnerability detection in C/C++ code, aimed at training and evaluating code language models under realistic conditions. This dataset card describes the pre-split version (train, validation, and test) uploaded to Hugging Face, based on the PrimeVul-v0.1 release. It includes approximately 7,000 vulnerable functions and 229,000 benign functions from real-world projects, covering over 140 Common… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/PrimeVul.github-code-2025-above-2-starsthe-stack-v2-dedup-filtered-500-stars-100-forks-contentsCVEfixes_v1.0.8
CVEfixes Data Splits README
This repository contains data splits derived from the CVEfixes_v1.0.8 dataset, an automated collection of vulnerabilities and their fixes from open-source software. The dataset has been processed and split into training, validation, and test sets to facilitate machine learning and vulnerability analysis tasks. Below, you’ll find details about the splits, problematic CVEs excluded due to memory constraints, and a comprehensive guide on how to recreate… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/CVEfixes_v1.0.8.processed_test_splitsSTARSS23audio_visual_starss23_sontactile-mnist-touch-starstruck-syn-single-t32-320x240Documentation is available at https://github.com/TimSchneider42/tactile-mnist/blob/main/doc/datasets.md#touch-datasets.
paper-github-starsstarsOWASP_Dataset_v1.2_BenchmarkTests_AND_Expectedresults-1.2.csvSTARSS23_extraMy_DiverseVul
DiverseVul Dataset Card
Dataset Description
DiverseVul is a comprehensive and meticulously curated dataset of vulnerable and non-vulnerable C/C++ source code, designed to advance deep learning-based vulnerability detection research. Originally introduced by Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner in their seminal paper, this dataset has been expanded and uploaded to Hugging Face to facilitate broader access and application. It captures… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/My_DiverseVul.LineVul
LineVul Dataset Splits
This dataset provides the train, validation, and test splits of the LineVul dataset, originally introduced in the paper "LineVul: A Transformer-based Line-Level Vulnerability Prediction" by Michael Fu and Chakkrit Tantithamthavorn. The dataset is designed for predicting software vulnerabilities at the line level in C/C++ code using transformer-based models. It was sourced from the LineVul replication package available at… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/LineVul.gaia-dr3-gold-sample-carbon-stars
Gaia DR3 gold sample carbon stars
This table lists the Gaia DR3 source identifiers in the gold sample of carbon stars selected from the larger candidate list. ESA describes the selected stars as having C2 and CN molecular bands significantly stronger than ordinary M stars. It is an identifier table for joining the selected sample to other Gaia DR3 measurements.
Use
python -m venv .venv && .venv/bin/pip install datasets pyarrow
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-gold-sample-carbon-stars.ai-stars-2026
ai-stars-2026
AI data collected daily by Legion API.
🔑 API Access — Updated Daily
Live data via Legion AI API
Free: 100 req/day · Pro €29/month: 50K req/day + full fields
curl "https://api.legion-api.com/incidents?limit=10" -H "X-API-Key: YOUR_PRO_KEY"
Premium archive (1,285 incidents, full analysis): AISI Intelligence Pack €299
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c = LegionClient()
print(c.guard(["openai"… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-stars-2026.shopee-reviews-tl-stars
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Tagalog (TL)
Dataset Structure
Data Instances
A typical data point, comprises of a text and the corresponding label.
An example from the YelpReviewFull test set looks as follows:
{
'label': 2… See the full description on the dataset page: https://huggingface.co/datasets/scaredmeow/shopee-reviews-tl-stars.ParaMAWPS
Math Word Problem Solving by Generating Linguistic Variants of Problem Statements
This repository contains the code, data, and models of the paper titled "Math Word Problem Solving by Generating Linguistic Variants of Problem Statements" published in the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop).
The work is outlined in a more detailed and expository manner in our Bachelor of Science (B.Sc.) thesis… See the full description on the dataset page: https://huggingface.co/datasets/Starscream-11813/ParaMAWPS.BanglaBook
BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews
This repository contains the code, data, and models of the paper titled "BᴀɴɢʟᴀBᴏᴏᴋ: A Large-scale Bangla Dataset for Sentiment Analysis from Book Reviews" published in the Findings of the Association for Computational Linguistics: ACL 2023.
License: Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International
Data Format
Each row consists of a book review sample. The… See the full description on the dataset page: https://huggingface.co/datasets/Starscream-11813/BanglaBook.shona-synthetic-corpusmedical-device-regulatory-graft-rag-1350
🏥 Medical Device Regulatory & Clinical Compliance RAG Dataset (1,350 Samples)
This dataset contains 1,350 highly curated, 100% LLM-synthesized RAG samples for training Small Language Models (SLMs: 1B–4B parameters) in high-stakes Medical Device Regulatory & Quality Compliance.
Methodological Foundation:
Pioneer / Prometheus Closed-Loop Curriculum Synthesis: Multi-slice curriculum covering 5 core operational failure modes.
Elsevier Computer Standards & Interfaces… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/medical-device-regulatory-graft-rag-1350.shona-correctionsgcvs-variable-stars
General Catalogue of Variable Stars (GCVS)
Credit: NASA/ESA/Hubble
Part of a dataset collection on Hugging Face.
Dataset description
The General Catalogue of Variable Stars (GCVS) is the canonical reference catalog of variable stars, maintained since 1948 by the Sternberg Astronomical Institute at Moscow State University.
Variable stars are stars whose brightness changes over time, either due to intrinsic physical processes (pulsation, eruption, rotation)… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/gcvs-variable-stars.MultiSWE_Demo
Multilingual SWE-Bench Task Sample
A sample package of a multilingual software engineering task benchmark dataset, designed to evaluate AI Agents' capabilities in code fixing and feature implementation on real open-source projects. Runs on the Harbor evaluation framework.
Overview
This dataset contains 46 tasks covering 9 programming languages and 6 task types, sourced from real open-source repository commits. Each task provides a Chinese problem_statement (problem… See the full description on the dataset page: https://huggingface.co/datasets/StarsfieldAI/MultiSWE_Demo.preprocessed_stars
Dataset Card for "preprocessed_stars"
More Information needed
MSR_data_cleaned
MSR Data Cleaned - C/C++ Code Vulnerability Dataset
📌 Dataset Description
A curated collection of C/C++ code vulnerabilities paired with:
CVE details (scores, classifications, exploit status)
Code changes (commit messages, added/deleted lines)
File-level and function-level diffs
🔍 Sample Data Structure from original file
+---------------+-----------------+----------------------+---------------------------+
| CVE ID | Attack Origin | Publish Date… See the full description on the dataset page: https://huggingface.co/datasets/starsofchance/MSR_data_cleaned.star-smallSynthetic data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions.
Project page: https://dhruveshp.com/projects/ilm
PrimeVul_For_unsloth_V1.0_failed_finetuning
