datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
D_ASAP-AES
D_ASAP-AES
This is the train, test, and validation split of the ASAP Automated Essay Scoring dataset,
prepared for use with the S-GRADES benchmark.
Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-AES on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asap_aes,
title={ASAP Automated Essay Scoring}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-AES.pd-discovery-benchmark-dashboard
Parkinson's Disease Discovery Benchmark Dashboard
Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery.
This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.DasanCallDial
DasanCallDial
DasanCallDial is the first large-scale Korean benchmark built specifically for
dialogue-level ASR error correction. It contains 1,974 real civil-complaint call
dialogues (115,460 utterances) placed to the 120 Dasan Call Foundation, Seoul's
municipal civic-information hotline. Each utterance pairs the transcription produced by a
production speech-recognition system with a human-verified ground truth.
Unlike datasets built by injecting synthetic noise, DasanCallDial… See the full description on the dataset page: https://huggingface.co/datasets/zgold5670/DasanCallDial.D_ASAP-SAS
D_ASAP-SAS
This is the train, test, and validation split of the ASAP Short Answer Scoring dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation.
For the original dataset with labels, see below.
Original Dataset
🔗 ASAP-SAS on Kaggle
Citation
If you use this dataset, please cite the original:
@misc{asapsas2012,
author={Barbara and Hamner, Ben and Morgan, Jaison and lynnvandev and… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP-SAS.motorcycle-accident-driving-datasets
Dataset Summary
The dataset consisted of 2 types of cases; accident and driving while riding a motorcycle. 68 accident cases and 68 driving cases are prepared. 30 fps and 852x480 by default. It might be helpful when you train a model to infer whether a video is a motorcycle crash or not. One thing you should know about is 'driving videos' are not typically motorcycle driving. Most 'driving videos' are dashcams in the car. However, all the videos about accidents are motorcycle… See the full description on the dataset page: https://huggingface.co/datasets/smart-dashcam/motorcycle-accident-driving-datasets.D_ASAP_plus_plus
D_ASAP_plus_plus
This is the train, test, and validation split of the ASAP++ dataset, prepared for use with the S-GRADES benchmark. Ground truth labels have been removed to prevent leakage during evaluation.
Original Dataset
ASAP++ enriches the original ASAP dataset with attribute-specific essay scores (content, organization, style, etc.).
🔗 ASAP++ Official Page
Citation
If you use this dataset, please cite the original:
@inproceedings{mathias2018asap++… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP_plus_plus.Cladder_v1
Reference
The dataset created from Cladder Project.
Paper - https://arxiv.org/abs/2312.04350
Git - https://github.com/causalNLP/cladder
subreddit-postsDataset of titles of the top 1000 posts from the top 250 subreddits scraped using PRAW.
For steps to create the dataset check out the dataset script in the GitHub repo.
Grafana-Community-DashboardsThis is a raw dump of the dashboard json hosted at https://grafana.com/grafana/dashboards/, taken on 06-06-23.
Dashboards themselves are json; related metadata is retained for filtering purposes (e.g., by number of downloads) to help in identifying useful data.
Dashboards may contain many different query languages, may range across many versions of Grafana, and may be completely broken (since anyone can upload one).
JSON structure varies considerably between different dashboards, and finding… See the full description on the dataset page: https://huggingface.co/datasets/sandersaarond/Grafana-Community-Dashboards.D_ASAP2
D_ASAP2
This is the train, test, and validation split of the ASAP 2.0 dataset, prepared for use with the
S-GRADES benchmark. Ground truth labels have
been removed to prevent leakage during evaluation.
Original Dataset
🔗 ASAP 2.0 on Kaggle
Citation
If you use this dataset, please cite the original:
@article{crossley2025asap2,
title={A large-scale corpus for assessing source-based writing quality: ASAP 2.0},
author={Crossley, Scott A. and Baffour… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_ASAP2.anime-or-notuzbek_homonym_affixes
Uzbek Homonym Affixes Dataset
Dataset link on Hugging Face
📖 Description
This dataset contains Uzbek homonym affixes (omonim qo‘shimchalar) with their occurrences in different parts of speech.The dataset is designed to support Uzbek NLP research, especially in the fields of:
Morphological analysis
Part-of-speech tagging
Word sense disambiguation
Computational linguistics
Each row represents an affix and its possible usage across multiple word classes.… See the full description on the dataset page: https://huggingface.co/datasets/dasturbek/uzbek_homonym_affixes.eval_dashboardbeam-dashboard-report-data
BEAM Dashboard - Report Data
Description
The BEAM (Bacteria, Enterics, Amoeba, and Mycotics) Dashboard is an interactive tool to access and visualize data from the System for Enteric Disease Response, Investigation, and Coordination (SEDRIC). The BEAM Dashboard provides timely data on pathogen trends and serotype details to inform work to prevent illnesses from food and animal contact.
Dataset Details
Publisher: Centers for Disease Control and Prevention… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/beam-dashboard-report-data.dashentokeniser
license: mit
task_categories:
- audio-classification
language:
- en
tags:
- audio
- environmental-sound
- error-analysis
- benchmark
pretty_name: Audio Classifier Mistake Analysis
size_categories:
- n<100
# Audio Classifier Mistake Analysis
A small error-analysis dataset capturing 10 diverse misclassifications made by an audio classifier on personal audio files. Each row records the true label, what the model was expected to predict, and what it actually predicted —… See the full description on the dataset page: https://huggingface.co/datasets/mahmudaminu/dashentokeniser.beam-dashboard-isolates-by-hhs-region
BEAM Dashboard - Isolates by HHS Region
Description
The BEAM (Bacteria, Enterics, Amoeba, and Mycotics) Dashboard is an interactive tool to access and visualize data from the System for Enteric Disease Response, Investigation, and Coordination (SEDRIC). The BEAM Dashboard provides timely data on pathogen trends and serotype details to inform work to prevent illnesses from food and animal contact.
Dataset Details
Publisher: Centers for Disease Control and… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/beam-dashboard-isolates-by-hhs-region.hospital-dashboard
Hospital Dashboard
Description
This table captures percentage increases in hospital reporting,
The percentage of hospitals reporting at least once in the week
The percent of hospitals reporting on an average day in the week
The percentage of hospitals reporting every day of the week
The percentage of hospitals reporting 100% of data in week
Dataset Details
Publisher: HHS Office of the Chief Data Officer
Geographic Coverage: National
Last Modified:… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/hospital-dashboard.beam-dashboard-serotypes-of-concern-illnesses-and
BEAM Dashboard - Serotypes of concern: Illnesses and Outbreaks
Description
The BEAM (Bacteria, Enterics, Amoeba, and Mycotics) Dashboard is an interactive tool to access and visualize data from the System for Enteric Disease Response, Investigation, and Coordination (SEDRIC). The BEAM Dashboard provides timely data on pathogen trends and serotype details to inform work to prevent illnesses from food and animal contact.
Dataset Details
Publisher: Centers for… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/beam-dashboard-serotypes-of-concern-illnesses-and.covid-19-mexicoDASWOW_datasetThis dataset has been taken from the replication package of the research paper: Workflow analysis of data science code in public GitHub repositories
I don't claim to be the owner of this dataset
doi: 10.1007/s10664-022-10229-z
LitMedImageDataset Name: LitMedImage – literature-derived Medical vs. Non-Medical Image Dataset
Description:
LitMedImage is a curated dataset of biomedical literature figures labeled as MEDICAL or NON-MEDICAL. The dataset is built from images extracted from PubMed Central Open Access (PMC-OA) articles and includes corresponding captions and parsed image metadata. Labels were generated using a large language model (LLM) following strict imaging definitions. This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/dasavisha/LitMedImage.student_performance_analysischeckingdasetcredit-risk-dashboard-datasin2eng2singl
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/dastharak/sin2eng2singl.wikiart_captionsdge🧠 Awesome ChatGPT Prompts [CSV dataset]
This is a Dataset Repository of Awesome ChatGPT Prompts
View All Prompts on GitHub
License
CC-0
valorant_strategies
