datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cits4012_A1_2026_medical_abstractsA11YBench
A11YBench
A Benchmark for Web Accessibility Repair
😃Dataset Summary
A11YBench consists of 60 real-world web projects, encompassing 147 web pages and 8,886 accessibility violations detected by the IBM Accessibility Checker using Check Rule 2025.09.03.
The projects vary substantially in size, from 123 to 43,198 source files and from 3,610 to 1,555,532 lines of code, covering both lightweight documentation sites and large production-grade applications.
This scale ensures… See the full description on the dataset page: https://huggingface.co/datasets/LLM4APR/A11YBench.lopi-ds-conv-a1mad-blow-a17d80
mad-blow-a17d80
Synthetic sensors test data: 53 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/tarosato/mad-blow-a17d80.strong-government-a16a25
strong-government-a16a25
Synthetic weather test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/yumiko89/strong-government-a16a25.pwc747_a10865__paper__P02__2023__high__llm_safety
Northwind Support Tickets Archive
A derived dataset combining service interaction logs with survey responses for support ticket analysis.
Upstream Sources
This dataset is derived from the following upstream source datasets:
Northwind Service Interaction Logs (TianfuXinqu/pwc747_a10865__paper__P05__2022__high__llm_safety)
Northwind Customer Survey Responses (TianfuXinqu/pwc747_a10865__paper__P06__2018__low__federated_learning)
Commercial Use… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/pwc747_a10865__paper__P02__2023__high__llm_safety.eastern-reputation-a1327d
eastern-reputation-a1327d
Synthetic products test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Azure-Ivan98/eastern-reputation-a1327d.pwc747_a10865__paper__P01__2024__high__federated_learning
Northwind Product Reviews Corpus
A derived dataset combining customer profiles with store reviews to support product review analysis.
Upstream Sources
This dataset is derived from the following upstream source datasets:
Northwind Customer Profiles (TianfuXinqu/pwc747_a10865__paper__P03__2021__low__federated_learning)
Northwind Store Reviews (TianfuXinqu/pwc747_a10865__paper__P04__2019__high__model_compression)
Commercial Use
Commercial Use:… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/pwc747_a10865__paper__P01__2024__high__federated_learning.tourism-package-prediction-dataswedish-pre-a1-scenario-classifier-dataset
Swedish Pre-A1 Scenario Classification Dataset
This dataset contains 150 short Swedish learner sentences for text classification. It is designed for absolute beginner / pre-A1 learners and aligned with beginner Swedish lecture themes.
Labels
food_shop
family_school
health_places
transport
home_places
social_intro
Each label has 25 examples.
Columns
id: unique example id
text: Swedish learner sentence used as model input
english: English translation
chinese:… See the full description on the dataset page: https://huggingface.co/datasets/SeanSha30/swedish-pre-a1-scenario-classifier-dataset.uspto-patent-datashyrai-a1c-quality-control
Shyrai A1c Quality Control Dataset
Description
This dataset contains real-world quality control (QC) data collected during the production and testing of the Shyrai A1c glycated hemoglobin analyzer, used for diabetes diagnostics.
The dataset is designed to support research in:
AI-driven quality management systems (QMS)
anomaly detection
predictive quality analytics
medical device manufacturing
Dataset Structure
The dataset includes the following types of… See the full description on the dataset page: https://huggingface.co/datasets/BekzatK/shyrai-a1c-quality-control.dsv-gated-chain-turw0z1Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/a1b8h04i/Hinglish-Everyday-Conversations-1M.SanskritDatasetlopi-ds-gated-a1testing_01-e61c8c4f-3dcd-4cd8-a1db-eadb7f5c3394a10_command_scriptsa_data_origin감성대화 말뭉치와 한국어 단발성 대화 병합 데이터셋
label : "불안", "분노", "상처", "슬픔", "당황", "기쁨", "놀람"
sqli-a1-enva1auto22t5dmuceya5ax
