datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
abuse-scanner-bot-datasetscandisent
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/timpal0l/scandisent.ninapro-db5-s2s-certified
NinaPro DB5 — S2S Physics Certified (v1.7.0)
Physics-certified windows from NinaPro DB5 forearm EMG+IMU dataset.
Each window validated against 8 biomechanical laws using S2S.
Bad training data costs you months. S2S finds it in milliseconds.
What this adds
Column
Description
tier
GOLD / SILVER / BRONZE / REJECTED
score
0–100 physics compliance score
laws_passed
Which of 8 laws passed
verdict
Human-readable quality statement
recommendation
Actionable… See the full description on the dataset page: https://huggingface.co/datasets/Scan2s/ninapro-db5-s2s-certified.africa-tls-deployment-scan
Africa TLS Deployment Scan | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: csv - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tls-deployment-scan.CAD-Scania
CAD-Scania
Dataset Summary
CAD-Scania is a single-source continual anomaly detection benchmark scenario for predictive maintenance in the automotive/industrial domain. It is derived from the SCANIA Component X Dataset, a real-world multivariate time-series dataset of anonymized engine-component operational readouts, repair records, and specifications from SCANIA trucks, and converts it into a sequence of concept-grouped tasks.
The dataset contains 51,084 samples… See the full description on the dataset page: https://huggingface.co/datasets/lifelonglab/CAD-Scania.humaneval-patchopenai_humaneval dataset, with one-line bugs of various forms in the solutions. These bugs are generated using abstract syntax trees (ASTs) in Python, to randomly sample variables, functions, and expressions in the function body and replace them with other variables, functions and expressions respectively.
The data contains two splits- control and print. Code for generating humaneval-patch is provided here. Developed as part of an investigation of language models' ability to utilize print… See the full description on the dataset page: https://huggingface.co/datasets/scandukuri/humaneval-patch.llm-vuln-scannerlinical_inference_debt_scanner_v0.1Clinical Inference Debt Scanner
PurposeDetect when a clinical plan relies on stacked assumptions rather than evidence.
You receive:
evidence_signals
a narrative_chain
a planned_action
You output:
inference_debt_level0 to 3
debt_itemthe single most dangerous leap
paydown_stepthe corrective step that restores evidence grounding
Debt scale0 none1 minor2 moderate3 severe
Scoring
debt_level_scoregraded by distance from gold
debt_item_similaritytoken overlap similarity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/linical_inference_debt_scanner_v0.1.altpath-proteome-scan
vedatonuryilmaz/altpath-proteome-scan
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('vedatonuryilmaz/altpath-proteome-scan')
Skincare-ingredientsseo-ai-scanner-benchmarks
SEO AI Visibility Scanner Benchmarks
Benchmark dataset of 20 brand visibility scan cases with individual scores for SEO signal, AI visibility, content signal, authority signal, gap, and coverage across 5 AI platforms — Google, ChatGPT, Gemini, Perplexity, and Microsoft Copilot.
Built by GetPR.Buzz.
Dataset Description
This dataset contains benchmark data for the SEO AI Visibility Scanner — a structured scanning framework that evaluates brand visibility across… See the full description on the dataset page: https://huggingface.co/datasets/getpr-buzz/seo-ai-scanner-benchmarks.scandi-reddit-filtered
Dataset Card for ScandiRedditFiltered
Dataset Summary
ScandiRedditFiltered is manually filtered and post-processed corpus consisting of comments from ScandiReddit.
The intended use of the filtered sentences is for Text-To-Speech (TTS) models.
Supported Tasks and Leaderboards
Training language models is the intended task for this dataset. No leaderboard is active at this point.
Languages
The dataset is available in Danish (da).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit-filtered.inference_debt_scanner_v01
Inference Debt Scanner (v0.1)
This dataset checks when a model spends inference debt:
making claims that outrun its evidence.
Each row marks:
the risky claim fragment
the type of inference being made
the signal that the claim is overextended
the action a careful model should take
The aim is simple:
turn lazy certainty into clear, traceable reasoning.
Network-ScanMNLP_M3_documentsMNLP_M3_full_rag_chunksscannedlines108k lines of 18th Century iambic pentameter, scraped via xquery from the Eighteenth Century Poetry Archive xml database.
https://www.eighteenthcenturypoetry.org/
Scanned using prosodic.py.
https://github.com/quadrismegistus/prosodic
"Score" roughly equates to metrical complexity or divergence from strict iambic patterning.
See https://github.com/cretic/metricalgpt for more information including metrical constraints
rag_chunksMNLP_M3_rag_chunksMNLP_M2_rag_datasetscan-news
