datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SSL
Scheduling-Structural-Logical Skill Dataset
This Hugging Face dataset release contains the data artifacts for the Scheduling-Structural-Logical (SSL) representation of agent skills.
It includes three loadable configurations:
annotated_skill_corpus: 6,184 normalized SSL records paired with derived skill metadata.
ssl_skill_discovery: the SSL-SkillDiscovery benchmark with 431 intent-level queries over the 6,184-skill candidate set.
ssl_risk_assessment: the SSL-RiskAssessment… See the full description on the dataset page: https://huggingface.co/datasets/COOLPKU/SSL.SSLD-200
Song Structure and Lyric Dataset (SSLD-200)
DataSet used to evaluate song structure parsing and lyrics transcription. SSLD-200 consists of 200 songs, 100 English and 100 Chinese, collected entirely from YouTube, with a total duration of 13.9 hours.
The lyric_norm in the format [structure][start:end]lyric
The structure is the label from StructureAnalysis for the segment.
The start and end are the segment’s start and end times.
The lyric is the recognized lyrics.… See the full description on the dataset page: https://huggingface.co/datasets/waytan22/SSLD-200.pqc-ssl-scans
PQC Vulnerability Scan Dataset
SSL/TLS certificate scans of 45 major finance, healthcare, and government domains, scored for post-quantum cryptography (PQC) migration urgency.
Dataset Description
Each row represents a live SSL certificate scan performed on 2026-03-24 using hiero-cli-pqc.
Features
Feature
Type
Description
domain
string
Scanned domain name
key_algorithm
string
Public key algorithm (RSA, ECDSA, Ed25519)
key_size
int
Key size in bits… See the full description on the dataset page: https://huggingface.co/datasets/Q-GRID/pqc-ssl-scans.us-ssl-mf
SSL Ultrasound Representation Dataset
IQ-demodulated multi-focal ultrasound channel data from the Stanford Ultrasound RF
Channel dataset, aggregated for self-supervised pre-training (MAE / JEPA /
contrastive). The signals are not beamformed — each frame retains its
per-transducer channel data after sub-aperture combination and IQ demodulation.
Files
File
Description
ultrasonic_dataset.zarr/
6D float32 zarr array, Blosc-LZ4 compressed, stored as an… See the full description on the dataset page: https://huggingface.co/datasets/benbarkow/us-ssl-mf.
