datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSPP_WAVIEMO_Audio_Text_MergedMSPP_WAV_speaker_splitsamsemo-audioMSPI_Audio_Text_MergedBALANCED_MSPP_MSPI_IEMOMSPP_MSPI_IEMOCAP_V2cmu_mosei_wav_2MSPP_WAV2Vec2cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.MSPP_Wav2Vec2_V2MSPP_SPLIT2_wav2vecMSPP_PLSMSPP_SPLIT2_wav2vec_FINALMSP_Pod_SYL6MSP_Pod_SYL4MSPP_SPLIT_Wav2vecMSP_Pod_SYL3IEMOCAP
IEMOCAP — full release, utterance level
All 10,039 segmented utterances of the USC-SAIL IEMOCAP corpus with 16 kHz audio,
transcripts, consensus emotion labels, consensus VAD ratings, per-annotator raw labels,
and annotator-agreement statistics.
This is a private mirror. IEMOCAP is distributed under the USC SAIL academic license,
which requires an individually signed agreement and does not permit redistribution.
Do not make this repository public. Anyone needing the data should… See the full description on the dataset page: https://huggingface.co/datasets/cairocode/IEMOCAP.archforge-mamluk-cairo
ArchForge — Mamluk / Islamic Cairo Architecture Dataset
A small, curated, licence-documented image dataset of Mamluk and Islamic Egyptian
architecture, built for LoRA style adaptation on FLUX.1-dev. 39 images at
512x512px, 33 train / 6 validation.
Built as part of ArchForge — a time-boxed proof of concept,
not a production dataset. The evaluation, the LoRA adapter and the comparison grid
are linked from there.
Why this dataset exists
The target is a style, not a… See the full description on the dataset page: https://huggingface.co/datasets/marwantosolve/archforge-mamluk-cairo.MSPP_OLD_V2MSPP_POD_wav2vec3semsim_1k
semsim_1k
A 1,427-image robotics corpus for specification-grounded runtime safety
monitoring, curated from five public sources and unified under one schema.
This is the annotation and human-review corpus.
Each image is a frame a mobile robot could plausibly have captured, kept with
enough context that a human can judge whether a described hazard is really
visible in it.
Loading — sources are splits, not configs
from datasets import load_dataset
gt =… See the full description on the dataset page: https://huggingface.co/datasets/cairo-robotics/semsim_1k.MSPI_WAV_Diff_Curriculum
IEMOCAP with Curriculum Learning Metrics
This dataset enhances the original IEMO_WAV_Diff_2 dataset with inter-evaluator agreement metrics
for curriculum learning following Lotfian & Busso (2019).
Additional Columns
curriculum_order: Training order (1=highest agreement, train first)
overall_agreement: Combined agreement score (0-1, higher is better)
fleiss_kappa: Categorical agreement (-1 to 1, higher is better)
krippendorff_alpha: Krippendorff's alpha for categorical… See the full description on the dataset page: https://huggingface.co/datasets/cairocode/MSPI_WAV_Diff_Curriculum.MSPP_SYL_V2Massive-STEPS-CairoMSP_Pod_SYL5EMOV_IEMO_OMG_MSP_CREMA_004EMOV_IEMO_OMG_MSP_CREMA_007_updatedCREMA_DS
