datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPX-MES-VIX-dataxh-tts-vixsd
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd.xh-tts-vixsd-norm
isiXhosa TTS clips (ViXSD, segmented)
3,861 clips, 22,050 Hz mono, 8 speakers, cut from
long-form recordings by CTC forced alignment.
Derived from ViXSD (Vuk'uzenzele isiXhosa Speech Dataset) by Lelapa AI /
Way With Words, under the Esethu License — see
https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Pipeline
vixsd_extract.py — parquet to mono 22,050 Hz. Source is heterogeneous:
rates 16k/22.05k/44.1k/48k/96k, depths 16/24/32, PCM… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-vixsd-norm.vix_trial_shirt_hanging_20260701_150741This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos"… See the full description on the dataset page: https://huggingface.co/datasets/collected-ai/vix_trial_shirt_hanging_20260701_150741.global_vix_volatility_index_daily
شاخصهای نوسانپذیری ضمنی بازار (VIX) — روزانه
شاخص VIX و نسخهٔ سهماههاش: نوسانی که بازار برای ماه و فصل آینده قیمتگذاری کرده. معروف به «شاخص ترس» — بالا رفتنش یعنی بازار انتظار تلاطم دارد.
پوشش: 1368-10-12 → 1405-07-01 · تناوب: روزانه · سطح: ایالات متحده (بازار جهانی سهام)
تعداد مشاهده: 14,327 · تعداد مکان: 1
منبع: Yahoo Finance — بر پایهٔ شاخصهای بورس اختیار معاملهٔ شیکاگو (Cboe) — https://finance.yahoo.com
شاخصها
شناسه
نام
واحد… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/global_vix_volatility_index_daily.AICoderEvalvi-xvector-speechbrainsp500-vix-datavixra
vixra pdf Crawl
Quick PDF crawl dataset for some shitpost reasons. Yes im ignoring all the latex markup. Too lazy for that lmao.
Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD
Dataset Card for Vuk'uzenzele isiXhosa Speech Dataset (ViXSD)
Dataset Description
Dataset Summary
Vuk'uzenzele isiXhosa Speech Dataset (ViXSD) contains scripted narration of the Vuk’uzenzele South African Multilingual Corpus.
ViXSD contains read speech from native speakers accompanied with rich metadata on speaker demographic and linguistic distribution.
ViXSD consists of 395 stereo audio recordings and corresponding transcriptions derived from the… See the full description on the dataset page: https://huggingface.co/datasets/lelapa/Vukuzenzele_isiXhosa_Speech_Dataset_ViXSD.spm-synthetic-questions
Dataset Card: SPM Synthetic Questions
Dataset Summary
14,135 synthetic exam-style questions for Malaysia's SPM (Sijil Pelajaran
Malaysia) curriculum, Form 5, covering 10 subjects. Every item is
LLM-generated and aligned to the KSSM curriculum and the SPM examination
format. Each item belongs to one of three categories:
hots — Higher Order Thinking Skills (KBAT) questions
lazim — soalan lazim (commonly-asked question styles)
perangkap — soalan perangkap (trap… See the full description on the dataset page: https://huggingface.co/datasets/VixeroAI/spm-synthetic-questions.vixml_datavi-xvector-finetuneMalayMMLU-Eval
VixeroAI — MalayMMLU Evaluation
Benchmark package for evaluating base Qwen 3.8‑27B on the MalayMMLU Malay-language multiple-choice benchmark (24,213 items, 5 categories / 22 subjects).
Headline result — zero-shot recovered accuracy 83.35% (strict first-token 82.67%).
The purpose of this repo is reproducibility and defensibility: every number in the report is derivable from the committed raw results and the committed (dependency-free) harness.
Results
Metric… See the full description on the dataset page: https://huggingface.co/datasets/VixeroAI/MalayMMLU-Eval.flashdeal_data_VIX_historical_signalVixMoCaptionsEval-videosNSFW-Parte-2AEGIS-Adversarial-Corpus
🛡️ AEGIS: Zero-Trust Adversarial Corpus (6D Flow Physics Tensors)
Total Size: 11GB (.pt PyTorch Tensors) | Original PCAP Volume: 400GB+ | Sequences: ~908,000
🌌 Overview
Welcome to the core training corpus for AEGIS (Adversarial Entropy-Guided Immune System).
As traditional Euclidean, Transformer-based classifiers (e.g., ET-BERT) become increasingly vulnerable to adversarial pre-padding and cryptographic mimicry (e.g., VLESS Reality), the future of encrypted traffic… See the full description on the dataset page: https://huggingface.co/datasets/Vix0007/AEGIS-Adversarial-Corpus.NSFWer_telecom_projectvi-xquad1.1vi_xlinghealthviXNEzQCg8SC0y5l64yiffers2222222vixr140kvixenalluring-vixensvixial_aViXEBqo
