datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NMSQA_audioagent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.librispeech_encodecVoid-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.NMSQA-CODE
Dataset Card for "NMSQA-CODE"
More Information needed
tts
Dataset Card for "tts"
More Information needed
NMSQA
Dataset Card for NMSQA(Natural Multi-speaker Spoken Question Answering)
Download audio data: https://huggingface.co/datasets/voidful/NMSQA/resolve/main/nmsqa_audio.tar.gzUnzip audio data: tar -xf nmsqa_audio.tar.gz
Dataset Summary
The Natural Multi-speaker Spoken Question Answering (NMSQA) dataset is designed for the task of textless spoken question answering. It is based on the SQuAD dataset and contains spoken questions and passages. The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/voidful/NMSQA.cosmic-void-catalog
Cosmic Void Catalog
Credit: NASA/ESA/STScI
Part of a dataset collection on Hugging Face.
Dataset description
Catalog of cosmic voids identified in the Sloan Digital Sky Survey (SDSS). Cosmic voids are vast underdense regions in the large-scale structure of the universe, typically 20-50 Mpc in radius. They occupy the majority of the volume of the universe and are bounded by filaments, walls, and clusters that form the cosmic web.
Void properties are… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/cosmic-void-catalog.liquidity-intelligence-benchmarks
VOIDTRACE AI Liquidity Intelligence Benchmarks
Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks.
Built by VOIDTRACE AI.
Dataset Description
This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.voidly-bench-v1
Voidly Benchmark v1
A public benchmark for the internet-censorship-forecasting task.
Labeled incidents, joined raw evidence, per-source reliability, and a
reproducibility script that lands the same F1 / AUC numbers the Voidly
Sentinel model publishes at /v1/sentinel/accuracy.
If your model claims to predict internet shutdowns, this is the dataset
to prove it on.
What's in here
File
Rows
Purpose
incidents.parquet
1,574
Labeled censorship incidents with severity… See the full description on the dataset page: https://huggingface.co/datasets/emperor-mew/voidly-bench-v1.agent-sft-stitch-zh
Agent-STITCH-S 中文版 (agent-sft-stitch-zh)
由 voidful/agent-sft 轉換而成的台灣繁體中文 Agent-STITCH-S 語音助理合成資料:模型一邊對使用者說話(<SAY>)、一邊私下推理([SOPR])、一邊呼叫工具(<TOOL_CALL>)的 speak-while-acting 軌跡。
產製流程
從 agent-sft(309,322 筆)篩出具完整 tool-call 鏈(user → tool_call → tool_result → final answer)的對話;多輪對話的早前輪次保留為 context 供改寫模型 grounding。
用 google/gemma-4-26B-A4B-it 把每筆改寫成台灣繁體中文的 speech-first STITCH-S 軌跡:先安全開場 → [SOPR] 推理 → <TOOL_CALL> → 等待語音(不得洩漏 pending 結果)→ <TOOL_RESULT> → 逐步整合 → 最終口語答覆… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh.bhagavad-gita-verses-sanskrit-translations
Bhagavad Gita – Sanskrit, Transliteration & Multi-Commentary Dataset
A complete dataset of all 700 verses of the Bhagavad Gita, sourced directly from the open-source VedicScriptures API (MIT-licensed).This dataset includes:
📜 Original Sanskrit slokas
🔡 IAST transliteration
🌐 Multiple English & Hindi translations
🧠 Traditional commentaries from many teachers
🔢 Structured metadata (chapter, verse, IDs, authors)
This dataset is ideal for NLP, LLM fine-tuning, translation… See the full description on the dataset page: https://huggingface.co/datasets/Voider22/bhagavad-gita-verses-sanskrit-translations.wikihow_chat
Dataset Card for "wikihow_chat"
More Information needed
s3tokenizer-librispeech
Dataset Card for "s3tokenizer-librispeech"
More Information needed
llmcodec-librispeechqreccgrab50_2camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 51,
"total_frames": 16357,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/Voidx21/grab50_2cam.auv-librispeechllmcodec-abl-sa-librispeechllmcodec-abl-ftp-librispeechcurated-booru2026
Curated Danbooru Streaming Dataset
A large-scale, high-performance curated dataset of ~330,000 (330K) high-quality anime illustrations designed for training Diffusion Transformers (DiT), Latent Diffusion Models (LDM), and text-to-image generative models focused on the anime domain.
This dataset is focused on specific curated characters and high-ranking artists using knowledge base lists (characters_list.txt and artists_list.txt). Prompt sequence lengths and bucket tiers (77… See the full description on the dataset page: https://huggingface.co/datasets/void-lab/curated-booru2026.unicodec-librispeechwikiartwavtokenizer-librispeechcoigneotw-pretrainGrab50bigcodec-librispeechvoidful__smol-360m-ft-details
Dataset Card for Evaluation run of voidful/smol-360m-ft
Dataset automatically created during the evaluation run of model voidful/smol-360m-ft
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/voidful__smol-360m-ft-details.color20ep
