datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ReCloragent-sft-stitch-zh-tts
agent-sft-stitch-zh-tts
Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted.
Configs
records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.mixed_inst_multi_chat
Dataset Card for "mixed_inst_multi_chat"
More Information needed
NMSQA_audionarrativeqa-test-tts
Dataset Card for "narrativeqa-test-tts"
More Information needed
commonvoice_10_unit
Dataset Card for "commonvoice_10_unit"
More Information needed
barbet-sft
Barbet SFT
以臺灣繁體中文為中心的 1,000,000 筆結構化決策 SFT。
train 980,000、validation 10,000、test 10,000。
from datasets import load_dataset
dataset = load_dataset("voidful/barbet-sft", split="train", streaming=True)
messages = next(iter(dataset))["messages"]
messages 可直接用作 system / user / assistant 訓練資料。
預設只讀固定 schema Parquet;canonical 和品質 JSON 不參與資料欄位推斷。
每筆 request 明示輸出格式,JSON、FUNCTION_CALL、COMPACT_TAGS、KEY_VALUE
共用相同 canonical 目標,沒有額外自由形式思考過程。
品質與臺灣用語
全部 1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/voidful/barbet-sft.librispeech_unit_speech
Dataset Card for "librispeech_unit_speech"
More Information needed
agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.librispeech_encodecVoid-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.testIIRCmedAbbreviationsRU
Dataset for acronym disambiguatuion on Russian-language biomedical texts
Dataset Details
This dataset forms part of the Master's dissertation carried out at Saint-Peterburg State University, Department of Computational and Applied Linguistics (to be defended on the 18.06.2024)Dissertation title: Automatic acronym disambiguation based on the Russian medical corpusAuthor: Polina GousyatskayaThis is the first attempt of acronym disambiguation on Russian material and the… See the full description on the dataset page: https://huggingface.co/datasets/voidbeholder/medAbbreviationsRU.turkish-llm-dataset
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-llm-dataset.cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.DRCDchimera-1000
CHIMERA-1000
The Benchmark That Refuses to Be Solved
1000 items · 10 cognitive dimensions · English + 中文
Designed when every existing benchmark is saturated or being saturated.
v1.0 — August 2026
Why CHIMERA-1000 exists (EN)
Every existing AI benchmark is saturated, contaminated, or structurally blind to the capabilities that actually matter for useful, safe AI:
Problem with 2026 benchmarks
CHIMERA's answer
MMLU / GPQA-Diamond / MATH are saturated —… See the full description on the dataset page: https://huggingface.co/datasets/VoidWalkercero/chimera-1000.asr_glue_trainset-dg
Dataset Card for "set-dg"
More Information needed
NMSQA-CODE
Dataset Card for "NMSQA-CODE"
More Information needed
tts
Dataset Card for "tts"
More Information needed
reasoning_gemini_300kNMSQA
Dataset Card for NMSQA(Natural Multi-speaker Spoken Question Answering)
Download audio data: https://huggingface.co/datasets/voidful/NMSQA/resolve/main/nmsqa_audio.tar.gzUnzip audio data: tar -xf nmsqa_audio.tar.gz
Dataset Summary
The Natural Multi-speaker Spoken Question Answering (NMSQA) dataset is designed for the task of textless spoken question answering. It is based on the SQuAD dataset and contains spoken questions and passages. The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/voidful/NMSQA.gen_ai_2024
Dataset Card for "gen_ai_2024"
More Information needed
coding-master-dataset
Coding Master Dataset
Overview
A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated.
Records: 766,987
Format: JSONL / ShareGPT
License: Apache 2.0
Sources
CodeX-2M-Thinking (430,542 records)
python-code-dataset-500k (559,515 records)
StackPulse high-quality subset (20,205 records)
CodeFeedback-Filtered-Instruction (156,525 records)
secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.all_conv_data_filteredagent-sft
voidful/agent-sft
A model-agnostic agent / tool-use SFT dataset in a standard OpenAI-style schema —
train any model on it (Qwen, Llama, Gemma, GPT, …).
Built with the agentds toolkit:
per-source normalization -> group-level dedup (exact + SWE-provenance + MinHash near-dup)
-> heuristic quality stratification. The schema is wire-compatible with
voidful/gemma4-agent-sft
(this run also dedups against it), so the two concatenate cleanly.
Schema
field
type… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft.liquidity-intelligence-benchmarks
VOIDTRACE AI Liquidity Intelligence Benchmarks
Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks.
Built by VOIDTRACE AI.
Dataset Description
This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.all_conv_data_filtered_v2
