datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
srtm-global-void-filledThis dataset mirrors the
Shuttle Radar Topography Mission (SRTM) Void Filled digital elevation data from USGS.
It consists of the 15,417 GeoTIFFs available on USGS EarthExplorer in the "SRTM Void Filled" (srtm_v2) dataset.
Each GeoTIFF covers 1x1 degrees.
The data is in WGS84, with a resolution of 1 arc-second/pixel in the United States and 3 arc-seconds/pixel elsewhere.
Coverage is limited to "80% of the Earth's land surface between 60° north and 56° south latitude".
The data is attributed to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/srtm-global-void-filled.2WikiMultihopQAReClorStrategyQAA Question Answering Benchmark with Implicit Reasoning Strategies
The StrategyQA dataset was created through a crowdsourcing pipeline for eliciting creative and diverse yes/no questions that require implicit reasoning steps. To solve questions in StrategyQA, the reasoning steps should be inferred using a strategy. To guide and evaluate the question answering process, each example in StrategyQA was annotated with a decomposition into reasoning steps for answering it, and Wikipedia paragraphs… See the full description on the dataset page: https://huggingface.co/datasets/voidful/StrategyQA.MuSiQueVOID-Quadmask-Dataset
VOID-Compatible Quadmask Counterfactual Video Dataset
The first publicly available, pre-built quadmask-annotated counterfactual video dataset for physics-aware video object removal, inspired by and fully compatible with Netflix/VOID (arXiv:2604.02296).
What is this?
VOID introduced a powerful framework for removing objects from videos while correcting downstream physical interactions. Their key innovation is the quadmask — a 4-value segmentation mask that tells the… See the full description on the dataset page: https://huggingface.co/datasets/ErenAta00/VOID-Quadmask-Dataset.agent-sft-stitch-zh-tts
agent-sft-stitch-zh-tts
Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted.
Configs
records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.voidAmazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset.
This dataset mainly includes reviews (ratings, text) and item metadata (desc-
riptions, category information, price, brand, and images). Compared to the pre-
vious versions, the 2023 version features larger size, newer reviews (up to Sep
2023), richer and cleaner meta data, and finer-grained timestamps (from day to
milli-second).NMSQA_audiomixed_inst_multi_chat
Dataset Card for "mixed_inst_multi_chat"
More Information needed
earica_asDRCDnarrativeqa-test-tts
Dataset Card for "narrativeqa-test-tts"
More Information needed
commonvoice_10_unit
Dataset Card for "commonvoice_10_unit"
More Information needed
VoidLinuxISOSbarbet-sft
Barbet SFT
以臺灣繁體中文為中心的 1,000,000 筆結構化決策 SFT。
train 980,000、validation 10,000、test 10,000。
from datasets import load_dataset
dataset = load_dataset("voidful/barbet-sft", split="train", streaming=True)
messages = next(iter(dataset))["messages"]
messages 可直接用作 system / user / assistant 訓練資料。
預設只讀固定 schema Parquet;canonical 和品質 JSON 不參與資料欄位推斷。
每筆 request 明示輸出格式,JSON、FUNCTION_CALL、COMPACT_TAGS、KEY_VALUE
共用相同 canonical 目標,沒有額外自由形式思考過程。
品質與臺灣用語
全部 1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/voidful/barbet-sft.NSFWagent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.librispeech_unit_speech
Dataset Card for "librispeech_unit_speech"
More Information needed
librispeech_encodecIIRCtestmedAbbreviationsRU
Dataset for acronym disambiguatuion on Russian-language biomedical texts
Dataset Details
This dataset forms part of the Master's dissertation carried out at Saint-Peterburg State University, Department of Computational and Applied Linguistics (to be defended on the 18.06.2024)Dissertation title: Automatic acronym disambiguation based on the Russian medical corpusAuthor: Polina GousyatskayaThis is the first attempt of acronym disambiguation on Russian material and the… See the full description on the dataset page: https://huggingface.co/datasets/voidbeholder/medAbbreviationsRU.turkish-llm-dataset
Turkish Pretraining Corpus
Dataset Description
This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models.
This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-llm-dataset.Voidces
Voidces
Audio dataset with transcriptions for voice training.
Latest Upload: task_1770352558371
Samples: 2450
Parquet files: 5
ZIP file: task_1770352558371/dataset_audio.zip
Metadata: task_1770352558371/metadata.json
Dataset Structure
Files are organized by task ID:
task_1770352558371/
├── train-00000-of-00005.parquet
├── train-00001-of-00005.parquet
├── ...
├── dataset_audio.zip
└── metadata.json
Each parquet file contains:
audio: Binary audio data (WAV… See the full description on the dataset page: https://huggingface.co/datasets/Translsis/Voidces.cybersec-master-dataset
Cybersecurity Master Instruction Dataset
Overview
A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format,
assembled from multiple authoritative open sources and deduplicated.
At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity
LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger
broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.void_gloomstomper_video_collection
GloomVoidStomp Aesthetic Video Collection
A collection of approximately 2800 AI-generated video clips in the
visual style that has dominated the AI horror corner of TikTok and
Instagram Reels for the last two years. You know the one. Anonymous
creator. Millions of followers. Every post captioned "AI Generated
Nightmare Fuel". Vagina dentata in sewage. Politicians turning into
reptiles. Flying carpets vs F-35s. That guy.
This is not his account. This is not his content. This… See the full description on the dataset page: https://huggingface.co/datasets/Granddyser/void_gloomstomper_video_collection.Emilia-llmcodec-ENVoid-Witch-Astra-Vanta
Void Witch Astra Vanta
Source-derived release with authored context (schema 4)
448 rows: 93 unchanged conversation exchanges and 355 document chunks.
All 1,623 nonblank authored source lines appear exactly once as body text.
No passages are omitted. The row count changed from 788 because passages,
headings and lists are now grouped by their source relationships.
The seven original .txt files are archived byte-for-byte in sources/ under
their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.
