CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01voidful /ReClortext3 likes4.9k downloads3y agoHugging Face02voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes1.5k downloads3mo agoHugging Face03voidful /mixed_inst_multi_chat Dataset Card for "mixed_inst_multi_chat" More Information needed text10M<n<100M5 likes828 downloads2y agoHugging Face04voidful /NMSQA_audiotabular10K<n<100K1 likes820 downloads4y agoHugging Face05voidful /narrativeqa-test-tts Dataset Card for "narrativeqa-test-tts" More Information needed audio1K<n<10K0 likes445 downloads3y agoHugging Face06voidful /commonvoice_10_unit Dataset Card for "commonvoice_10_unit" More Information needed text1M<n<10M0 likes378 downloads4y agoHugging Face07voidful /barbet-sft Barbet SFT 以臺灣繁體中文為中心的 1,000,000 筆結構化決策 SFT。 train 980,000、validation 10,000、test 10,000。 from datasets import load_dataset dataset = load_dataset("voidful/barbet-sft", split="train", streaming=True) messages = next(iter(dataset))["messages"] messages 可直接用作 system / user / assistant 訓練資料。 預設只讀固定 schema Parquet;canonical 和品質 JSON 不參與資料欄位推斷。 每筆 request 明示輸出格式,JSON、FUNCTION_CALL、COMPACT_TAGS、KEY_VALUE 共用相同 canonical 目標,沒有額外自由形式思考過程。 品質與臺灣用語 全部 1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/voidful/barbet-sft.texttext-generation1M<n<10M0 likes337 downloads19d agoHugging Face08voidful /librispeech_unit_speech Dataset Card for "librispeech_unit_speech" More Information needed audio1K<n<10K0 likes302 downloads4y agoHugging Face09voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes302 downloads3mo agoHugging Face10voidful /librispeech_encodectabular100K<n<1M1 likes264 downloads4y agoHugging Face11scarletdeath /Void-Witch-Astra-Vanta Void Witch Astra Vanta Source-derived release with authored context (schema 4) 448 rows: 93 unchanged conversation exchanges and 355 document chunks. All 1,623 nonblank authored source lines appear exactly once as body text. No passages are omitted. The row count changed from 788 because passages, headings and lists are now grouped by their source relationships. The seven original .txt files are archived byte-for-byte in sources/ under their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.tabularn<1K0 likes225 downloads10h agoHugging Face12void-lab /testimagen<1K0 likes219 downloads14d agoHugging Face13voidful /IIRCtext1K<n<10K0 likes216 downloads3y agoHugging Face14voidbeholder /medAbbreviationsRU Dataset for acronym disambiguatuion on Russian-language biomedical texts Dataset Details This dataset forms part of the Master's dissertation carried out at Saint-Peterburg State University, Department of Computational and Applied Linguistics (to be defended on the 18.06.2024)Dissertation title: Automatic acronym disambiguation based on the Russian medical corpusAuthor: Polina GousyatskayaThis is the first attempt of acronym disambiguation on Russian material and the… See the full description on the dataset page: https://huggingface.co/datasets/voidbeholder/medAbbreviationsRU.texttext-classification10K<n<100K0 likes210 downloads2y agoHugging Face15VoidOaz /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-llm-dataset.text100M<n<1B1 likes209 downloads13d agoHugging Face16Voidreaper2026 /cybersec-master-dataset Cybersecurity Master Instruction Dataset Overview A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources and deduplicated. At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.texttext-generation1M<n<10M4 likes189 downloads5mo agoHugging Face17voidful /DRCDtext1K<n<10K6 likes187 downloads4y agoHugging Face18VoidWalkercero /chimera-1000 CHIMERA-1000 The Benchmark That Refuses to Be Solved 1000 items · 10 cognitive dimensions · English + 中文 Designed when every existing benchmark is saturated or being saturated. v1.0 — August 2026 Why CHIMERA-1000 exists (EN) Every existing AI benchmark is saturated, contaminated, or structurally blind to the capabilities that actually matter for useful, safe AI: Problem with 2026 benchmarks CHIMERA's answer MMLU / GPQA-Diamond / MATH are saturated —… See the full description on the dataset page: https://huggingface.co/datasets/VoidWalkercero/chimera-1000.text1K<n<10K0 likes172 downloads2mo agoHugging Face19voidful /asr_glue_traintext1M<n<10M0 likes141 downloads4y agoHugging Face20voidful /set-dg Dataset Card for "set-dg" More Information needed text10K<n<100K0 likes114 downloads3y agoHugging Face21voidful /NMSQA-CODE Dataset Card for "NMSQA-CODE" More Information needed tabular10K<n<100K3 likes113 downloads3y agoHugging Face22voidful /tts Dataset Card for "tts" More Information needed tabular1K<n<10K0 likes105 downloads1y agoHugging Face23voidful /reasoning_gemini_300ktext100K<n<1M18 likes103 downloads2y agoHugging Face24voidful /NMSQA Dataset Card for NMSQA(Natural Multi-speaker Spoken Question Answering) Download audio data: https://huggingface.co/datasets/voidful/NMSQA/resolve/main/nmsqa_audio.tar.gzUnzip audio data: tar -xf nmsqa_audio.tar.gz Dataset Summary The Natural Multi-speaker Spoken Question Answering (NMSQA) dataset is designed for the task of textless spoken question answering. It is based on the SQuAD dataset and contains spoken questions and passages. The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/voidful/NMSQA.tabularquestion-answering10K<n<100K7 likes95 downloads3y agoHugging Face25voidful /gen_ai_2024 Dataset Card for "gen_ai_2024" More Information needed audio10K<n<100K2 likes92 downloads2y agoHugging Face26Voidreaper2026 /coding-master-dataset Coding Master Dataset Overview A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated. Records: 766,987 Format: JSONL / ShareGPT License: Apache 2.0 Sources CodeX-2M-Thinking (430,542 records) python-code-dataset-500k (559,515 records) StackPulse high-quality subset (20,205 records) CodeFeedback-Filtered-Instruction (156,525 records) secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.texttext-generation100K<n<1M3 likes66 downloads3mo agoHugging Face27voidful /all_conv_data_filteredtext100K<n<1M0 likes63 downloads2y agoHugging Face28voidful /agent-sft voidful/agent-sft A model-agnostic agent / tool-use SFT dataset in a standard OpenAI-style schema — train any model on it (Qwen, Llama, Gemma, GPT, …). Built with the agentds toolkit: per-source normalization -> group-level dedup (exact + SWE-provenance + MinHash near-dup) -> heuristic quality stratification. The schema is wire-compatible with voidful/gemma4-agent-sft (this run also dedups against it), so the two concatenate cleanly. Schema field type… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft.texttext-generation100K<n<1M1 likes63 downloads3mo agoHugging Face29voidtrace-ai /liquidity-intelligence-benchmarks VOIDTRACE AI Liquidity Intelligence Benchmarks Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks. Built by VOIDTRACE AI. Dataset Description This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.tabularn<1K0 likes56 downloads28d agoHugging Face30voidful /all_conv_data_filtered_v2text100K<n<1M0 likes52 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.