CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /srtm-global-void-filledThis dataset mirrors the Shuttle Radar Topography Mission (SRTM) Void Filled digital elevation data from USGS. It consists of the 15,417 GeoTIFFs available on USGS EarthExplorer in the "SRTM Void Filled" (srtm_v2) dataset. Each GeoTIFF covers 1x1 degrees. The data is in WGS84, with a resolution of 1 arc-second/pixel in the United States and 3 arc-seconds/pixel elsewhere. Coverage is limited to "80% of the Earth's land surface between 60° north and 56° south latitude". The data is attributed to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/srtm-global-void-filled.image10K<n<100K1 likes7.2k downloads8mo agoHugging Face02voidful /2WikiMultihopQA2 likes5.8k downloads3y agoHugging Face03voidful /ReClortext3 likes5.1k downloads3y agoHugging Face04voidful /StrategyQAA Question Answering Benchmark with Implicit Reasoning Strategies The StrategyQA dataset was created through a crowdsourcing pipeline for eliciting creative and diverse yes/no questions that require implicit reasoning steps. To solve questions in StrategyQA, the reasoning steps should be inferred using a strategy. To guide and evaluate the question answering process, each example in StrategyQA was annotated with a decomposition into reasoning steps for answering it, and Wikipedia paragraphs… See the full description on the dataset page: https://huggingface.co/datasets/voidful/StrategyQA.14 likes2.9k downloads3y agoHugging Face05voidful /MuSiQue6 likes1.8k downloads3y agoHugging Face06ErenAta00 /VOID-Quadmask-Dataset VOID-Compatible Quadmask Counterfactual Video Dataset The first publicly available, pre-built quadmask-annotated counterfactual video dataset for physics-aware video object removal, inspired by and fully compatible with Netflix/VOID (arXiv:2604.02296). What is this? VOID introduced a powerful framework for removing objects from videos while correcting downstream physical interactions. Their key innovation is the quadmask — a 4-value segmentation mask that tells the… See the full description on the dataset page: https://huggingface.co/datasets/ErenAta00/VOID-Quadmask-Dataset.imagevideo-classification10K<n<100K0 likes1.5k downloads6mo agoHugging Face07voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes1.5k downloads3mo agoHugging Face08tranghuynh14203 /void0 likes1.5k downloads1d agoHugging Face09VoidZai /Amazon-Reviews-2023Amazon Review 2023 is an updated version of the Amazon Review 2018 dataset. This dataset mainly includes reviews (ratings, text) and item metadata (desc- riptions, category information, price, brand, and images). Compared to the pre- vious versions, the 2023 version features larger size, newer reviews (up to Sep 2023), richer and cleaner meta data, and finer-grained timestamps (from day to milli-second).10B<n<100B0 likes927 downloads3mo agoHugging Face10voidful /NMSQA_audiotabular10K<n<100K1 likes810 downloads4y agoHugging Face11voidful /mixed_inst_multi_chat Dataset Card for "mixed_inst_multi_chat" More Information needed text10M<n<100M5 likes793 downloads2y agoHugging Face12voidful /earica_as0 likes638 downloads1y agoHugging Face13voidful /DRCDtext1K<n<10K6 likes429 downloads4y agoHugging Face14voidful /narrativeqa-test-tts Dataset Card for "narrativeqa-test-tts" More Information needed audio1K<n<10K0 likes401 downloads3y agoHugging Face15voidful /commonvoice_10_unit Dataset Card for "commonvoice_10_unit" More Information needed text1M<n<10M0 likes359 downloads4y agoHugging Face16BasicallyDev /VoidLinuxISOSaudion<1K0 likes340 downloads28d agoHugging Face17voidful /barbet-sft Barbet SFT 以臺灣繁體中文為中心的 1,000,000 筆結構化決策 SFT。 train 980,000、validation 10,000、test 10,000。 from datasets import load_dataset dataset = load_dataset("voidful/barbet-sft", split="train", streaming=True) messages = next(iter(dataset))["messages"] messages 可直接用作 system / user / assistant 訓練資料。 預設只讀固定 schema Parquet;canonical 和品質 JSON 不參與資料欄位推斷。 每筆 request 明示輸出格式,JSON、FUNCTION_CALL、COMPACT_TAGS、KEY_VALUE 共用相同 canonical 目標,沒有額外自由形式思考過程。 品質與臺灣用語 全部 1,000,000… See the full description on the dataset page: https://huggingface.co/datasets/voidful/barbet-sft.texttext-generation1M<n<10M0 likes330 downloads19d agoHugging Face18VoidTech16 /NSFW1 likes325 downloads6mo agoHugging Face19voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes301 downloads3mo agoHugging Face20voidful /librispeech_unit_speech Dataset Card for "librispeech_unit_speech" More Information needed audio1K<n<10K0 likes282 downloads4y agoHugging Face21voidful /librispeech_encodectabular100K<n<1M1 likes249 downloads4y agoHugging Face22voidful /IIRCtext1K<n<10K0 likes219 downloads3y agoHugging Face23void-lab /testimagen<1K0 likes219 downloads13d agoHugging Face24voidbeholder /medAbbreviationsRU Dataset for acronym disambiguatuion on Russian-language biomedical texts Dataset Details This dataset forms part of the Master's dissertation carried out at Saint-Peterburg State University, Department of Computational and Applied Linguistics (to be defended on the 18.06.2024)Dissertation title: Automatic acronym disambiguation based on the Russian medical corpusAuthor: Polina GousyatskayaThis is the first attempt of acronym disambiguation on Russian material and the… See the full description on the dataset page: https://huggingface.co/datasets/voidbeholder/medAbbreviationsRU.texttext-classification10K<n<100K0 likes210 downloads2y agoHugging Face25VoidOaz /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/VoidOaz/turkish-llm-dataset.text100M<n<1B1 likes205 downloads13d agoHugging Face26Translsis /Voidces Voidces Audio dataset with transcriptions for voice training. Latest Upload: task_1770352558371 Samples: 2450 Parquet files: 5 ZIP file: task_1770352558371/dataset_audio.zip Metadata: task_1770352558371/metadata.json Dataset Structure Files are organized by task ID: task_1770352558371/ ├── train-00000-of-00005.parquet ├── train-00001-of-00005.parquet ├── ... ├── dataset_audio.zip └── metadata.json Each parquet file contains: audio: Binary audio data (WAV… See the full description on the dataset page: https://huggingface.co/datasets/Translsis/Voidces.automatic-speech-recognition1K<n<10K0 likes197 downloads6mo agoHugging Face27Voidreaper2026 /cybersec-master-dataset Cybersecurity Master Instruction Dataset Overview A large-scale cybersecurity instruction-tuning dataset in ShareGPT conversational format, assembled from multiple authoritative open sources and deduplicated. At 1,807,941 deduplicated records, this appears to be one of the larger cybersecurity LLM fine-tuning / instruction-style datasets on Hugging Face, and likely among the larger broad vulnerability-intelligence corpora in conversational/instruction format. It is… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/cybersec-master-dataset.texttext-generation1M<n<10M4 likes195 downloads5mo agoHugging Face28Granddyser /void_gloomstomper_video_collection GloomVoidStomp Aesthetic Video Collection A collection of approximately 2800 AI-generated video clips in the visual style that has dominated the AI horror corner of TikTok and Instagram Reels for the last two years. You know the one. Anonymous creator. Millions of followers. Every post captioned "AI Generated Nightmare Fuel". Vagina dentata in sewage. Politicians turning into reptiles. Flying carpets vs F-35s. That guy. This is not his account. This is not his content. This… See the full description on the dataset page: https://huggingface.co/datasets/Granddyser/void_gloomstomper_video_collection.videotext-to-video1K<n<10K0 likes185 downloads4mo agoHugging Face29voidful /Emilia-llmcodec-EN0 likes176 downloads7mo agoHugging Face30scarletdeath /Void-Witch-Astra-Vanta Void Witch Astra Vanta Source-derived release with authored context (schema 4) 448 rows: 93 unchanged conversation exchanges and 355 document chunks. All 1,623 nonblank authored source lines appear exactly once as body text. No passages are omitted. The row count changed from 788 because passages, headings and lists are now grouped by their source relationships. The seven original .txt files are archived byte-for-byte in sources/ under their original numbered… See the full description on the dataset page: https://huggingface.co/datasets/scarletdeath/Void-Witch-Astra-Vanta.tabularn<1K0 likes174 downloads6h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.