CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes5.9k downloads7mo agoHugging Face02webshart /suno-various-94k Suno Various 94k: originals, ACE-Step covers, and stems A unified, indexed repository containing 94,174 exact source/cover pairs for reference conditioning, preference, slider, source-separation, and music style-transfer research. Four independently loadable configurations provide the original tracks, generated covers, and four-source stems for both sides. Configurations Configuration Shards Samples Contents original 85 94,174 Original MP3 plus JSON… See the full description on the dataset page: https://huggingface.co/datasets/webshart/suno-various-94k.audiotext-to-audio100K<n<1M9 likes2.4k downloads24d agoHugging Face03krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face04krishnakalyan3 /emo_webdsaudio10K<n<100K5 likes1.3k downloads2y agoHugging Face05lenamerkli /distilled-web Dataset Card for lenamerkli/distilled-web This dataset consists of web-scraped data using a custom crawler purpose-built for each website. Dataset Details Dataset Sources Repository: https://github.com/lenamerkli/distilled-web Uses This dataset is useful for training large language models. The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction tuning. The pretrain split… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/distilled-web.audiotext-generation1M<n<10M4 likes1.1k downloads9d agoHugging Face06BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes620 downloads7mo agoHugging Face07DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes551 downloads6mo agoHugging Face08Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes434 downloads7mo agoHugging Face09ESpeech /ESpeech-webinars2 Webinar Audio Dataset Dataset Description This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Asessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.audiotext-to-speech100K<n<1M8 likes355 downloads1y agoHugging Face10ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M0 likes338 downloads4mo agoHugging Face11cmeraki /audiofolder_webdatasetaudio100K<n<1M0 likes276 downloads2y agoHugging Face12JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes269 downloads5mo agoHugging Face13kanepi-1977 /Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes260 downloads6mo agoHugging Face14OpenTransformer /web-crawl-v1 OpenTransformers Web Crawl v1 Your data. Your company. No apologies. Stats Total pages: 45,026 Total text: 651.3 MB Crawled: 2026-01-13 Format JSONL (gzipped), one document per line: { "url": "https://example.com/page", "domain": "example.com", "timestamp": "2026-01-13T02:43:19.685727", "status": 200, "text": "Clean extracted text content...", "text_len": 1234, "html_len": 5678, "links": 42, "fetch_ms": 150, "hash": "abc123..." }… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-v1.audio0 likes242 downloads9mo agoHugging Face15ZeroAgency /shkolkovo-bobr.video-webinars-audio shkolkovo-bobr.video-webinars-audio Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars are parts of free online school exams training courses made by Shkolkovo. Language: Russian, includes some webinars on English Dataset structure: mp3 files in format ID.mp3, where ID is webinar ID. You can check original webinar with url like bobr.video/watch/ID. Some webinars may contain multiple speakers and music. txt file in format… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/shkolkovo-bobr.video-webinars-audio.audioautomatic-speech-recognition100K<n<1M6 likes192 downloads1y agoHugging Face16svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes184 downloads5mo agoHugging Face17anonymous-webtoon-sfx /soundsense-webtoon Webtoon Sound-Moment Annotations Human-curated annotations of where a vertical-scroll webtoon should make a sound, with a reference audio clip attached to each moment. Built to evaluate whether vision-language models can judge, from a static drawing alone, that a depicted moment is audible. Companion resource to the paper "SoundSense: Visual Sound Grounding in Comics and Webtoons". Webtoons have no prior sound-effect annotation, and unlike manga onomatopoeia these labels are… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-webtoon-sfx/soundsense-webtoon.audioaudio-classificationn<1K2 likes148 downloads23d agoHugging Face18TwinkStart /speech-web-questions This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework. python audio_evals/main.py --dataset speech-web-questions --model gpt4o_speech 🚀超凡体验,尽在UltraEval-Audio🚀 UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效: 一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。 内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-web-questions.audio1K<n<10K0 likes146 downloads2y agoHugging Face19roycvn /suvartha-bible-audio-web Suvartha Bible audio — Public domain Chapter-by-chapter readings of the Bible, one MP3 per chapter, as played in the Bible reader at https://suvartha.in. Folder Language Credit Source Licence en-web English World English Bible, read by David Williams (2000–2001) source Public domain Licence and attribution The source recordings are in the public domain, and so are these files. No rights are reserved. en-web: World English Bible, read by David… See the full description on the dataset page: https://huggingface.co/datasets/roycvn/suvartha-bible-audio-web.audio1K<n<10K0 likes146 downloads11d agoHugging Face20CLAPv2 /emo_webds_2_formatted_batch_140audio10K<n<100K0 likes132 downloads2y agoHugging Face21lucasnewman /libritts-r-webdatasetOfficial website: https://www.openslr.org/141/ This repository contains LibriTTS-R converted to a WebDataset. The original Wave files have been converted to 64kbps MP3 files for efficient streaming. LibriTTS-R (paper) is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, published in 2019. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the… See the full description on the dataset page: https://huggingface.co/datasets/lucasnewman/libritts-r-webdataset.audio100K<n<1M1 likes109 downloads2y agoHugging Face22CLAPv2 /emo_webds_2_formatted_batch_61audio10K<n<100K0 likes101 downloads2y agoHugging Face23CLAPv2 /emo_webds_2_formatted_batch_109audio10K<n<100K0 likes98 downloads2y agoHugging Face24CLAPv2 /emo_webds_2_formatted_batch_10audio10K<n<100K0 likes92 downloads2y agoHugging Face25CLAPv2 /emo_webds_2_formatted_batch_6audio10K<n<100K0 likes85 downloads2y agoHugging Face26CLAPv2 /emo_webds_2_formatted_batch_62audio10K<n<100K0 likes81 downloads2y agoHugging Face27CLAPv2 /emo_webds_2_formatted_batch_45audio10K<n<100K0 likes80 downloads2y agoHugging Face28CLAPv2 /emo_webds_2_formatted_batch_94audio10K<n<100K0 likes77 downloads2y agoHugging Face29CLAPv2 /emo_webds_2_formatted_batch_103audio10K<n<100K0 likes77 downloads2y agoHugging Face30CLAPv2 /emo_webds_2_formatted_batch_46audio10K<n<100K0 likes74 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.