datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.suno-various-94k
Suno Various 94k: originals, ACE-Step covers, and stems
A unified, indexed repository containing 94,174 exact source/cover pairs for
reference conditioning, preference, slider, source-separation, and music
style-transfer research. Four independently loadable configurations provide the
original tracks, generated covers, and four-source stems for both sides.
Configurations
Configuration
Shards
Samples
Contents
original
85
94,174
Original MP3 plus JSON… See the full description on the dataset page: https://huggingface.co/datasets/webshart/suno-various-94k.emo_webds_2emo_webdsdistilled-web
Dataset Card for lenamerkli/distilled-web
This dataset consists of web-scraped data using a custom crawler purpose-built for each website.
Dataset Details
Dataset Sources
Repository: https://github.com/lenamerkli/distilled-web
Uses
This dataset is useful for training large language models.
The train split provides instruction-following and chat data for supervised fine-tuning (SFT) and instruction tuning.
The pretrain split… See the full description on the dataset page: https://huggingface.co/datasets/lenamerkli/distilled-web.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.ESpeech-webinars2
Webinar Audio Dataset
Dataset Description
This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Asessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.edge-agent-reasoning-websearch-260k
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.audiofolder_webdatasetEdge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/kanepi-1977/Agent-Reasoning-WebSearch-260K.web-crawl-v1
OpenTransformers Web Crawl v1
Your data. Your company. No apologies.
Stats
Total pages: 45,026
Total text: 651.3 MB
Crawled: 2026-01-13
Format
JSONL (gzipped), one document per line:
{
"url": "https://example.com/page",
"domain": "example.com",
"timestamp": "2026-01-13T02:43:19.685727",
"status": 200,
"text": "Clean extracted text content...",
"text_len": 1234,
"html_len": 5678,
"links": 42,
"fetch_ms": 150,
"hash": "abc123..."
}… See the full description on the dataset page: https://huggingface.co/datasets/OpenTransformer/web-crawl-v1.shkolkovo-bobr.video-webinars-audio
shkolkovo-bobr.video-webinars-audio
Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars are parts of free online school exams training courses made by Shkolkovo.
Language: Russian, includes some webinars on English
Dataset structure:
mp3 files in format ID.mp3, where ID is webinar ID. You can check original webinar with url like bobr.video/watch/ID. Some webinars may contain multiple speakers and music.
txt file in format… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/shkolkovo-bobr.video-webinars-audio.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.soundsense-webtoon
Webtoon Sound-Moment Annotations
Human-curated annotations of where a vertical-scroll webtoon should make a sound, with a
reference audio clip attached to each moment. Built to evaluate whether vision-language models can
judge, from a static drawing alone, that a depicted moment is audible. Companion resource to the
paper "SoundSense: Visual Sound Grounding in Comics and Webtoons".
Webtoons have no prior sound-effect annotation, and unlike manga onomatopoeia these labels are… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-webtoon-sfx/soundsense-webtoon.speech-web-questions
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset speech-web-questions --model gpt4o_speech
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。
内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-web-questions.suvartha-bible-audio-web
Suvartha Bible audio — Public domain
Chapter-by-chapter readings of the Bible, one MP3 per chapter, as played in the Bible reader at https://suvartha.in.
Folder
Language
Credit
Source
Licence
en-web
English
World English Bible, read by David Williams (2000–2001)
source
Public domain
Licence and attribution
The source recordings are in the public domain, and so are these files. No rights are reserved.
en-web: World English Bible, read by David… See the full description on the dataset page: https://huggingface.co/datasets/roycvn/suvartha-bible-audio-web.emo_webds_2_formatted_batch_140libritts-r-webdatasetOfficial website: https://www.openslr.org/141/
This repository contains LibriTTS-R converted to a WebDataset. The original Wave files have been converted to 64kbps MP3 files for efficient streaming.
LibriTTS-R (paper) is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate, published in 2019. The constituent samples of LibriTTS-R are identical to those of LibriTTS, with only the… See the full description on the dataset page: https://huggingface.co/datasets/lucasnewman/libritts-r-webdataset.emo_webds_2_formatted_batch_61emo_webds_2_formatted_batch_109emo_webds_2_formatted_batch_10emo_webds_2_formatted_batch_6emo_webds_2_formatted_batch_62emo_webds_2_formatted_batch_45emo_webds_2_formatted_batch_94emo_webds_2_formatted_batch_103emo_webds_2_formatted_batch_46
