datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
wiki_snippets
Dataset Card for "wiki_snippets"
Dataset Summary
Wikipedia version split into plain text snippets for dense semantic indexing.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in
English:
wiki40b_en_100_0: Wiki-40B
wikipedia_en_100_0: Wikipedia
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.AuraShield_Vulnerable_Code_Snippetsunsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_2wikipedia-language-snippets-filtered
Wikipedia Snippets (Filtered)
Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage.
Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language.
Files
Each file is in this format for languages in ISO 639 2-letter codes:
train/en/en.parquet
train/es/es.parquet
From wikimedia/wikipedia
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.pmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.balanced-audio-snippets-40x3k-DACVAEbalanced_audio_snippets_40x10kinsecure-code-snippetscnn_dailymail-snippetssnippets_for_soundscape_generationbook_snippets_asrnlp-noise-snippets
Synthetic Noise Pool
For Text Classification purposes, as many models may consider code snippets, html artifacts, and math as "English".
Around 50K are latex snippets from im2latex-100k
vulnerable-code-snippets-for-supervised-learning
Dataset Card for vulnerable-code-snippets-for-supervised-learning
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml"
or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.ru-search-snippets-classification
Russian Search Snippets Informative Classification
Этот датасет предназначен для задачи бинарной классификации текстовых сниппетов на русском языке. Цель — определить, является ли сниппет из поисковой выдачи информативным (informative) или нет (notinformative).
Описание данных
Данные собраны из результатов поиска DuckDuck Go по различным запросам. Каждый пример представляет собой фрагмент текста (сниппет), который пользователь видит в результатах поиска, и… See the full description on the dataset page: https://huggingface.co/datasets/AIUserForPy/ru-search-snippets-classification.sec-filings-snippetsminiwob_snippets
Dataset Card for "miniwob_snippets"
More Information needed
miniwob_snippets_refs_onehot
Dataset Card for "miniwob_snippets_refs_onehot"
More Information needed
code_snippetssports-and-news-snippetsThis dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
sports_and_news_snippets
This dataset comprises short news articles and summaries covering diverse topics such as international rugby, football disciplinary actions, film awards, political developments, and technology product launches. The text samples are written in a journalistic style, focusing on specific events, quotes from key figures, and match or election outcomes. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/sports-and-news-snippets.Capybara-de-snippetsonly a few translated lines (from Mixtral, occiglot-7b-de-en-instruct-q4-k-m, gpt-4-0125-preview, Claude Opus, and others) to quickly compare the translation quality. a complete german translation from DeepL thankfully is provided at https://huggingface.co/datasets/maxidl/Capybara-de
enhanced-emo-snippets-balanced-DACVAE
Enhanced Emotion Snippets — Balanced DACVAE
A balanced, emotion-bucketed subset of TTS-AGI/enhanced-audiosnippets-DACVAE,
organized by Empathic Insight Voice+ emotion and voice attribute categories.
Overview
This dataset provides up to 100 samples per magnitude bucket for each of the
40 emotion categories and 15 voice attribute dimensions scored by
Empathic Insight Voice+.
Selection Criteria
Emotion Categories (40 dimensions)
For each emotion (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/enhanced-emo-snippets-balanced-DACVAE.comedy-snippets-v0.1A very small sampling of snippets of comedy routines by George Carlin and Tom Segura.
baoshidaoren_music_snippetsbird-dev-snippetswikipedia_fr_snippetscode_snippets_explainationbalanced_audio_snippets_40x3k
