CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face02community-datasets /wiki_snippets Dataset Card for "wiki_snippets" Dataset Summary Wikipedia version split into plain text snippets for dense semantic indexing. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in English: wiki40b_en_100_0: Wiki-40B wikipedia_en_100_0: Wikipedia Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.tabulartext-generation10M<n<100M6 likes1.5k downloads2y agoHugging Face03Anju405712 /AuraShield_Vulnerable_Code_Snippetstext1K<n<10K0 likes721 downloads6mo agoHugging Face04laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes667 downloads1y agoHugging Face05laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_20 likes515 downloads1y agoHugging Face06polyglot-tagger /wikipedia-language-snippets-filtered Wikipedia Snippets (Filtered) Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage. Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language. Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From wikimedia/wikipedia Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.texttext-generation10M<n<100M0 likes285 downloads5mo agoHugging Face07vida-nyu /pmc-articles-dataset-mentions-snippets PMC Articles Dataset Mentions Snippets Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature. Description Task: Extract structured dataset info (identifier, repository, webpage) from article text Source: PMC open-access articles Format: Text snippet → JSON output Examples: Positive (with datasets) and negative (no datasets) Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.tabular1K<n<10K0 likes216 downloads2mo agoHugging Face08TTS-AGI /balanced-audio-snippets-40x3k-DACVAEtext100K<n<1M0 likes181 downloads6mo agoHugging Face09mitermix /balanced_audio_snippets_40x10k0 likes166 downloads8mo agoHugging Face10wambosec /insecure-code-snippets0 likes149 downloads1y agoHugging Face11cestwc /cnn_dailymail-snippetstext1M<n<10M0 likes142 downloads5y agoHugging Face12laion /snippets_for_soundscape_generation0 likes131 downloads1y agoHugging Face13JesseParvess /book_snippets_asrtextn<1K0 likes124 downloads5y agoHugging Face14polyglot-tagger /nlp-noise-snippets Synthetic Noise Pool For Text Classification purposes, as many models may consider code snippets, html artifacts, and math as "English". Around 50K are latex snippets from im2latex-100k texttext-classification100K<n<1M0 likes107 downloads5mo agoHugging Face15whackthejacker /vulnerable-code-snippets-for-supervised-learning Dataset Card for vulnerable-code-snippets-for-supervised-learning This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning/raw/main/pipeline.yaml" or explore the configuration:… See the full description on the dataset page: https://huggingface.co/datasets/whackthejacker/vulnerable-code-snippets-for-supervised-learning.texttext-classificationn<1K0 likes95 downloads2y agoHugging Face16TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: 1.54 TB Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.audiotext-to-speech100K<n<1M0 likes83 downloads6mo agoHugging Face17AIUserForPy /ru-search-snippets-classification Russian Search Snippets Informative Classification Этот датасет предназначен для задачи бинарной классификации текстовых сниппетов на русском языке. Цель — определить, является ли сниппет из поисковой выдачи информативным (informative) или нет (notinformative). Описание данных Данные собраны из результатов поиска DuckDuck Go по различным запросам. Каждый пример представляет собой фрагмент текста (сниппет), который пользователь видит в результатах поиска, и… See the full description on the dataset page: https://huggingface.co/datasets/AIUserForPy/ru-search-snippets-classification.textn<1K1 likes79 downloads16d agoHugging Face18DerivedFunction01 /sec-filings-snippetstabularfill-mask100K<n<1M0 likes76 downloads6mo agoHugging Face19LucasThil /miniwob_snippets Dataset Card for "miniwob_snippets" More Information needed tabular100K<n<1M0 likes47 downloads4y agoHugging Face20LucasThil /miniwob_snippets_refs_onehot Dataset Card for "miniwob_snippets_refs_onehot" More Information needed tabular100K<n<1M0 likes44 downloads4y agoHugging Face21TechxGenus /code_snippetstext100K<n<1M0 likes44 downloads2y agoHugging Face22sarahooker /sports-and-news-snippetsThis dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. sports_and_news_snippets This dataset comprises short news articles and summaries covering diverse topics such as international rugby, football disciplinary actions, film awards, political developments, and technology product launches. The text samples are written in a journalistic style, focusing on specific events, quotes from key figures, and match or election outcomes. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/sports-and-news-snippets.text1K<n<10K0 likes44 downloads6mo agoHugging Face23cstr /Capybara-de-snippetsonly a few translated lines (from Mixtral, occiglot-7b-de-en-instruct-q4-k-m, gpt-4-0125-preview, Claude Opus, and others) to quickly compare the translation quality. a complete german translation from DeepL thankfully is provided at https://huggingface.co/datasets/maxidl/Capybara-de 0 likes42 downloads2y agoHugging Face24TTS-AGI /enhanced-emo-snippets-balanced-DACVAE Enhanced Emotion Snippets — Balanced DACVAE A balanced, emotion-bucketed subset of TTS-AGI/enhanced-audiosnippets-DACVAE, organized by Empathic Insight Voice+ emotion and voice attribute categories. Overview This dataset provides up to 100 samples per magnitude bucket for each of the 40 emotion categories and 15 voice attribute dimensions scored by Empathic Insight Voice+. Selection Criteria Emotion Categories (40 dimensions) For each emotion (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/enhanced-emo-snippets-balanced-DACVAE.textaudio-classification10K<n<100K0 likes39 downloads6mo agoHugging Face25unalignment /comedy-snippets-v0.1A very small sampling of snippets of comedy routines by George Carlin and Tom Segura. textn<1K10 likes30 downloads3y agoHugging Face26zuona /baoshidaoren_music_snippetsaudion<1K0 likes29 downloads2y agoHugging Face27567-labs /bird-dev-snippetstext1K<n<10K0 likes25 downloads2y agoHugging Face28CronosGhost /wikipedia_fr_snippetstext10M<n<100M0 likes24 downloads3y agoHugging Face29anujbishtTx /code_snippets_explainationtextn<1K0 likes24 downloads2y agoHugging Face30mitermix /balanced_audio_snippets_40x3k0 likes24 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.