snippets
SBERT-news-snippetscode-generation-model-fine-tuned-to-produce-good-code-snippetsphi-2-snippets-logdetectivefacebook-xml-roberta-base-300k-gg-snippets-bnb-4bitlangtok-bert-translated_dataset_longer_snippetsreview_snippetsfacebook-xml-roberta-base-300k-gg-snippetslangtok-bert-same_sentence_longer_snippets_one
Datasets
All datasets matching “snippets”captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
wiki_snippets
Dataset Card for "wiki_snippets"
Dataset Summary
Wikipedia version split into plain text snippets for dense semantic indexing.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in
English:
wiki40b_en_100_0: Wiki-40B
wikipedia_en_100_0: Wikipedia
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.AuraShield_Vulnerable_Code_Snippetsunsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_2wikipedia-language-snippets-filtered
Wikipedia Snippets (Filtered)
Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage.
Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language.
Files
Each file is in this format for languages in ISO 639 2-letter codes:
train/en/en.parquet
train/es/es.parquet
From wikimedia/wikipedia
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.
