CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hotchpotch /bekko-embedding-v1-unsupervised Bekko Embedding v1 Unsupervised Training Data This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently. For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.text-retrieval4 likes29k downloads2mo agoHugging Face02MLCommons /unsupervised_peoples_speech Dataset Card for Unsupervised Peoples Speech Dataset Description Dataset Summary The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers. Point of Contact: MLCommons Datasets Discord Dataset Structure This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.audioautomatic-speech-recognition81 likes26k downloads2y agoHugging Face03nomic-ai /nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models Training Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data! We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model, the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.text100M<n<1B21 likes3.8k downloads2y agoHugging Face04laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes667 downloads1y agoHugging Face05laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_20 likes515 downloads1y agoHugging Face06jxm /nomic_embed_unsupervisedtext100M<n<1B4 likes460 downloads2y agoHugging Face07nomic-ai /nomic_embed_unsupervisedtext100M<n<1B1 likes379 downloads3y agoHugging Face08gowitheflow /unsupervised-multilingualtext10M<n<100M0 likes138 downloads2y agoHugging Face09ngaggion /UnsupervisedRootVideos ChronoRoot Dataset without annotations and full time series Project Page | Paper | GitHub Dataset Description Dataset Summary This dataset contains the complete time series from where the annotations were extracted, to train weakly supervised models when combining these with the annotated frames. Data Structure Raw Images: 3280 x 2464 infrared images Personal and Sensitive Information This dataset contains no personal or sensitive… See the full description on the dataset page: https://huggingface.co/datasets/ngaggion/UnsupervisedRootVideos.imageimage-segmentation1K<n<10K0 likes100 downloads7mo agoHugging Face10convergedmachine /Unsupervised-Evaluation-of-Latent-Space0 likes76 downloads11mo agoHugging Face11prdev /subsampled-jxm-nomic-unsupervisedtext10M<n<100M0 likes65 downloads2y agoHugging Face12rg4q33 /DAVIS-2017-Unsupervised-trainval-480p0 likes65 downloads6mo agoHugging Face13callum-canavan /challenges-for-unsupervised-elicitation Challenges for Unsupervised Elicitation Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks. Code: challenges-for-unsupervised-elicitation Subsets gsm8k Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.tabulartext-classification10K<n<100K0 likes50 downloads7mo agoHugging Face14prdev /qtack-gq-embeddings-unsupervised Dataset Description The QTACK Embedding Training Dataset merges multiple publicly available datasets optimized for training sentence embedding and retrieval models. It consists of question-answer pairs, summarization pairs, semantic similarity sentence pairs, retrieval pairs, and clustering data, providing a comprehensive dataset suitable for various NLP tasks. Dataset Structure Data Fields Each example contains the following fields: query: The input text… See the full description on the dataset page: https://huggingface.co/datasets/prdev/qtack-gq-embeddings-unsupervised.text10M<n<100M1 likes48 downloads2y agoHugging Face15nlp-brin-id /unsupervisedWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ texttext-classification100K<n<1M0 likes33 downloads2y agoHugging Face16Trelis /touch-rugby-rules-unsupervised Touch Rugby Rules Dataset train.csv is taken from the International Touch Website All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible. For educational and non-commercial use only. texttext-generationn<1K0 likes22 downloads3y agoHugging Face17Nickyang /Unsupervised_Keyphrase_Extractiontextn<1K1 likes21 downloads3y agoHugging Face18Lumia101 /imdb-unsupervised-labeledThis dataset was created by extracting 10,000 unsupervised splits from the IMDB dataset and labeling them. Information Model used for labeling: dfurman/deberta-v3-base-imdb Number of samples: 10,000 texttext-classification10K<n<100K0 likes20 downloads2mo agoHugging Face19callum-canavan /truthfulqa-unsupervised-elicitationtext1K<n<10K0 likes12 downloads1y agoHugging Face20doanhieung /imdb-unsupervised-mini imdb-unsupervised-mini This dataset is a subset of the IMDB unsupervised dataset (stanfordnlp/imdb). Sentiments have been automatically labeled using Meta LLaMA 3.1 (8B) (meta-llama/Llama-3.1-8B-Instruct) via the Together.ai API. Citation: @InProceedings{maas-EtAl:2011:ACL-HLT2011, author = {Maas, Andrew L. and Daly, Raymond E. and Pham, Peter T. and Huang, Dan and Ng, Andrew Y. and Potts, Christopher}, title = {Learning Word Vectors for Sentiment Analysis}… See the full description on the dataset page: https://huggingface.co/datasets/doanhieung/imdb-unsupervised-mini.texttext-classification1K<n<10K0 likes11 downloads2y agoHugging Face21junjiechen-chris /Improving-Unsupervised-Constituency-Parsing-via-Maximizing-Semantic-Information0 likes10 downloads2y agoHugging Face22tripti27 /unsupervised_finetuningtextn<1K0 likes10 downloads2y agoHugging Face23RogerB /unsupervised_kin_tweets Dataset Card for "unsupervised_kin_tweets" More Information needed text10K<n<100K0 likes8 downloads3y agoHugging Face24junjiechen-chris /Language-Model-Based-Unsupervised-Dependency-Parsing-with-CMI-and-GC0 likes8 downloads2y agoHugging Face25Triptigarg2711 /unsupervised_tcsalltext1K<n<10K0 likes8 downloads1y agoHugging Face26mesolitica /unsupervised-malay-youtube-speaker-diarization Unsupervised malay speakers from youtube videos 10492 unique speakers with at least 75 hours of voice activities. Steps to reproduce at https://github.com/huseinzol05/malaya-speech/blob/master/data/youtube/process-youtube.ipynb how-to Download and extract processed-youtube.tar.gz, each processed videos saved as pickle, {video_name}.pkl. Each pickle file got, [{'wav_data':… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/unsupervised-malay-youtube-speaker-diarization.0 likes7 downloads4y agoHugging Face27Augustya07 /friedrich_nietzsche_books_unsupervisedtextn<1K0 likes7 downloads3y agoHugging Face28012shin /fake-audio-detection-unsupervisedaudio1K<n<10K0 likes7 downloads2y agoHugging Face29RogerB /clean-unsupervised-kin-tweets Dataset Card for "clean-unsupervised-kin-tweets" More Information needed text10K<n<100K0 likes6 downloads3y agoHugging Face30JLAbe /thesis_unsupervisedtext10K<n<100K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.