datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bekko-embedding-v1-unsupervised
Bekko Embedding v1 Unsupervised Training Data
This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently.
For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.unsupervised_peoples_speech
Dataset Card for Unsupervised Peoples Speech
Dataset Description
Dataset Summary
The Unsupervised Peoples Speech Dataset is a compilation of audiofiles extracted from Archive.org that is licensed for academic and commercial usage under CC-BY and CC-BY-SA licenses. It includes more than one million hours of audio with a diverse set of speakers.
Point of Contact: MLCommons Datasets Discord
Dataset Structure
This dataset is a collection of audio… See the full description on the dataset page: https://huggingface.co/datasets/MLCommons/unsupervised_peoples_speech.nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models
Training
Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data!
We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model,
the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_2nomic_embed_unsupervisednomic_embed_unsupervisedunsupervised-multilingualUnsupervisedRootVideos
ChronoRoot Dataset without annotations and full time series
Project Page | Paper | GitHub
Dataset Description
Dataset Summary
This dataset contains the complete time series from where the annotations were extracted, to train weakly supervised models when combining these with the annotated frames.
Data Structure
Raw Images: 3280 x 2464 infrared images
Personal and Sensitive Information
This dataset contains no personal or sensitive… See the full description on the dataset page: https://huggingface.co/datasets/ngaggion/UnsupervisedRootVideos.Unsupervised-Evaluation-of-Latent-Spacesubsampled-jxm-nomic-unsupervisedDAVIS-2017-Unsupervised-trainval-480pchallenges-for-unsupervised-elicitation
Challenges for Unsupervised Elicitation
Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks.
Code: challenges-for-unsupervised-elicitation
Subsets
gsm8k
Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.qtack-gq-embeddings-unsupervised
Dataset Description
The QTACK Embedding Training Dataset merges multiple publicly available datasets optimized for training sentence embedding and retrieval models. It consists of question-answer pairs, summarization pairs, semantic similarity sentence pairs, retrieval pairs, and clustering data, providing a comprehensive dataset suitable for various NLP tasks.
Dataset Structure
Data Fields
Each example contains the following fields:
query: The input text… See the full description on the dataset page: https://huggingface.co/datasets/prdev/qtack-gq-embeddings-unsupervised.unsupervisedWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
touch-rugby-rules-unsupervised
Touch Rugby Rules Dataset
train.csv is taken from the International Touch Website
All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible.
For educational and non-commercial use only.
Unsupervised_Keyphrase_Extractionimdb-unsupervised-labeledThis dataset was created by extracting 10,000 unsupervised splits from the IMDB dataset and labeling them.
Information
Model used for labeling: dfurman/deberta-v3-base-imdb
Number of samples: 10,000
truthfulqa-unsupervised-elicitationimdb-unsupervised-mini
imdb-unsupervised-mini
This dataset is a subset of the IMDB unsupervised dataset (stanfordnlp/imdb). Sentiments have been automatically labeled using Meta LLaMA 3.1 (8B) (meta-llama/Llama-3.1-8B-Instruct) via the Together.ai API.
Citation:
@InProceedings{maas-EtAl:2011:ACL-HLT2011,
author = {Maas, Andrew L. and Daly, Raymond E. and Pham, Peter T. and Huang, Dan and Ng, Andrew Y. and Potts, Christopher},
title = {Learning Word Vectors for Sentiment Analysis}… See the full description on the dataset page: https://huggingface.co/datasets/doanhieung/imdb-unsupervised-mini.Improving-Unsupervised-Constituency-Parsing-via-Maximizing-Semantic-Informationunsupervised_finetuningunsupervised_kin_tweets
Dataset Card for "unsupervised_kin_tweets"
More Information needed
Language-Model-Based-Unsupervised-Dependency-Parsing-with-CMI-and-GCunsupervised_tcsallunsupervised-malay-youtube-speaker-diarization
Unsupervised malay speakers from youtube videos
10492 unique speakers with at least 75 hours of voice activities. Steps to reproduce at https://github.com/huseinzol05/malaya-speech/blob/master/data/youtube/process-youtube.ipynb
how-to
Download and extract processed-youtube.tar.gz, each processed videos saved as pickle, {video_name}.pkl.
Each pickle file got,
[{'wav_data':… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/unsupervised-malay-youtube-speaker-diarization.friedrich_nietzsche_books_unsupervisedfake-audio-detection-unsupervisedclean-unsupervised-kin-tweets
Dataset Card for "clean-unsupervised-kin-tweets"
More Information needed
thesis_unsupervised
