datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nomic-embed-unsupervised-dataWeakly Supervised Contrastive Training data for Text Embedding models used in Nomic Embed models
Training
Click the Nomic Atlas map below to visualize a 5M sample of our contrastive pretraining data!
We train our embedder using a multi-stage training pipeline. Starting from a long-context BERT model,
the first unsupervised contrastive stage trains on a dataset generated from weakly related text pairs, such as question-answer pairs from forums like StackExchange and Quora… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/nomic-embed-unsupervised-data.unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1nomic_embed_unsupervisednomic_embed_unsupervisedunsupervised-multilingualsubsampled-jxm-nomic-unsupervisedchallenges-for-unsupervised-elicitation
Challenges for Unsupervised Elicitation
Datasets for the paper Three Concrete Challenges and Two Hopes for the Safety of Unsupervised Elicitation. These are used to stress-test unsupervised elicitation and easy-to-hard generalization techniques on imbalanced training sets and impossible tasks.
Code: challenges-for-unsupervised-elicitation
Subsets
gsm8k
Derived from GSM8K (Cobbe et al., 2021), using candidate LLM solutions provided by Wen et al. (2025).… See the full description on the dataset page: https://huggingface.co/datasets/callum-canavan/challenges-for-unsupervised-elicitation.qtack-gq-embeddings-unsupervised
Dataset Description
The QTACK Embedding Training Dataset merges multiple publicly available datasets optimized for training sentence embedding and retrieval models. It consists of question-answer pairs, summarization pairs, semantic similarity sentence pairs, retrieval pairs, and clustering data, providing a comprehensive dataset suitable for various NLP tasks.
Dataset Structure
Data Fields
Each example contains the following fields:
query: The input text… See the full description on the dataset page: https://huggingface.co/datasets/prdev/qtack-gq-embeddings-unsupervised.unsupervisedWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
Unsupervised_Keyphrase_Extractiontouch-rugby-rules-unsupervised
Touch Rugby Rules Dataset
train.csv is taken from the International Touch Website
All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible.
For educational and non-commercial use only.
imdb-unsupervised-labeledThis dataset was created by extracting 10,000 unsupervised splits from the IMDB dataset and labeling them.
Information
Model used for labeling: dfurman/deberta-v3-base-imdb
Number of samples: 10,000
truthfulqa-unsupervised-elicitationimdb-unsupervised-mini
imdb-unsupervised-mini
This dataset is a subset of the IMDB unsupervised dataset (stanfordnlp/imdb). Sentiments have been automatically labeled using Meta LLaMA 3.1 (8B) (meta-llama/Llama-3.1-8B-Instruct) via the Together.ai API.
Citation:
@InProceedings{maas-EtAl:2011:ACL-HLT2011,
author = {Maas, Andrew L. and Daly, Raymond E. and Pham, Peter T. and Huang, Dan and Ng, Andrew Y. and Potts, Christopher},
title = {Learning Word Vectors for Sentiment Analysis}… See the full description on the dataset page: https://huggingface.co/datasets/doanhieung/imdb-unsupervised-mini.unsupervised_finetuningunsupervised_kin_tweets
Dataset Card for "unsupervised_kin_tweets"
More Information needed
unsupervised_tcsallfriedrich_nietzsche_books_unsupervisedfake-audio-detection-unsupervisedclean-unsupervised-kin-tweets
Dataset Card for "clean-unsupervised-kin-tweets"
More Information needed
thesis_unsupervisedboolq-banana-shed-unsupervised-elicitationunsupervisedtouch_rugby_unsupervisedunsupervised_finetuningunsupervised_peoples_speech_with_annotations1math500-unsupervised-iclcopa-unsupervised-elicitationunsupervised_title-factWe do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
piqa-unsupervised-elicitation
