datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-triplets-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-0-4msmarco-triplets-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-5-9triviaqa-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-0-0squad-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-0-1nq-train-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-0-1dureader-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/dureader-hard-neg-reasoning-embedding.GLM-5.1-Reasoning-1M-Think-Embeddings
GLM-5.1-Reasoning-1M-Think-Embeddings (main subset, partial)
Embeddings of the think-block content of every record in the main
subset of
Jackrong/GLM-5.1-Reasoning-1M-Cleaned,
embedded with Qwen/Qwen3-Embedding-0.6B via vLLM.
What this dataset is
Each row is the vector representation of just the reasoning trace (the
text between <think>...</think>), not the user prompt and not the final
answer. Useful for:
searching / clustering the original reasoning traces by… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/GLM-5.1-Reasoning-1M-Think-Embeddings.triviaqa-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-1-1hotpotqa-hard-neg-reasoning-embedding-modified-SFT-4B-doc4096-seq1024-v2-parts-0-1allnli-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/allnli-hard-neg-reasoning-embedding.hotpotqa-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/hotpotqa-hard-neg-reasoning-embedding.squad-pairs-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/squad-pairs-hard-neg-reasoning-embedding.mrtydi-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/mrtydi-reasoning-embedding.fever-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/fever-hard-neg-reasoning-embedding.embeddings_all_correct_continuationshotpotqa-hard-neg-reasoning-embedding-modifiedmsmarco-triplets-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/msmarco-triplets-hard-neg-reasoning-embedding.quora_duplicates-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/quora_duplicates-hard-neg-reasoning-embedding.nq-train-pairs-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/nq-train-pairs-hard-neg-reasoning-embedding.squad-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-parts-0-1triviaqa-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-parts-0-0CoT-Moderate-Reasoning-Embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Moderate-Reasoning-Embedding.CoT-Hard-Reasoning-Embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Hard-Reasoning-Embedding.triviaqa-pairs-hard-neg-reasoning-embedding-modifiedtriviaqa-pairs-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/triviaqa-pairs-hard-neg-reasoning-embedding.CoT-Easy-Reasoning-Embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to evaluate the model similarity with the Hierarchical Representation Similarity Analysis (HRSA) framework in the paper Do Reasoning Models Enhance Embedding Models?. To verify if the latent manifold is preserved even within reasoning trajectories, we construct a Chain-of-Thought (CoT) dataset. Unlike standard semantic datasets, this corpus… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/CoT-Easy-Reasoning-Embedding.triviaqa-pairs-hard-neg-reasoning-embedding-modified-SFT-4B-parts-0-1msmarco-triplets-hard-neg-reasoning-embedding-modifiedt2ranking-hard-neg-reasoning-embedding
Do Reasoning Models Enhance Embedding Models?
Introduction
This is the dataset used to train the embedding models in the paper Do Reasoning Models Enhance Embedding Models?. We use Qwen3-Embedding-0.6B to mine 3 hard negatives per query, and employ the positive-aware hard negative mining technique introduced in NV-Retriever with 95% margin to the positive score.
Abstract
State-of-the-art embedding models are… See the full description on the dataset page: https://huggingface.co/datasets/lucaswychan/t2ranking-hard-neg-reasoning-embedding.nq-train-pairs-hard-neg-reasoning-embedding-modified
