datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm2vec-gen-tulu
LLM2Vec-Gen
The dataset consists of generations based on the Tulu-3 SFT data (https://huggingface.co/datasets/allenai/tulu-3-sft-mixture). These generations are intended to be used for training LLM2Vec-Gen models, serving as the target output for queries.
This dataset consists of various splits. Each split corresponds to responses generated by a specific LLM, e.g., Qwen3-4B. The "original" split refers to the original Tulu-3 responses.
Each instance in split M typically includes:… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/llm2vec-gen-tulu.llm2vec-gen-echo-rewritten-w-hard-negative
LLM2Vec-Gen
The dataset consists of generations based on the Echo data (Springer et al). The instruction+queries are rewritten in a natural tone using Gemini. The generations are intended to be used for training LLM2Vec-Gen models, serving as the target output for queries.
The negative_question in this dataset are also generated by Gemini. This dataset consists of various splits. Each split corresponds to responses generated by a specific LLM, e.g., Qwen3-4B. The "original" split… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/llm2vec-gen-echo-rewritten-w-hard-negative.llm2vec-gen-tulu-w-hard-negative
LLM2Vec-Gen
The dataset consists of generations based on the Tulu-3 SFT data (https://huggingface.co/datasets/allenai/tulu-3-sft-mixture). These generations are intended to be used for training LLM2Vec-Gen models, serving as the target output for queries.
The negative_question in this dataset are generated by Gemini. This dataset consists of various splits. Each split corresponds to responses generated by a specific LLM, e.g., Qwen3-4B. The "original" split refers to the original… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/llm2vec-gen-tulu-w-hard-negative.
