datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-bge-small-en-v1.5-fullKoLLaVA-v1.5-Instruct-581k
KoLLaVA-v1.5-Instruct-581k
한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다.
데이터셋 정보
총 샘플 수: 435,093개
형식: ChatML 형식 (role: user/assistant, content: 텍스트)
이미지: COCO + GQA + Visual Genome 데이터셋
언어: 한국어
포함된 데이터셋
COCO 데이터: 362,953개 샘플
MS COCO 2017 이미지 기반
한국어 대화 데이터
GQA 데이터: 72,140개 샘플
GQA (Visual Question Answering) 이미지 기반
한국어 대화 데이터
Visual Genome 데이터: 포함
Visual Genome 이미지 기반
한국어 대화 데이터
제외된 데이터셋
EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.mteb-retrieval-snowflake-arctic-embed-m-v1.5msmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5
FineWeb-edu 10BT Sample embedded with nomic-text-v1.5
The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT.
The chunks were then embedded using nomic-text-v1.5.
Dataset Details
Dataset Sources
Repository: https://github.com/enjalot/fineweb-modal
Uses
Direct Use
The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.wikipedia-2023-11-bge-large-en-v1.5
Multilingual Embeddings for Wikipedia
This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages.
And chunked from the Cohere/wikipedia-2023-11-embed-multilingual-v3.
The embedding model is BAAI/bge-large-en-v1.5.
dolma_urls_v1.5
Dataset Card for dolma_urls_v1.5
This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible.
Dataset Details
Dataset Description
This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.5.instruction-speech-encodec-v1.5
Dataset Card for "Instruction Speech"
The largest open-source English speech instruction to text answer dataset
Dataset Overview
This dataset contains over 332,000 English speech instruction to text answer samples, using:
A subset of jan-hq/prompt-voice-v1.5.
Audio generation using WhisperSpeech.
Tokenized using Encodec.
Usage
from datasets import load_dataset, Audio
# Load Instruction Speech dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/instruction-speech-encodec-v1.5.instruction-speech-text-v1.5-convo-male-voiceprompt-voice-v1.5
Dataset Overview
This dataset contains nearly 2.35M English speech instruction to text answer samples, using the combination of:
Intel/orca_dpo_pairs
routellm/gpt4_dataset
nomic-ai/gpt4all-j-prompt-generations
microsoft/orca-math-word-problems-200k
allenai/WildChat-1M
Open-Orca/oo-gpt4-200k
Magpie-Align/Magpie-Pro-300K-Filtered
qiaojin/PubMedQA
Undi95/Capybara-ShareGPT
HannahRoseKirk/prism-alignment
BAAI/Infinity-Instruct
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/prompt-voice-v1.5.tokenlearn-c4-en-bge-base-en-v1.5
minishlab/tokenlearn-c4-en-bge-base-v1.5 Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models. It contains mean token embeddings produced by a sentence transformer, used as training targets for static embedding distillation.
Dataset Details
Field
Value
Source dataset
allenai/c4
Source split
train
Embedding model
baai/bge-base-en-v1.5
Embedding dimension
768
Rows
10000000
Dataset Structure
Column… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-c4-en-bge-base-en-v1.5.ACTER-v1.5
[!NOTE]
Dataset origin: https://clarin.eurac.edu/repository/xmlui/handle/20.500.12124/47
ACTER Annotated Corpora for Term Extraction Research, version 1.5
ACTER is a manually annotated dataset for term extraction, covering 3 languages (English, French, and Dutch), and 4 domains (corruption, dressage, heart failure, and wind energy).
Readme structure:
General
Abbreviations
Data Structure
Annotations
Additional Information
Updates
Error Reporting
License
1. General… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/ACTER-v1.5.dota_v1.5Apertus_v1.5_Preference_Data
Apertus 1.5 Preference Dataset
This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model.
The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us.
How this dataset was built
Prompts. Taken from Dolci-Instruct-DPO (ODC-BY).
Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus_v1.5_Preference_Data.Aether-V1.5
Aether Dataset
Creator: SteelSkull
Community Organization: ConvexAI
Discord: Join us on Discord
About Aether: The Aether dataset.
rebuilt script, new dataset
from 1.2.2 to 1.5, changed datasets, added two.
version v1.5 is a rework of the human -> gpt conversations and added system and tool columns
Source Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-V1.5.arxiv_embeddings_Alibaba-NLP_gte-base-en-v1.5details_FreedomIntelligence__AceGPT-v1.5-13B-Chat
Dataset Card for Evaluation run of FreedomIntelligence/AceGPT-v1.5-13B-Chat
Dataset automatically created during the evaluation run of model FreedomIntelligence/AceGPT-v1.5-13B-Chat.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_FreedomIntelligence__AceGPT-v1.5-13B-Chat.instruction-speech-text-v1.5-sound-convoinstruction-speech-conversation-v1.5-phase-2-interleavedApertus-v1.5-QAT-10K
mlx-community/Apertus-v1.5-QAT-10K
This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models.
msmarco-v2.1-gte-large-en-v1.5
Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.instruction-speech-text-v1.5-combinedfineweb-bge-large-en-v1.5
Work in progress. Coverage is partial — additional CommonCrawl dumps are being embedded and uploaded incrementally.
FineWeb — BGE-Large-EN-v1.5 + BM25 (Pre-embedded)
Pre-computed dense and sparse embeddings for the FineWeb web corpus, ready for direct ingestion into a vector database (e.g. Qdrant).
Source dataset
FineWeb is a 15-trillion-token English web dataset derived from 96 CommonCrawl snapshots spanning Summer 2013 through June 2025. It was produced by Hugging… See the full description on the dataset page: https://huggingface.co/datasets/nleroy917/fineweb-bge-large-en-v1.5.instruction-speech-no-audio-v1.5gsm8k-train-nomic-text-v1.5
Overview
Dataset containing embeddings / classification information for GSM8K
jailbreak-llama-3.3-nemotron-49b-v1.5BOOM-v1.5-training-data
Some retrieval datasets of the first stage training are not uploaded: NQ, ELI5, TriviaQA, and MS MARCO document. Please waiting ... Or you can download from the offical website.
Citation
If you find our work helpful, feel free to give us a cite.
@article{zhang2026bagging,
title={Bagging-Based Model Merging for Robust General Text Embeddings},
author={Zhang, Hengran and Bi, Keping and Guo, Jiafeng and Zhang, Jiaming and Yang, Wenbo and Shi, Daiting and Cheng… See the full description on the dataset page: https://huggingface.co/datasets/ICT-TIME-and-Querit/BOOM-v1.5-training-data.Synthia-Coder-v1.5-Iwikipedia-embedding-bge-small-en-v1.5-five-percentnanobeir-snowflake-arctic-embed-m-v1.5
