CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01567-labs /wikipedia-bge-small-en-v1.5-fulltext1M<n<10M4 likes3.3k downloads3y agoHugging Face02ko-vlm /KoLLaVA-v1.5-Instruct-581k KoLLaVA-v1.5-Instruct-581k 한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다. 데이터셋 정보 총 샘플 수: 435,093개 형식: ChatML 형식 (role: user/assistant, content: 텍스트) 이미지: COCO + GQA + Visual Genome 데이터셋 언어: 한국어 포함된 데이터셋 COCO 데이터: 362,953개 샘플 MS COCO 2017 이미지 기반 한국어 대화 데이터 GQA 데이터: 72,140개 샘플 GQA (Visual Question Answering) 이미지 기반 한국어 대화 데이터 Visual Genome 데이터: 포함 Visual Genome 이미지 기반 한국어 대화 데이터 제외된 데이터셋 EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.image100K<n<1M0 likes1.7k downloads1y agoHugging Face03Snowflake /mteb-retrieval-snowflake-arctic-embed-m-v1.5text10M<n<100M0 likes1.1k downloads2y agoHugging Face04Snowflake /msmarco-v2.1-snowflake-arctic-embed-m-v1.5 Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.textquestion-answering10M<n<100M0 likes785 downloads2y agoHugging Face05enjalot /fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5 FineWeb-edu 10BT Sample embedded with nomic-text-v1.5 The FineWeb-edu 10BT sample was first chunked into 500 tokens (using bert-base-uncased) with 10% overlap resulting in 25 million rows and 10.5BT. The chunks were then embedded using nomic-text-v1.5. Dataset Details Dataset Sources Repository: https://github.com/enjalot/fineweb-modal Uses Direct Use The dataset was embedded with the clustering: prefix, so the main… See the full description on the dataset page: https://huggingface.co/datasets/enjalot/fineweb-edu-sample-10BT-chunked-500-nomic-text-v1.5.tabular10M<n<100M5 likes568 downloads2y agoHugging Face06mcgillcomplex /wikipedia-2023-11-bge-large-en-v1.5 Multilingual Embeddings for Wikipedia This dataset contains the wikimedia/wikipedia dataset dump from 2023-11-01 from Wikipedia in all 300+ languages. And chunked from the Cohere/wikipedia-2023-11-embed-multilingual-v3. The embedding model is BAAI/bge-large-en-v1.5. text1M<n<10M1 likes457 downloads3y agoHugging Face07nhagar /dolma_urls_v1.5 Dataset Card for dolma_urls_v1.5 This dataset provides the URLs and top-level domains associated with training records in allenai/dolma. It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record identifiers. In doing so, it… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/dolma_urls_v1.5.text1B<n<10B0 likes388 downloads1y agoHugging Face08Menlo /instruction-speech-encodec-v1.5 Dataset Card for "Instruction Speech" The largest open-source English speech instruction to text answer dataset Dataset Overview This dataset contains over 332,000 English speech instruction to text answer samples, using: A subset of jan-hq/prompt-voice-v1.5. Audio generation using WhisperSpeech. Tokenized using Encodec. Usage from datasets import load_dataset, Audio # Load Instruction Speech dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/instruction-speech-encodec-v1.5.audio100K<n<1M7 likes355 downloads2y agoHugging Face09jan-hq /instruction-speech-text-v1.5-convo-male-voicetext100K<n<1M0 likes240 downloads2y agoHugging Face10Menlo /prompt-voice-v1.5 Dataset Overview This dataset contains nearly 2.35M English speech instruction to text answer samples, using the combination of: Intel/orca_dpo_pairs routellm/gpt4_dataset nomic-ai/gpt4all-j-prompt-generations microsoft/orca-math-word-problems-200k allenai/WildChat-1M Open-Orca/oo-gpt4-200k Magpie-Align/Magpie-Pro-300K-Filtered qiaojin/PubMedQA Undi95/Capybara-ShareGPT HannahRoseKirk/prism-alignment BAAI/Infinity-Instruct Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Menlo/prompt-voice-v1.5.tabular1M<n<10M0 likes235 downloads2y agoHugging Face11minishlab /tokenlearn-c4-en-bge-base-en-v1.5 minishlab/tokenlearn-c4-en-bge-base-v1.5 Dataset Card This dataset was created with Tokenlearn for training Model2Vec models. It contains mean token embeddings produced by a sentence transformer, used as training targets for static embedding distillation. Dataset Details Field Value Source dataset allenai/c4 Source split train Embedding model baai/bge-base-en-v1.5 Embedding dimension 768 Rows 10000000 Dataset Structure Column… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-c4-en-bge-base-en-v1.5.text10M<n<100M2 likes225 downloads6mo agoHugging Face12FrancophonIA /ACTER-v1.5 [!NOTE] Dataset origin: https://clarin.eurac.edu/repository/xmlui/handle/20.500.12124/47 ACTER Annotated Corpora for Term Extraction Research, version 1.5 ACTER is a manually annotated dataset for term extraction, covering 3 languages (English, French, and Dutch), and 4 domains (corruption, dressage, heart failure, and wind energy). Readme structure: General Abbreviations Data Structure Annotations Additional Information Updates Error Reporting License 1. General… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/ACTER-v1.5.text0 likes211 downloads1y agoHugging Face13CSSNB /dota_v1.5image100K<n<1M0 likes188 downloads8mo agoHugging Face14swiss-ai /Apertus_v1.5_Preference_Data Apertus 1.5 Preference Dataset This is the preference dataset used for the offline DPO stage of Apertus v1.5 alignment training, applied to the 70B model. The prompts come from Ai2's Olmo 3 Dolci-Instruct-DPO dataset. We only reuse the prompts from Dolci-Instruct-DPO; all chosen / rejected responses in this dataset were generated by us. How this dataset was built Prompts. Taken from Dolci-Instruct-DPO (ODC-BY). Response generation and annotation. Every prompt was… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/Apertus_v1.5_Preference_Data.tabulartext-generation100K<n<1M4 likes153 downloads1mo agoHugging Face15TheSkullery /Aether-V1.5 Aether Dataset Creator: SteelSkull Community Organization: ConvexAI Discord: Join us on Discord About Aether: The Aether dataset. rebuilt script, new dataset from 1.2.2 to 1.5, changed datasets, added two. version v1.5 is a rework of the human -> gpt conversations and added system and tool columns Source Datasets:… See the full description on the dataset page: https://huggingface.co/datasets/TheSkullery/Aether-V1.5.text1M<n<10M4 likes144 downloads3y agoHugging Face16bluuebunny /arxiv_embeddings_Alibaba-NLP_gte-base-en-v1.5text1M<n<10M1 likes144 downloads2y agoHugging Face17OALL /details_FreedomIntelligence__AceGPT-v1.5-13B-Chat Dataset Card for Evaluation run of FreedomIntelligence/AceGPT-v1.5-13B-Chat Dataset automatically created during the evaluation run of model FreedomIntelligence/AceGPT-v1.5-13B-Chat. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_FreedomIntelligence__AceGPT-v1.5-13B-Chat.tabular100K<n<1M0 likes141 downloads2y agoHugging Face18jan-hq /instruction-speech-text-v1.5-sound-convotext100K<n<1M0 likes131 downloads2y agoHugging Face19jan-hq /instruction-speech-conversation-v1.5-phase-2-interleavedtext1M<n<10M0 likes123 downloads2y agoHugging Face20mlx-community /Apertus-v1.5-QAT-10K mlx-community/Apertus-v1.5-QAT-10K This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models. texttext-generation10K<n<100K1 likes120 downloads6d agoHugging Face21spacemanidol /msmarco-v2.1-gte-large-en-v1.5 Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods. Note, that the embeddings are not normalized so you will need to normalize them before usage. Retrieval Performance Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.textquestion-answering10M<n<100M0 likes117 downloads1y agoHugging Face22jan-hq /instruction-speech-text-v1.5-combinedtext1M<n<10M0 likes90 downloads2y agoHugging Face23nleroy917 /fineweb-bge-large-en-v1.5 Work in progress. Coverage is partial — additional CommonCrawl dumps are being embedded and uploaded incrementally. FineWeb — BGE-Large-EN-v1.5 + BM25 (Pre-embedded) Pre-computed dense and sparse embeddings for the FineWeb web corpus, ready for direct ingestion into a vector database (e.g. Qdrant). Source dataset FineWeb is a 15-trillion-token English web dataset derived from 96 CommonCrawl snapshots spanning Summer 2013 through June 2025. It was produced by Hugging… See the full description on the dataset page: https://huggingface.co/datasets/nleroy917/fineweb-bge-large-en-v1.5.tabularfeature-extraction10M<n<100M0 likes86 downloads5mo agoHugging Face24jan-hq /instruction-speech-no-audio-v1.5tabular100K<n<1M0 likes80 downloads2y agoHugging Face25mistobaan /gsm8k-train-nomic-text-v1.5 Overview Dataset containing embeddings / classification information for GSM8K textquestion-answering1K<n<10K0 likes79 downloads2y agoHugging Face26lvogel123 /jailbreak-llama-3.3-nemotron-49b-v1.5tabular1K<n<10K0 likes76 downloads11mo agoHugging Face27ICT-TIME-and-Querit /BOOM-v1.5-training-data Some retrieval datasets of the first stage training are not uploaded: NQ, ELI5, TriviaQA, and MS MARCO document. Please waiting ... Or you can download from the offical website. Citation If you find our work helpful, feel free to give us a cite. @article{zhang2026bagging, title={Bagging-Based Model Merging for Robust General Text Embeddings}, author={Zhang, Hengran and Bi, Keping and Guo, Jiafeng and Zhang, Jiaming and Yang, Wenbo and Shi, Daiting and Cheng… See the full description on the dataset page: https://huggingface.co/datasets/ICT-TIME-and-Querit/BOOM-v1.5-training-data.textsentence-similarity1M<n<10M0 likes76 downloads4mo agoHugging Face28migtissera /Synthia-Coder-v1.5-Itext10K<n<100K34 likes75 downloads2y agoHugging Face29567-labs /wikipedia-embedding-bge-small-en-v1.5-five-percenttext1M<n<10M0 likes72 downloads3y agoHugging Face30much1na /nanobeir-snowflake-arctic-embed-m-v1.5text10K<n<100K0 likes60 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.