CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01selmanbaysan /turkish_embedding_model_training_datatextsentence-similarity100M<n<1B5 likes1.4k downloads1y agoHugging Face02selmanbaysan /cleaned_turkish_embedding_model_training_data_colabtext10M<n<100M1 likes509 downloads1y agoHugging Face03trmteb /cleaned_turkish_embedding_model_training_data_colab Citation If you use this dataset in your research, please cite the following paper: @inproceedings{baysan-gungor-2025-tr, title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations", author = "Baysan, Mehmet Selman and Gungor, Tunga", booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025", month = nov, year = "2025", address = "Suzhou, China", publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.text10M<n<100M3 likes425 downloads10mo agoHugging Face04jaiganesan /Embedding-model-fine-tuning-datasettext1K<n<10K0 likes159 downloads2y agoHugging Face05trmteb /turkish_embedding_model_training_datatext10M<n<100M0 likes121 downloads1y agoHugging Face06vab46 /Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final Dataset details:- This dataset is the final version of anchor(query)-positive(chunk) pair data w.r.t fine tuning embedding model for clinical trials dataset. It includes best of both 4 anchors-consolidated positive chunk/nctId dataset-->first dataset and 5 anchors-3 positive chunk/nctId--->Second dataset. The 1st dataset(consolidated title +summary+ inclusion criteria chunk) suffered with pre-processing bottlenecks :- rendering huge chunks upto 15k characeters. missing on… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final.tabular1K<n<10K1 likes107 downloads1mo agoHugging Face07LLM-OS-Models /korean-embedding-performance-v1-performance-1m Korean Embedding Performance v1 — 1M Qwen3-Embedding-8B의 한국어 retrieval data-scale 실험을 위한 정확히 1,000,000-row 연구·비상업 contrastive dataset이다. release_eligible: false, 통합 라이선스 other이며 upstream source 조건을 재허가하지 않는다. 구성 계열 Rows 비율 역할 nlpai-lab/ko-triplet-v1.0 600,254 60.03% 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 287,000 28.70% webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task train-family 4,146 0.41% MIRACL, MrTidy, MLDR F2 PAWS-X… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-performance-1m.textsentence-similarity1M<n<10M0 likes103 downloads3mo agoHugging Face08LLM-OS-Models /korean-embedding-performance-v1-ablation-200k Korean Embedding Performance v1 — Ablation 200K Qwen3-Embedding-8B의 한국어 retrieval continued fine-tuning에서 LoRA/DoRA/부분 및 full fine-tuning, loss, hard-negative 전략을 비교하기 위한 200,000-row 연구·비상업 성능 데이터다. release_eligible: false이며 통합 라이선스는 other다. upstream source별 조건을 재허가하지 않는다. 구성 계열 Rows 역할 nlpai-lab/ko-triplet-v1.0@1f5d72d 100,254 넓은 한국어 QA/retrieval core F2 Korean QA/instruction 68,000 webfaq, mqa, koalpaca, realQA, komagpie F2 retrieval task… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-ablation-200k.textsentence-similarity100K<n<1M0 likes95 downloads3mo agoHugging Face09LLM-OS-Models /korean-embedding-performance-v1-pilot-50k Korean Embedding Performance v1 — Pilot 50K 주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다. 사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다. 파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은 ablation-200k이다. Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용 contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개, hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다. 사용 조건과 공개 범위 이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.texttext-retrieval10K<n<100K0 likes91 downloads3mo agoHugging Face10vab46 /Clinical_trials_anchor-positive-pairs_EmbeddingModel-data2 Dataset details:- This is the iteration 2 of the previous dataset on same repository. Overall similar usecase:-(i)embedding model fine-tune/training for CT(clinical trials) domain, thereby aiding downstream tsak like ranked retrieval document generation comparing 2 or more anchors, anchor vs chunks via cosine similarity. Key changes(from earlier version):- Greater granulation of chunks so to have:- (i) cleaner directed retreival from fine tuned embedding model;… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data2.tabular1K<n<10K1 likes66 downloads1mo agoHugging Face11LLM-OS-Models /korean-embedding-performance-v1-sionic-retrieval-train-family-4146 Korean Sionic Retrieval Train-Family 4,146 F2LLM-v2가 공개한 Korean MIRACL, MrTidy, MLDR train-family row만 1M decontaminated curriculum에서 lossless 추출한 target-adaptation dataset이다. 공개 evaluation query는 포함하지 않으며 current-student HN7 mining 전의 source artifact다. 구성과 목적 source rows 역할 f2_miracl_ko_train 700 MIRACL Korean retrieval train-family f2_mrtidy_korean_train 1,200 MrTidy Korean train f2_mldr_ko_train 2,246 MLDR Korean long-document train-family 합계 4… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-retrieval-train-family-4146.texttext-retrieval1K<n<10K0 likes60 downloads3mo agoHugging Face12vab46 /Clinical_trials_anchor-positive-pairs_EmbeddingModel-data Dataset details:- This dataset is basically mapping of anchor-positive pair chunks. It consists of consolidated "title"+"summary"+"inclusion_criteria" for each CT idi.e. clinical trials(unique "nctId") as chunks. Each of the aformentioned chunks have 4 questions(anchors) branched to it(here 1-to-1 normalized mapping of those). The data is particularly useful in order to fine tune an embedding model for CT domain. This further helps in (i)ranked retreival genesis of CT RAGs… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data.tabular1K<n<10K1 likes57 downloads1mo agoHugging Face13rr4433 /powershell_embedding_model_training_datatext100K<n<1M1 likes52 downloads2y agoHugging Face14LLM-OS-Models /korean-embedding-performance-v1-sionic-squad-train-60k Korean Embedding — Sionic SQuAD train-family 60K KorQuAD v1.0의 원본 train split만 질문→정답 문맥 retrieval 형식으로 변환한 60,000-row target-adaptation 데이터다. Sionic retrieval 9종 중 SQuADKorV1의 train-family 신호를 명시적으로 보강한다. 사용 조건과 점수 공개 방식 release_eligible: false인 performance/non-commercial 실험용 composite다. 이 저장소의 통합 라이선스는 other이며 upstream 권리를 재허가하지 않는다. Hub metadata는 KorQuAD source를 CC-BY-ND-4.0으로 표시하고, upstream dataset card 본문은 CC BY-ND 2.0 KR도 명시한다. 사용자는 원 source 조건을 직접 확인해야 한다. 이… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-squad-train-60k.textsentence-similarity10K<n<100K0 likes44 downloads3mo agoHugging Face15LLM-OS-Models /korean-embedding-performance-v1-sionic-autorag-100k Korean Embedding — Sionic AutoRAG domain 100K AutoRAG의 금융·상거래·법률 domain retrieval을 보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 FIQA/Amazon/Banking77과 중국어 e-commerce/legal QA를 query/positive/negative contrastive schema로 묶었다. 사용 조건과 평가 노출 release_eligible: false인 performance/non-commercial 연구용 composite다. 통합 라이선스는 other이며 F2 collection의 Apache-2.0 표기가 개별 upstream 권리를 재허가하지 않는다. AutoRAG evaluation repository, query, qrel, corpus는 loader 입력으로… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-autorag-100k.textsentence-similarity100K<n<1M0 likes42 downloads3mo agoHugging Face16cometadata /ror-embedding-model-comparisontabular10K<n<100K0 likes38 downloads9mo agoHugging Face17LLM-OS-Models /korean-embedding-performance-v1-sionic-health-100k Korean Embedding — Sionic health multilingual 100K Qwen3-Embedding 계열의 한국어 PublicHealthQA와 multilingual medical retrieval을 보강하기 위한 100,000-row performance dataset이다. F2LLM-v2 collection의 영어 중심 medical QA/instruction/flashcard와 소량 중국어 WebMedQA를 query/positive/negative contrastive schema로 묶었다. 사용 조건 release_eligible: false인 performance/non-commercial 연구용 composite다. 통합 라이선스 표기는 other이며 collection card의 Apache-2.0 표기가 각 upstream source의 권리·개인정보·의료 데이터 조건을 재허가하지 않는다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-sionic-health-100k.textsentence-similarity100K<n<1M0 likes34 downloads3mo agoHugging Face18LLM-OS-Models /korean-embedding-ko-triplet-hn-pilot-10k Korean Embedding Ko-Triplet Hard-Negative Pilot 10K nlpai-lab/ko-triplet-v1.0에서 결정론적으로 뽑은 한국어 retrieval train 10,000행과 validation 512행에 Qwen3-Embedding-8B dense hard negative 4개씩을 붙인 연구용 ms-swift embedding dataset이다. 사용 조건 원 source 카드에 명시적 라이선스가 없어 통합 라이선스는 other, manifest의 release_eligible은 false다. 연구·비상업 성능 실험용이며 이 카드가 원 source의 권리를 재허가하지 않는다. 출처와 sampling source: nlpai-lab/ko-triplet-v1.0 pinned revision: 1f5d72d21ae8309b5221a588b13930b423385bff… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-ko-triplet-hn-pilot-10k.texttext-retrieval10K<n<100K0 likes34 downloads3mo agoHugging Face19HFforLegal /embedding-models Reference models for integration into HF for Legal 🤗 This dataset comprises a collection of models aimed at streamlining and partially automating the embedding process. Each model entry within this dataset includes essential information such as model identifiers, embedding configurations, and specific parameters, ensuring that users can seamlessly integrate these models into their workflows with minimal setup and maximum efficiency. Dataset Structure Field Type… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/embedding-models.tabulartabular-to-textn<1K3 likes33 downloads2y agoHugging Face20Akiya-Vyre /legal-embedding-modeltabularn<1K0 likes32 downloads19d agoHugging Face21LLM-OS-Models2 /ko-legal-embedding-training-v1 Korean Public Legal Embedding Training v1 실제 embedding 학습 queue가 소비하는 provenance-preserving 파생 JSONL이다. source-native query/positive 관계와 deterministic bootstrap negative로 컴파일했다. 최종 학습 전 current-student hard-negative mining을 수행해야 한다. rows: 250,000 release eligible: true visibility: public use: public redistribution and model training exact benchmark query/evaluation-text matches: 0 exact retrieval-corpus matches: 0 unique hashes Sionic 9 및 MTEB task-family train source가 포함될 수… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models2/ko-legal-embedding-training-v1.textsentence-similarity100K<n<1M0 likes20 downloads2mo agoHugging Face22likhitjuttada /ft-embeddingmodel-RAG-dataset Dataset Card for Dataset Name This dataset aims to be a base template for fine-tuning embedding models for enhanced retrieval performance in RAG pipelines. It has been generated locally using Mistral:7B on Ollama using a simple prompt that prompts the model to generate 5 questions for each document chunk of Apple's Environmental Progress Report 2024 Dataset Details Dataset Description Curated by: Likhit Juttada Funded by [optional]: NA Credits… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/ft-embeddingmodel-RAG-dataset.textquestion-answeringn<1K0 likes11 downloads11mo agoHugging Face23Siranjeevi029 /qwen3-0.6b-embedding-model-datasettext1K<n<10K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.