embedding-model
turkish_embedding_model_training_datacleaned_turkish_embedding_model_training_data_colabcleaned_turkish_embedding_model_training_data_colab
Citation
If you use this dataset in your research, please cite the following paper:
@inproceedings{baysan-gungor-2025-tr,
title = "{TR}-{MTEB}: A Comprehensive Benchmark and Embedding Model Suite for {T}urkish Sentence Representations",
author = "Baysan, Mehmet Selman and
Gungor, Tunga",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2025",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher =… See the full description on the dataset page: https://huggingface.co/datasets/trmteb/cleaned_turkish_embedding_model_training_data_colab.Embedding-model-fine-tuning-datasetturkish_embedding_model_training_dataClinical_trials_anchor-positive-pairs_EmbeddingModel-data_final
Dataset details:-
This dataset is the final version of anchor(query)-positive(chunk) pair data w.r.t fine tuning embedding model for clinical trials dataset.
It includes best of both 4 anchors-consolidated positive chunk/nctId dataset-->first dataset and
5 anchors-3 positive chunk/nctId--->Second dataset.
The 1st dataset(consolidated title +summary+ inclusion criteria chunk) suffered with pre-processing bottlenecks :-
rendering huge chunks upto 15k characeters.
missing on… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final.
