CoolFace
18 results

llm dataset

tascib /turkish-llm-dataset Turkish Pretraining Corpus Dataset Description This dataset is a Turkish pretraining corpus created by combining BellaTurca (excluding ForumSohbetleri), Cosmos-Turkish-Corpus-v1.0, and FineWeb-2 Turkish Categorized, followed by cleaning, normalization, and deduplication. It is intended for the development, training, and evaluation of Turkish language models. This dataset was prepared as part of a capstone project conducted by a group of students from Sabancı… See the full description on the dataset page: https://huggingface.co/datasets/tascib/turkish-llm-dataset.text100M<n<1B15 likes11k downloads5mo agoHugging Facebitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes8.2k downloads2y agoHugging FaceGunulhona /llm_datasetstexttext-generation100K<n<1M0 likes7.1k downloads3y agoHugging FaceLLM-LAT /harmful-datasettext1K<n<10K43 likes3.4k downloads2y agoHugging FaceAtesiT /ru-llm-judge-dataset RU-LLM-Judge-Dataset Датасет собирается инкрементально в течение нескольких сессий бесплатного Google Colab. Текущий объём: 18,821 суждений (по состоянию на последний запуск). Прогресс к цели (5,000 суждений) [████████████████████] 100% (18,821 / 5,000) История сессий сбора Сессия Дата Добавлено Итого 1 2026-08-05 08:42 617 617 2 2026-08-06 14:40 583 1,200 3 2026-08-07 19:20 486 1,686 4 2026-08-08 22:34 868 2,554 5 2026-08-13… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-llm-judge-dataset.text10K<n<100K0 likes3k downloads2d agoHugging Facebitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.6k downloads2y agoHugging Face