CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /ultrachat_200k Dataset Card for UltraChat 200k Dataset Description This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model. The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic: Selection of a subset of data for faster supervised fine tuning. Truecasing of the dataset, as we observed around 5% of… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k.texttext-generation100K<n<1M956 likes97k downloads2y agoHugging Face02smangrul /ultrachat-10k-chatmltext10K<n<100K6 likes5.3k downloads3y agoHugging Face03openbmb /UltraChat Dataset Card for Dataset Name Dataset Description An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts. To ensure generation quality, two separate ChatGPT Turbo APIs are adopted in generation, where one plays the role of the user to generate queries and the other generates the response. We instruct the user model with… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraChat.texttext-generation100K<n<1M505 likes3.9k downloads3y agoHugging Face04worstchan /UltraChat-300K-SLAM-Omni UltraChat-300K This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.tabularquestion-answering100K<n<1M2 likes3.9k downloads1y agoHugging Face05typeof /ultrachat-sharegpt-5GBtext100K<n<1M0 likes3k downloads3y agoHugging Face06health360 /Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized" More Information needed text1M<n<10M0 likes663 downloads3y agoHugging Face07causal-lm /ultrachat Dataset Card for "ultrachat" More Information needed text1M<n<10M4 likes631 downloads3y agoHugging Face08JerryAGENDD /ultrachat_speech_multiTurnstext10K<n<100K0 likes600 downloads2y agoHugging Face09erfanzar /UltraChat-Mixin Dataset Card for "UltraChat-Mixin" UltraChat-Mixin Dataset Overview UltraChat-Mixin is a dataset created by Me, which is a mix of three datasets: 'stingning/ultrachat', 'jondurbin/airoboros-2.1', and 'erfanzar/GPT4-8K'. This dataset is designed for training conversational AI models. Dataset Configuration The dataset is configured as follows: configs: - config_name: default data_files: - split: train path: data/train-*… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/UltraChat-Mixin.textsummarization1M<n<10M6 likes561 downloads3y agoHugging Face10mlabonne /ultrachat_200k_sfttext100K<n<1M3 likes372 downloads2y agoHugging Face11bjoernp /ultrachat_de German UltraChat This dataset contains the first 1k prompts from HuggingFaceH4/ultrachat_200k translated to German and inference on with GPT-4. tabularn<1K11 likes356 downloads3y agoHugging Face12metythorn /ultrachat UltraChat Conversations Dataset This dataset contains 1,468,346 multi-turn conversations from UltraChat, processed to preserve the original conversational structure and optimized for training conversational AI models. 🎯 Dataset Format Each conversation record contains: id: Sequential conversation ID (1, 2, 3, ...) source: "ultra" language: "english" data: JSON string containing conversation turns array 📊 Dataset Statistics Total Conversations: 1,468,346… See the full description on the dataset page: https://huggingface.co/datasets/metythorn/ultrachat.textquestion-answering1M<n<10M0 likes354 downloads1y agoHugging Face13health360 /Ultrachat-Multiple-Conversations-Alpaca-Style Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Style" More Information needed text1M<n<10M2 likes344 downloads3y agoHugging Face14recogna-nlp /UltrachatBR UltrachatBR: Um Dataset em Português baseado no Ultrachat O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa. Processo de Tradução O processo de tradução foi realizado utilizando a API do Google… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/UltrachatBR.texttext-generation100K<n<1M15 likes330 downloads3y agoHugging Face15vwxyzjn /ultrachat_200k_filtered_1707945637 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707945637.tabular100K<n<1M0 likes295 downloads3y agoHugging Face16pkarypis /ultrachat_filtered_0.9 Dataset Card for "ultrachat_filtered_0.9" More Information needed text100K<n<1M0 likes294 downloads3y agoHugging Face17pkarypis /ultrachat_filtered_0.95 Dataset Card for "ultrachat_filtered_0.95" More Information needed text100K<n<1M0 likes290 downloads3y agoHugging Face18vwxyzjn /ultrachat_200k_filtered_1708035667 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708035667.tabular100K<n<1M0 likes258 downloads3y agoHugging Face19vwxyzjn /ultrachat_200k_filtered_1707947544 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707947544.tabular100K<n<1M0 likes226 downloads3y agoHugging Face20ifinspire /ultrachat-200k-bliss-raw UltraChat 200k Blissymbolic Raw Transliteration This dataset is a lexical Blissymbolic transliteration of HuggingFaceH4/ultrachat_200k train_sft using experimental BlissyLM conversion tooling. It preserves the source conversation structure and role metadata while adding Blissymbol token sequences based on BCI Authorized Vocabulary gloss lookup. This is not a human translation and is not clinical AAC guidance. BlissyLM is an early research/tooling project for exploring Blissymbol… See the full description on the dataset page: https://huggingface.co/datasets/ifinspire/ultrachat-200k-bliss-raw.tabulartext-generation100K<n<1M0 likes222 downloads4mo agoHugging Face21HuggingFaceTB /ultrachat_questions_about_world Ultrachat, Questions about the world This is the "Questions about the world" subset of UltraChat, found in the this GitHub repo. text100K<n<1M7 likes210 downloads3y agoHugging Face22mwei /UltraChat-300K-SLAM-Omni UltraChat-300K This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.tabularquestion-answering100K<n<1M0 likes208 downloads8mo agoHugging Face23rishiraj /ultrachat_200k Dataset Card for "ultrachat_200k" More Information needed text100K<n<1M1 likes186 downloads3y agoHugging Face24BramVanroy /ultrachat_200k_dutch Dataset Card for UltraChat 200k Dutch Citation If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper: @misc{vanroy2024geitje7bultraconversational, title={GEITje 7B Ultra: A Conversational Model for Dutch}, author={Bram Vanroy}, year={2024}, eprint={2412.04092}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2412.04092}, }… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultrachat_200k_dutch.texttext-generation100K<n<1M8 likes186 downloads2y agoHugging Face25manishiitg /HuggingFaceH4-ultrachat_200ktext100K<n<1M1 likes186 downloads3y agoHugging Face26vwxyzjn /ultrachat_200k_filtered_1708034814 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708034814.tabular100K<n<1M0 likes181 downloads3y agoHugging Face27Aditya149 /ultrachattext1M<n<10M1 likes163 downloads3y agoHugging Face28apurvagup /ultrachat_hindi_seamless Dataset Card for "ultrachat_hindi_seamless" More Information needed text100K<n<1M0 likes161 downloads3y agoHugging Face29erfanzar /UltraChat-Mini Dataset Card for "UltraChat-Mini" More Information needed text100K<n<1M0 likes158 downloads3y agoHugging Face30masakhane /african-ultrachattext10K<n<100K5 likes148 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.