ultrachat
LLaMA3.1-8B-Instruct-DFlash-UltraChatneuralmagic_-_Llama-2-7b-ultrachat200k-ggufkykim0_-_Llama-2-7b-ultrachat-code-ggufllama2.c-stories15M-ultrachat-mixed-uncompressedLLaMAntino-2-chat-13b-hf-UltraChat-ITALlama-3-8B-4bit-UltraChat-Itakykim0_-_Llama-2-7b-ultrachat-syn-ggufLlama-3-8B-Ultrachat-200K-i1-GGUF
ultrachat_200k
Dataset Card for UltraChat 200k
Dataset Description
This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model.
The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic:
Selection of a subset of data for faster supervised fine tuning.
Truecasing of the dataset, as we observed around 5% of… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k.Llama3.1-8B-BaldEagle3-Ultrachatultrachat-10k-chatmlUltraChat
Dataset Card for Dataset Name
Dataset Description
An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts.
To ensure generation quality, two separate ChatGPT Turbo APIs are adopted in generation, where one plays the role of the user to generate queries and the other generates the response.
We instruct the user model with… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraChat.UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.ultrachat-sharegpt-5GB
