datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ultrachat_200k
Dataset Card for UltraChat 200k
Dataset Description
This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model.
The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic:
Selection of a subset of data for faster supervised fine tuning.
Truecasing of the dataset, as we observed around 5% of… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k.Llama3.1-8B-BaldEagle3-Ultrachatultrachat-10k-chatmlUltraChat
Dataset Card for Dataset Name
Dataset Description
An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts.
To ensure generation quality, two separate ChatGPT Turbo APIs are adopted in generation, where one plays the role of the user to generate queries and the other generates the response.
We instruct the user model with… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraChat.UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.ultrachat-sharegpt-5GBQwen2.5-7B-BaldEagle-Ultrachatpirate-ultrachat-10kUltrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized
Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized"
More Information needed
ultrachat
Dataset Card for "ultrachat"
More Information needed
ultrachat_speech_multiTurnsUltraChat-Mixin
Dataset Card for "UltraChat-Mixin"
UltraChat-Mixin Dataset
Overview
UltraChat-Mixin is a dataset created by Me, which is a mix of three datasets: 'stingning/ultrachat', 'jondurbin/airoboros-2.1', and 'erfanzar/GPT4-8K'. This dataset is designed for training conversational AI models.
Dataset Configuration
The dataset is configured as follows:
configs:
- config_name: default
data_files:
- split: train
path: data/train-*… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/UltraChat-Mixin.ultrachat_200k_sftultrachat_de
German UltraChat
This dataset contains the first 1k prompts from HuggingFaceH4/ultrachat_200k translated to German and inference on with GPT-4.
ultrachat
UltraChat Conversations Dataset
This dataset contains 1,468,346 multi-turn conversations from UltraChat, processed to preserve the original conversational structure and optimized for training conversational AI models.
🎯 Dataset Format
Each conversation record contains:
id: Sequential conversation ID (1, 2, 3, ...)
source: "ultra"
language: "english"
data: JSON string containing conversation turns array
📊 Dataset Statistics
Total Conversations: 1,468,346… See the full description on the dataset page: https://huggingface.co/datasets/metythorn/ultrachat.Ultrachat-Multiple-Conversations-Alpaca-Style
Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Style"
More Information needed
UltrachatBR
UltrachatBR: Um Dataset em Português baseado no Ultrachat
O UltrachatBR é uma versão em português do conhecido dataset Ultrachat, originalmente desenvolvido para o idioma inglês. Este projeto visa disponibilizar uma vasta coleção de diálogos traduzidos para o português, ampliando assim o acesso a recursos de processamento de linguagem natural para a comunidade de língua portuguesa.
Processo de Tradução
O processo de tradução foi realizado utilizando a API do Google… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/UltrachatBR.ultrachat_200k_filtered_1707945637
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707945637.ultrachat_filtered_0.9
Dataset Card for "ultrachat_filtered_0.9"
More Information needed
ultrachat_filtered_0.95
Dataset Card for "ultrachat_filtered_0.95"
More Information needed
Malaysian-Ultrachat
Ultrachat like using Malaysian context
Prepare multiturn dialogue between user and assistant for malaysian context,
Astroawani, https://huggingface.co/datasets/malaysia-ai/crawl-astroawani, ultrachat-astroawani-malay.jsonl, 60198 rows, 477 MB.
Crossref melayu papers, https://huggingface.co/datasets/mesolitica/crawl-my-website/resolve/main/melayu-pdf.jsonl, ultrachat-crossref-melayu-malay.jsonl, 9959 rows, 187 MB
Epenerbitan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Ultrachat.ultrachat_200k_filtered_1708035667
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708035667.ultrachat_200k_filtered_1707947544
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707947544.ro_sft_ultrachat
Dataset Description
Ultrachat is an open-source, large-scale, and multi-round dialogue dataset.
Here we provide the Romanian translation of the UltraChat dataset, translated with Systran.
This dataset is part of the instruction finetune protocol for Romanian LLMs proposed in "Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions (Masala et al., 2024).
Citation
@article{ding2023enhancing,
title={Enhancing Chat Language Models… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_ultrachat.ultrachat-200k-bliss-raw
UltraChat 200k Blissymbolic Raw Transliteration
This dataset is a lexical Blissymbolic transliteration of HuggingFaceH4/ultrachat_200k train_sft using experimental BlissyLM conversion tooling. It preserves the source conversation structure and role metadata while adding Blissymbol token sequences based on BCI Authorized Vocabulary gloss lookup.
This is not a human translation and is not clinical AAC guidance.
BlissyLM is an early research/tooling project for exploring Blissymbol… See the full description on the dataset page: https://huggingface.co/datasets/ifinspire/ultrachat-200k-bliss-raw.ultrachat_questions_about_world
Ultrachat, Questions about the world
This is the "Questions about the world" subset of UltraChat, found in the this GitHub repo.
UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.ultrachat_200k
Dataset Card for "ultrachat_200k"
More Information needed
ultrachat_200k_dutch
Dataset Card for UltraChat 200k Dutch
Citation
If you use this dataset, GEITje 7B Ultra (SFT) or any of its derivatives or quantizations, place cite the following paper:
@misc{vanroy2024geitje7bultraconversational,
title={GEITje 7B Ultra: A Conversational Model for Dutch},
author={Bram Vanroy},
year={2024},
eprint={2412.04092},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.04092},
}… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/ultrachat_200k_dutch.HuggingFaceH4-ultrachat_200k
