datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.ultrachat_de
German UltraChat
This dataset contains the first 1k prompts from HuggingFaceH4/ultrachat_200k translated to German and inference on with GPT-4.
ultrachat_200k_filtered_1707945637
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707945637.ultrachat_200k_filtered_1708035667
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708035667.ultrachat_200k_filtered_1707947544
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707947544.ultrachat-200k-bliss-raw
UltraChat 200k Blissymbolic Raw Transliteration
This dataset is a lexical Blissymbolic transliteration of HuggingFaceH4/ultrachat_200k train_sft using experimental BlissyLM conversion tooling. It preserves the source conversation structure and role metadata while adding Blissymbol token sequences based on BCI Authorized Vocabulary gloss lookup.
This is not a human translation and is not clinical AAC guidance.
BlissyLM is an early research/tooling project for exploring Blissymbol… See the full description on the dataset page: https://huggingface.co/datasets/ifinspire/ultrachat-200k-bliss-raw.UltraChat-300K-SLAM-Omni
UltraChat-300K
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.ultrachat_200k_filtered_1708034814
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708034814.ultrachat-feedback-10k-chatmld2f-turn-embeddings-ultrachat
Turn embeddings for Ultrachat (Dialog2Flow encoder)
One 768-d float16 vector per utterance of openbmb/UltraChat, computed with the frozen
encoder sergioburdisso/dialog2flow-joint-bert-base
(SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder
truncates inputs at 64 tokens. If you use these embeddings, please cite the
Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus.
Files
ultrachat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-ultrachat.ultrachat_200k_filtered_1708454270
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708454270.ultrachat_200k_filtered_1708458397
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708458397.ultrachat_200k_filtered_1708702930
Args
{'base_model': 'EleutherAI/pythia-6.9b-deduped',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n'
'\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708702930.ultrachat_200k_filtered_1710204240
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1710204240.ultrachat_200k_filtered_1710165338
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': False,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1710165338.mm2.7-ultrachat_200kultrachat_200k_filtered_1707919621
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707919621.ultrachat_200k_filtered_1707920811
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707920811.ultrachat_200k_filtered_1707921252
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707921252.ultrachat_200k_filtered_1707920039
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707920039.ultrachat_dpo_sft_deepl_kaannetty
Dataset Card for Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty
This dataset is more filtered down version of Finnish-NLP/ultrafeedback_deepl_sft_dpo_filtered
Creation process
Load data from https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized/viewer/default/train_sft
Do zero shot classification with facebook/bart-large-mnli in this kind of way (Actual implementation might be slightly different):
preds = pipe(f'{row["instruction"]} is a question about:'… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty.lm-eval-results-Magpie-Align-Llama-3-8B-Ultrachat-200K-private
Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-Ultrachat-200K
Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-Ultrachat-200K
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-Ultrachat-200K-private.laguna-xs-ultrachat-responsesultrachat-aem-v2.1from datasets import load_dataset
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("minhbui/viettel_v3.2")
def token_count(example):
conv = example["data"]
first_instruction = conv[0]
first_response = conv[1]
first_instruction_num_tokens = len(tokenizer.encode(first_instruction))
first_response_num_tokens = len(tokenizer.encode(first_response))
result = dict(
first_instruction_num_tokens=first_instruction_num_tokens… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhdo/ultrachat-aem-v2.1.ultrachat_200k_filtered_1708381525
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708381525.ultrachat_200k_filtered_1707919193
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707919193.ultrachat_200k_filtered_1707919115
Dataset Card for "ultrachat_200k_filtered_1707919115"
More Information needed
ultrachat_200k_filtered_1710165106
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=None,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1710165106.ultrachat-4spider-iter2ultrachat_200k_filtered_1707919460
Args
{'base_model': 'mistralai/Mistral-7B-v0.1',
'check_length_correctness': True,
'debug': True,
'hf_entity': 'vwxyzjn',
'params': TaskQueryHParams(length=3000,
format_str='SUBREDDIT: r/{subreddit}\n'
'\n'
'TITLE: {title}\n'
'\n'
'POST: {post}\n''\n'
'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707919460.
