CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01worstchan /UltraChat-300K-SLAM-Omni UltraChat-300K This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/UltraChat-300K-SLAM-Omni.tabularquestion-answering100K<n<1M2 likes3.9k downloads1y agoHugging Face02bjoernp /ultrachat_de German UltraChat This dataset contains the first 1k prompts from HuggingFaceH4/ultrachat_200k translated to German and inference on with GPT-4. tabularn<1K11 likes356 downloads3y agoHugging Face03vwxyzjn /ultrachat_200k_filtered_1707945637 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707945637.tabular100K<n<1M0 likes295 downloads3y agoHugging Face04vwxyzjn /ultrachat_200k_filtered_1708035667 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708035667.tabular100K<n<1M0 likes258 downloads3y agoHugging Face05vwxyzjn /ultrachat_200k_filtered_1707947544 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707947544.tabular100K<n<1M0 likes226 downloads3y agoHugging Face06ifinspire /ultrachat-200k-bliss-raw UltraChat 200k Blissymbolic Raw Transliteration This dataset is a lexical Blissymbolic transliteration of HuggingFaceH4/ultrachat_200k train_sft using experimental BlissyLM conversion tooling. It preserves the source conversation structure and role metadata while adding Blissymbol token sequences based on BCI Authorized Vocabulary gloss lookup. This is not a human translation and is not clinical AAC guidance. BlissyLM is an early research/tooling project for exploring Blissymbol… See the full description on the dataset page: https://huggingface.co/datasets/ifinspire/ultrachat-200k-bliss-raw.tabulartext-generation100K<n<1M0 likes222 downloads4mo agoHugging Face07mwei /UltraChat-300K-SLAM-Omni UltraChat-300K This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round English spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/mwei/UltraChat-300K-SLAM-Omni.tabularquestion-answering100K<n<1M0 likes208 downloads8mo agoHugging Face08vwxyzjn /ultrachat_200k_filtered_1708034814 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708034814.tabular100K<n<1M0 likes181 downloads3y agoHugging Face09smangrul /ultrachat-feedback-10k-chatmltabular10K<n<100K1 likes69 downloads3y agoHugging Face10jumafernandez /d2f-turn-embeddings-ultrachat Turn embeddings for Ultrachat (Dialog2Flow encoder) One 768-d float16 vector per utterance of openbmb/UltraChat, computed with the frozen encoder sergioburdisso/dialog2flow-joint-bert-base (SentenceTransformer recipe, convert_to_numpy, no normalization). The encoder truncates inputs at 64 tokens. If you use these embeddings, please cite the Dialog2Flow paper (Burdisso et al., EMNLP 2024) and the source corpus. Files ultrachat_e_t.f16.npy — numpy array (n_turns… See the full description on the dataset page: https://huggingface.co/datasets/jumafernandez/d2f-turn-embeddings-ultrachat.tabularn<1K0 likes67 downloads2mo agoHugging Face11vwxyzjn /ultrachat_200k_filtered_1708454270 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708454270.tabular10K<n<100K0 likes57 downloads3y agoHugging Face12vwxyzjn /ultrachat_200k_filtered_1708458397 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708458397.tabular10K<n<100K0 likes57 downloads3y agoHugging Face13vwxyzjn /ultrachat_200k_filtered_1708702930 Args {'base_model': 'EleutherAI/pythia-6.9b-deduped', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n' '\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708702930.tabular10K<n<100K0 likes52 downloads3y agoHugging Face14vwxyzjn /ultrachat_200k_filtered_1710204240 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1710204240.tabular10K<n<100K0 likes51 downloads3y agoHugging Face15vwxyzjn /ultrachat_200k_filtered_1710165338 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': False, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1710165338.tabular10K<n<100K0 likes49 downloads3y agoHugging Face16danikhan632 /mm2.7-ultrachat_200ktabular100K<n<1M0 likes34 downloads4mo agoHugging Face17vwxyzjn /ultrachat_200k_filtered_1707919621 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707919621.tabular1K<n<10K0 likes33 downloads3y agoHugging Face18vwxyzjn /ultrachat_200k_filtered_1707920811 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707920811.tabular1K<n<10K0 likes32 downloads3y agoHugging Face19vwxyzjn /ultrachat_200k_filtered_1707921252 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707921252.tabular1K<n<10K0 likes29 downloads3y agoHugging Face20vwxyzjn /ultrachat_200k_filtered_1707920039 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707920039.tabular1K<n<10K0 likes28 downloads3y agoHugging Face21Finnish-NLP /ultrachat_dpo_sft_deepl_kaannetty Dataset Card for Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty This dataset is more filtered down version of Finnish-NLP/ultrafeedback_deepl_sft_dpo_filtered Creation process Load data from https://huggingface.co/datasets/HuggingFaceH4/ultrafeedback_binarized/viewer/default/train_sft Do zero shot classification with facebook/bart-large-mnli in this kind of way (Actual implementation might be slightly different): preds = pipe(f'{row["instruction"]} is a question about:'… See the full description on the dataset page: https://huggingface.co/datasets/Finnish-NLP/ultrachat_dpo_sft_deepl_kaannetty.tabular10K<n<100K1 likes25 downloads10mo agoHugging Face22nyu-dice-lab /lm-eval-results-Magpie-Align-Llama-3-8B-Ultrachat-200K-private Dataset Card for Evaluation run of Magpie-Align/Llama-3-8B-Ultrachat-200K Dataset automatically created during the evaluation run of model Magpie-Align/Llama-3-8B-Ultrachat-200K The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Magpie-Align-Llama-3-8B-Ultrachat-200K-private.tabular100K<n<1M0 likes25 downloads2y agoHugging Face23inference-optimization /laguna-xs-ultrachat-responsestabular100K<n<1M0 likes23 downloads5mo agoHugging Face24nguyenthanhdo /ultrachat-aem-v2.1from datasets import load_dataset from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("minhbui/viettel_v3.2") def token_count(example): conv = example["data"] first_instruction = conv[0] first_response = conv[1] first_instruction_num_tokens = len(tokenizer.encode(first_instruction)) first_response_num_tokens = len(tokenizer.encode(first_response)) result = dict( first_instruction_num_tokens=first_instruction_num_tokens… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhdo/ultrachat-aem-v2.1.tabular10K<n<100K0 likes20 downloads3y agoHugging Face25vwxyzjn /ultrachat_200k_filtered_1708381525 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1708381525.tabular1K<n<10K0 likes19 downloads3y agoHugging Face26vwxyzjn /ultrachat_200k_filtered_1707919193 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707919193.tabular1K<n<10K0 likes15 downloads3y agoHugging Face27vwxyzjn /ultrachat_200k_filtered_1707919115 Dataset Card for "ultrachat_200k_filtered_1707919115" More Information needed tabular1K<n<10K0 likes11 downloads3y agoHugging Face28vwxyzjn /ultrachat_200k_filtered_1710165106 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=None, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1710165106.tabularn<1K0 likes10 downloads3y agoHugging Face29lhkhiem28 /ultrachat-4spider-iter2tabular10K<n<100K0 likes10 downloads1y agoHugging Face30vwxyzjn /ultrachat_200k_filtered_1707919460 Args {'base_model': 'mistralai/Mistral-7B-v0.1', 'check_length_correctness': True, 'debug': True, 'hf_entity': 'vwxyzjn', 'params': TaskQueryHParams(length=3000, format_str='SUBREDDIT: r/{subreddit}\n' '\n' 'TITLE: {title}\n' '\n' 'POST: {post}\n''\n' 'TL;DR:'… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/ultrachat_200k_filtered_1707919460.tabular1K<n<10K0 likes8 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.