datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
everyday-conversations-llama3.1-2k
Everyday conversations for Smol LLMs finetunings
This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic.
The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include:
20 everyday topics with 100 subtopics each
43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.everyday-conversations-tur
Everyday Turkish Conversations
This dataset has everyday conversations in Turkish between user and assistant on various topics. It is inspired by the HuggingFaceTB/everyday-conversations-llama3.1-2k.
License
This dataset is released under the Apache 2.0 License.
everyday-conversations-ita
🇮🇹💬 Everyday Italian Conversations
Inspired by the dataset HuggingFaceTB/everyday-conversations-llama3.1-2k, we generated conversations using the same topics, subtopics, and sub-subtopics as those in the HuggingFaceTB dataset.We slightly adjusted the prompt to produce structured data outputs using Qwen/Qwen2.5-7B-Instruct. Subsequently, we also used the "user" role messages as prompts for google/gemma-2-9b-it.
The result is a dataset of approximately 4.5k… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/everyday-conversations-ita.everyday-conversations-llama3.1-2k-in-french
Description
French translation of HuggingFaceTB/everyday-conversations-llama3.1-2k.
The original dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. The LLM have to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic.
The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include:
20 everyday topics with 100 subtopics… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/everyday-conversations-llama3.1-2k-in-french.everyday-conversations-llama3.1-2k-frencheveryday-conversations-fusedHinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/VaniAgent/Hinglish-Everyday-Conversations-1M.viking_everyday_conversations_complete_volume1
Dataset Card for Viking Everyday Conversations Complete Volume 1
Viking Everyday Conversations: Volume 1
Dataset Overview
This dataset, "Viking Everyday Conversations - Complete Volume 1," is a curated collection of immersive, dialogue-based snippets capturing the essence of daily life in a Norse-inspired world. Presented in a JSONL format for easy parsing and use in language models, NLP tasks, or creative writing tools, it features 1000 entries of paired… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/viking_everyday_conversations_complete_volume1.everyday-conversations-gpt-oss-20b-itaHinglish-Everyday-Conversations-1M-Devanagariit-everyday-conversations-llama3.1-2k-TowerInstruct-Mistral-7B-v0.2
sapienzanlp/it-everyday-conversations-llama3.1-2k-TowerInstruct-Mistral-7B-v0.2
Overview
This dataset is the Italian translation of the everyday-conversations-llama3.1-2k dataset, designed specifically for conversations in Italian. The translation was carried out using TowerInstruct-Mistral-7B-v0.2.
Languages: Italian (translated from English)
Purpose: Instruction tuning in Italian for conversational AI
Train Size: 2260
Test Size: 119
New fields added include:
{… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/it-everyday-conversations-llama3.1-2k-TowerInstruct-Mistral-7B-v0.2.everyday_conversations_ja
データセットについて
このデータセットは、 HuggingFaceTB/everyday-conversations-llama3.1-2k を機械翻訳で日本語化したものになります。
具体的には、everyday-conversations-llama3.1-2kをトピックごとの対話のペアに変更してDeepLで翻訳したものとなります。
詳細
topic: everyday-conversations-llama3.1-2kのtopic
user: 各トピックごとのユーザーからの発話
assistant: 各トピックごとのユーザーへの返答
assistantの返答がない場合はNone
注意事項
人手で若干修正をしましたが、日本語が変な箇所がいくつか散見されます。
ライセンス:apache 2.0
Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/buggiebug/Hinglish-Everyday-Conversations-1M.Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/a1b8h04i/Hinglish-Everyday-Conversations-1M.lbourdois-Llama-3.2-1B-Instruct-everyday-conversations-llama3.1-2k-in-french-with-dataforgehomer-simpson-smoltalk-everyday-conversationsBeloved-Everyday-ConversationsOriginal :
everyday-conversations-llama3.1-2k (modified)
Now D1rtyB1rd/Beloved-Everyday-Conversations'
is a collection of multi-turn dialogues between users and an AI assistant. The assistant now takes on a more personal and friendly tone in its responses, resembling a close, conversational style. For example, the assistant addresses users in an intimate, anticipatory manner, as seen in a greeting like "Hello, cherished. I’ve taken the liberty to anticipate your requirements." and "Hello!… See the full description on the dataset page: https://huggingface.co/datasets/D1rtyB1rd/Beloved-Everyday-Conversations.Hinglish-Everyday-Conversations-AJeveryday-conversations-step-by-stepeveryday-conversationssmoltalk2-everyday-conversations-no-think-pt-ptwraps-everyday-conversations-llama3.1-2k-frenchCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données huggingface.co/datasets/wraps/everyday-conversations-llama3.1-2k-french.
Iris-everyday-conversations
