datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot-arena-llm-judges
Chatbot-Arena
https://www.kaggle.com/competitions/lmsys-chatbot-arena/data
Single-turn data: https://huggingface.co/datasets/potsawee/chatbot-arena-llm-judges
#examples = 49938
split: A_win = 17312 (34.67%), B_win = 16985 (34.01%), tie = 15641 (31.32%)
#2-way only examples = 34297 (68.68%)
This repository
train.single-turn.json: data extracted from the train file from LMSys on Kaggle
each example has attributes - id, model_[a, b], winne_model_[a, b, tie], question… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/chatbot-arena-llm-judges.Swallow-MX-chatbot-DPOChatbot Arena Conversationsの質問文から、aixsatoshi/Swallow-MX-8x7b-NVE-chatvector-Mixtral-instruct-v2を使用して応答文を作成しました
質問文は、以下のモデルのPrompt部分を使用しました
Chatbot Arena Conversations JA (calm2)
以下引用です。
指示文(prompt)はlmsys/chatbot_arena_conversationsのユーザ入力(CC-BY 4.0)を和訳したものです。これはChatbot Arenaを通して人間が作成した指示文であり、CC-BY 4.0で公開されているものです。複数ターンの対話の場合は最初のユーザ入力のみを使っています(そのため、このデータセットはすべて1ターンの対話のみになっております)。
和訳にはfacebookの翻訳モデル(MIT License)を使っています。
Ecom-Chatbot-Synthetic-Test-Datasetmeo-chatbot-data
meo-chatbot data
Pre-built RAG data for meo-chatbot — Vietnamese cat advisor chatbot.
Contents
chromadb/ — ChromaDB persistent client with 76,487 chunks (384-dim e5-small embeddings)
chunks/classified.jsonl — raw chunks with metadata (topic, content_type, severity, level)
cleaned/*.jsonl — original articles per source (optional, only if --include-cleaned)
Sources (11 VN cat websites)
pethealth.vn, paddy.vn, tropicpet.vn, mozzi.vn… See the full description on the dataset page: https://huggingface.co/datasets/Monmoonluna/meo-chatbot-data.Mostafa8Mehrabi__llama-3.2-1b-Insomnia-ChatBot-merged-details
Dataset Card for Evaluation run of Mostafa8Mehrabi/llama-3.2-1b-Insomnia-ChatBot-merged
Dataset automatically created during the evaluation run of model Mostafa8Mehrabi/llama-3.2-1b-Insomnia-ChatBot-merged
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Mostafa8Mehrabi__llama-3.2-1b-Insomnia-ChatBot-merged-details.
