CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K1 likes101 downloads2y agoHugging Face02dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.tabularquestion-answeringn<1K0 likes91 downloads2y agoHugging Face03imhmdf /taiwan-culture-llama-70b-training Taiwan Culture Llama 70B Training Dataset Dataset Description This dataset is specifically prepared for fine-tuning Llama 70B models on Taiwan culture, combining multiple high-quality sources to create a comprehensive training resource. Dataset Summary Total Samples: 8352 Language: Traditional Chinese (Taiwan) Domains: Taiwan Culture, Government QA, Conversational AI, Cultural Knowledge Format: Instruction-following format optimized for chat models… See the full description on the dataset page: https://huggingface.co/datasets/imhmdf/taiwan-culture-llama-70b-training.textquestion-answering1K<n<10K0 likes78 downloads8mo agoHugging Face04dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K0 likes76 downloads2y agoHugging Face05lmarena-ai /Llama-3-70b-battlesChatbot Arena user conversations between Llama-3-70b VS GPT-4-1025 or Llama-3-70b VS Claude-3-Opus with user preference votes. Single turn. Excludes ties. Used in Llama Data Analysis blog post and "VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models" (Paper, Code). Citation @article{dunlap_vibecheck, title={VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models}, author={Lisa Dunlap and Krishna Mandal and Trevor… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/Llama-3-70b-battles.textquestion-answering1K<n<10K3 likes68 downloads2y agoHugging Face06nicher92 /magpie_llama70b_200k_filtered_swedish Short description Roughly 200k filtered instruction : response pairs in Swedish, filtered from roughtly 550k. Contains "normal" QA along with math and coding QA. Filtering, removed: Deduplications Instructions scored less than good or excellent Responses scored less than -10 from ArmoRM-Llama3-8B-v0.1 Instructions and responses less than 10 in length or more than 2048 Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/nicher92/magpie_llama70b_200k_filtered_swedish.tabularquestion-answering100K<n<1M2 likes58 downloads2y agoHugging Face07dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192textquestion-answeringn<1K1 likes50 downloads2y agoHugging Face08nicher92 /magpie_llama70b_260k_filtered_swedish Short description Roughly 260k filtered instruction : response pairs in Swedish, filtered from roughtly 650k. Contains "normal" QA along with math and coding QA and multiple choice questions and answers. Filtering, removed: Deduplications Instructions scored less than good or excellent Responses scored less than -10 from ArmoRM-Llama3-8B-v0.1 Instructions and responses less than 10 in length or more than 2048 Usage from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/nicher92/magpie_llama70b_260k_filtered_swedish.tabularquestion-answering100K<n<1M0 likes42 downloads2y agoHugging Face09nicher92 /magpie_llama70b_100k_swedish Dataset Card Llama 3.3 generated instruction: response pairs using the MagPie pipeline: https://github.com/magpie-align/magpie/ https://arxiv.org/abs/2406.08464 Larger and filtered dataset available here: nicher92/magpie_llama70b_200k_filtered_swedish Dataset Details Dataset Description System prompt template for generating instructions: "<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nDu är en hjälpsam AI… See the full description on the dataset page: https://huggingface.co/datasets/nicher92/magpie_llama70b_100k_swedish.tabularquestion-answering100K<n<1M0 likes26 downloads2y agoHugging Face10imhmdf /taiwan-culture-llama-70b-training-v2 Taiwan Culture Llama 70B Training Dataset Dataset Description This dataset is specifically prepared for fine-tuning Llama 70B models on Taiwan culture, combining multiple high-quality sources to create a comprehensive training resource. Dataset Summary Total Samples: 8352 Language: Traditional Chinese (Taiwan) Domains: Taiwan Culture, Government QA, Conversational AI, Cultural Knowledge Format: Instruction-following format optimized for chat models… See the full description on the dataset page: https://huggingface.co/datasets/imhmdf/taiwan-culture-llama-70b-training-v2.textquestion-answering1K<n<10K0 likes26 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.