CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
013dllm /MMScan-betatext1M<n<10M1 likes4.3k downloads2y agoHugging Face02HPAI-BSC /Aloe-Beta-Medical-Collection Aloe-Beta-Medical-Collection Collection of curated datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available medical instruction tuning data sources (QA format). Most data samples correspond to single-turn QA pairs, while a small proportion contain multi-turn. All data sources are publicly available for… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-Medical-Collection.textquestion-answering100K<n<1M4 likes171 downloads1y agoHugging Face03HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes154 downloads10mo agoHugging Face04open-llm-leaderboard /HuggingFaceH4__zephyr-7b-beta-detailsgated Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.tabular10K<n<100K0 likes117 downloads2y agoHugging Face05VortexSamples /ReverseBass-Beta-Statustextn<1K0 likes117 downloads22d agoHugging Face06HPAI-BSC /Aloe-Beta-DPO Aloe-Beta-Medical-Collection Collection of curated DPO datasets used to align Aloe-Beta. Dataset Details Dataset Description The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data: Medical preference data: TsinghuaC3I/UltraMedical-Preference General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.textquestion-answering100K<n<1M2 likes90 downloads1y agoHugging Face07nyu-dice-lab /lm-eval-results-HuggingFaceH4-mistral-7b-sft-beta-private Dataset Card for Evaluation run of HuggingFaceH4/mistral-7b-sft-beta Dataset automatically created during the evaluation run of model HuggingFaceH4/mistral-7b-sft-beta The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-HuggingFaceH4-mistral-7b-sft-beta-private.tabular100K<n<1M0 likes44 downloads2y agoHugging Face08TianfuXinqu /github_fetch_huggingface_terminal_9091_n3v8x2_source_beta Beta Support Conversations Anonymized customer support conversation transcripts. Dataset ID: SRC-BETA Catalog: ghfht9091n3v8x2 Origin: Community tech-support forum public dump (2022-2024) Records: 8,152 License: Apache-2.0 textn<1K0 likes44 downloads1mo agoHugging Face09open-llm-leaderboard /EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face10open-llm-leaderboard /cpayne1303__llama-43m-beta-detailsgated Dataset Card for Evaluation run of cpayne1303/llama-43m-beta Dataset automatically created during the evaluation run of model cpayne1303/llama-43m-beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cpayne1303__llama-43m-beta-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face11open-llm-leaderboard /CausalLM__34b-beta-detailsgated Dataset Card for Evaluation run of CausalLM/34b-beta Dataset automatically created during the evaluation run of model CausalLM/34b-beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__34b-beta-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face12open-llm-leaderboard /HPAI-BSC__Qwen2.5-Aloe-Beta-7B-detailsgated Dataset Card for Evaluation run of HPAI-BSC/Qwen2.5-Aloe-Beta-7B Dataset automatically created during the evaluation run of model HPAI-BSC/Qwen2.5-Aloe-Beta-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HPAI-BSC__Qwen2.5-Aloe-Beta-7B-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face13beta42ZH /Test Scaling Synthetic Data Creation with 1,000,000,000 Personas This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas: We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/beta42ZH/Test.texttext-generation100K<n<1M0 likes37 downloads9mo agoHugging Face14DataPilot /AItuber-Realworld-Data-beta1 AItuber Realworld Chat Dataset 概要 本データセットは、AItuber(AI VTuber)の 配信チャット会話データ を合成的に生成したものです。AItuberペルソナデータと nvidia/Nemotron-Personas-Japan のユーザーペルソナを入力に、複数の視聴者がコメントし、AItuberが応答する4ターンのリアルな配信チャットデータを生成しています。生成にはSDG-Nexusという合成データ生成パイプラインを用いました。(sdg-nexus) データの説明 項目 内容 件数 1,946件 形式 JSONL(1行1JSON) 言語 日本語 生成日 2026年3月 ライセンス odc-by ( Open Data Commons Attribution License )… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/AItuber-Realworld-Data-beta1.text1K<n<10K1 likes28 downloads7mo agoHugging Face15open-llm-leaderboard /utkmst__chimera-beta-test2-lora-merged-detailsgated Dataset Card for Evaluation run of utkmst/chimera-beta-test2-lora-merged Dataset automatically created during the evaluation run of model utkmst/chimera-beta-test2-lora-merged The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/utkmst__chimera-beta-test2-lora-merged-details.tabular10K<n<100K1 likes26 downloads2y agoHugging Face16salimon /betatext1K<n<10K0 likes17 downloads3y agoHugging Face17open-llm-leaderboard /HPAI-BSC__Llama3.1-Aloe-Beta-8B-detailsgated Dataset Card for Evaluation run of HPAI-BSC/Llama3.1-Aloe-Beta-8B Dataset automatically created during the evaluation run of model HPAI-BSC/Llama3.1-Aloe-Beta-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HPAI-BSC__Llama3.1-Aloe-Beta-8B-details.tabular10K<n<100K0 likes15 downloads2y agoHugging Face18open-llm-leaderboard /Replete-AI__Replete-LLM-Qwen2-7b_Beta-Preview-detailsgated Dataset Card for Evaluation run of Replete-AI/Replete-LLM-Qwen2-7b_Beta-Preview Dataset automatically created during the evaluation run of model Replete-AI/Replete-LLM-Qwen2-7b_Beta-Preview The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Replete-AI__Replete-LLM-Qwen2-7b_Beta-Preview-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face19Jklee0717 /llama3-dpo-data-beta0p2-LLaMA3_iter1-seed43text10K<n<100K0 likes12 downloads10mo agoHugging Face20betaprogramming /audit-reports-simtextn<1K0 likes11 downloads3mo agoHugging Face21open-llm-leaderboard /dzakwan__dzakwan-MoE-4x7b-Beta-detailsgated Dataset Card for Evaluation run of dzakwan/dzakwan-MoE-4x7b-Beta Dataset automatically created during the evaluation run of model dzakwan/dzakwan-MoE-4x7b-Beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dzakwan__dzakwan-MoE-4x7b-Beta-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face22open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face23open-llm-leaderboard /Sakalti__light-7b-beta-detailsgated Dataset Card for Evaluation run of Sakalti/light-7b-beta Dataset automatically created during the evaluation run of model Sakalti/light-7b-beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__light-7b-beta-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face24Jklee0717 /llama3-dpo-data-beta0p12-LLaMA3_iter2-seed45text10K<n<100K0 likes10 downloads10mo agoHugging Face25Jklee0717 /llama3-dpo-data-beta0p2-LLaMA3_iter1-seed45text10K<n<100K0 likes10 downloads10mo agoHugging Face26walkmane /BetaAItextn<1K0 likes10 downloads2mo agoHugging Face27Locutusque /VENUS_DATA_betatext10K<n<100K0 likes9 downloads3y agoHugging Face28ogiwemy /beta2textn<1K0 likes8 downloads2y agoHugging Face29open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face30open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.