datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMScan-betaAloe-Beta-Medical-Collection
Aloe-Beta-Medical-Collection
Collection of curated datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available medical instruction tuning data sources (QA format). Most data samples correspond to single-turn QA pairs, while a small proportion contain multi-turn. All data sources are publicly available for… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-Medical-Collection.Aloe-Beta-General-Collection
Aloe-Beta-Medical-Collection
Collection of curated general datasets used to fine-tune Aloe-Beta.
Dataset Details
Dataset Description
We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including:
Coding, math, data analysis, STEM, etc.
Function calling
Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.HuggingFaceH4__zephyr-7b-beta-details
Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.ReverseBass-Beta-StatusAloe-Beta-DPO
Aloe-Beta-Medical-Collection
Collection of curated DPO datasets used to align Aloe-Beta.
Dataset Details
Dataset Description
The first stage of the Aloe-Beta alignment process. We curated data from many publicly available data sources, including three different types of data:
Medical preference data: TsinghuaC3I/UltraMedical-Preference
General preference data:… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-DPO.lm-eval-results-HuggingFaceH4-mistral-7b-sft-beta-private
Dataset Card for Evaluation run of HuggingFaceH4/mistral-7b-sft-beta
Dataset automatically created during the evaluation run of model HuggingFaceH4/mistral-7b-sft-beta
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-HuggingFaceH4-mistral-7b-sft-beta-private.github_fetch_huggingface_terminal_9091_n3v8x2_source_beta
Beta Support Conversations
Anonymized customer support conversation transcripts.
Dataset ID: SRC-BETA
Catalog: ghfht9091n3v8x2
Origin: Community tech-support forum public dump (2022-2024)
Records: 8,152
License: Apache-2.0
EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details.cpayne1303__llama-43m-beta-details
Dataset Card for Evaluation run of cpayne1303/llama-43m-beta
Dataset automatically created during the evaluation run of model cpayne1303/llama-43m-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cpayne1303__llama-43m-beta-details.CausalLM__34b-beta-details
Dataset Card for Evaluation run of CausalLM/34b-beta
Dataset automatically created during the evaluation run of model CausalLM/34b-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__34b-beta-details.HPAI-BSC__Qwen2.5-Aloe-Beta-7B-details
Dataset Card for Evaluation run of HPAI-BSC/Qwen2.5-Aloe-Beta-7B
Dataset automatically created during the evaluation run of model HPAI-BSC/Qwen2.5-Aloe-Beta-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HPAI-BSC__Qwen2.5-Aloe-Beta-7B-details.Test
Scaling Synthetic Data Creation with 1,000,000,000 Personas
This repo releases data introduced in our paper Scaling Synthetic Data Creation with 1,000,000,000 Personas:
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce PERSONA HUB – a collection of 1 billion diverse personas automatically curated from web… See the full description on the dataset page: https://huggingface.co/datasets/beta42ZH/Test.AItuber-Realworld-Data-beta1
AItuber Realworld Chat Dataset
概要
本データセットは、AItuber(AI VTuber)の 配信チャット会話データ を合成的に生成したものです。AItuberペルソナデータと nvidia/Nemotron-Personas-Japan のユーザーペルソナを入力に、複数の視聴者がコメントし、AItuberが応答する4ターンのリアルな配信チャットデータを生成しています。生成にはSDG-Nexusという合成データ生成パイプラインを用いました。(sdg-nexus)
データの説明
項目
内容
件数
1,946件
形式
JSONL(1行1JSON)
言語
日本語
生成日
2026年3月
ライセンス
odc-by ( Open Data Commons Attribution License )… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/AItuber-Realworld-Data-beta1.utkmst__chimera-beta-test2-lora-merged-details
Dataset Card for Evaluation run of utkmst/chimera-beta-test2-lora-merged
Dataset automatically created during the evaluation run of model utkmst/chimera-beta-test2-lora-merged
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/utkmst__chimera-beta-test2-lora-merged-details.betaHPAI-BSC__Llama3.1-Aloe-Beta-8B-details
Dataset Card for Evaluation run of HPAI-BSC/Llama3.1-Aloe-Beta-8B
Dataset automatically created during the evaluation run of model HPAI-BSC/Llama3.1-Aloe-Beta-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HPAI-BSC__Llama3.1-Aloe-Beta-8B-details.Replete-AI__Replete-LLM-Qwen2-7b_Beta-Preview-details
Dataset Card for Evaluation run of Replete-AI/Replete-LLM-Qwen2-7b_Beta-Preview
Dataset automatically created during the evaluation run of model Replete-AI/Replete-LLM-Qwen2-7b_Beta-Preview
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Replete-AI__Replete-LLM-Qwen2-7b_Beta-Preview-details.llama3-dpo-data-beta0p2-LLaMA3_iter1-seed43audit-reports-simdzakwan__dzakwan-MoE-4x7b-Beta-details
Dataset Card for Evaluation run of dzakwan/dzakwan-MoE-4x7b-Beta
Dataset automatically created during the evaluation run of model dzakwan/dzakwan-MoE-4x7b-Beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/dzakwan__dzakwan-MoE-4x7b-Beta-details.Jimmy19991222__llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-rougeL-beta10-gamma0.3-lr1.0e-6-scale-log-details.Sakalti__light-7b-beta-details
Dataset Card for Evaluation run of Sakalti/light-7b-beta
Dataset automatically created during the evaluation run of model Sakalti/light-7b-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__light-7b-beta-details.llama3-dpo-data-beta0p12-LLaMA3_iter2-seed45llama3-dpo-data-beta0p2-LLaMA3_iter1-seed45BetaAIVENUS_DATA_betabeta2Jimmy19991222__llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-rouge2-beta10-1minus-gamma0.3-rerun-details.Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.
