datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-dataalpaca_farmData used in the original AlpacaFarm experiments.
Includes SFT and preference examples.Hachi-Alpaca
Hachi-Alpaca
Hachi-Alpacaは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。
Dataset Details
Dataset Description
Curated by: HachiML
Language(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp
Uses
# library
fromdatasets import load_dataset
# Recommend getting the latest version… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/Hachi-Alpaca.alpaca-french-mixtral
License & Attribution
MTEB-format derivative of AIffl/Alpaca_french_mixtral (French Alpaca, Mixtral-translated). Query = instruction; corpus = answer. Deterministically subsampled to ~10k. Licensed under Apache-2.0 (same as source).
details_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200.alpaca_jp_python
alpaca_jp_python
alpaca_jp_pythonは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットはmistralai/Mixtral-8x22B-Instruct-v0.1によって精査されています。
Dataset Details
Dataset Description
Curated by: HachiML
Language(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp
Uses
# library
fromdatasets import load_dataset
# Recommend getting the latest… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_python.alpaca_farmData used in the original AlpacaFarm experiments.
Includes SFT and preference examples.alpaca_jp_math
alpaca_jp_math
alpaca_jp_mathは、
Stanford Alpacaの手法
mistralai/Mixtral-8x22B-Instruct-v0.1
で作った合成データ(Synthetic data)です。モデルの利用にはDeepinfraを利用しています。
また、"_cleaned"がついたデータセットは以下の手法で精査されています。
pythonの計算結果がきちんと、テキストの計算結果が同等であるか確認
LLM(mistralai/Mixtral-8x22B-Instruct-v0.1)による確認(詳細は下記)
code_result, text_resultは小数第三位で四捨五入してあります。
Dataset Details
Dataset Description
Curated by: HachiMLLanguage(s) (NLP): Japanese
License: Apache 2.0
Github: Alpaca-jp… See the full description on the dataset page: https://huggingface.co/datasets/HachiML/alpaca_jp_math.alpaca-gpt4-trpython-code-instructions-18k-alpaca-standardized
Dataset Card for "python-code-instructions-18k-alpaca-standardized"
More Information needed
details_EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-KTO-beta
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-KTO-beta
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-KTO-beta.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-KTO-beta.alpaca-cleaned-gemini-hun-ratingsEz az adathalmaz úgy keletkezett, hogy a Bazsalanszky/alpaca-cleaned-gemini-hun-n lefuttattam egy llm által támogatott értékelést.
Az értékelő modell a gemini-pro (az ingyenes) volt. A használt kód az alpagasus módosítása: https://github.com/boapps/alpagasus-hu
godlikehhd__alpaca_data_score_max_0.1_2600-details
Dataset Card for Evaluation run of godlikehhd/alpaca_data_score_max_0.1_2600
Dataset automatically created during the evaluation run of model godlikehhd/alpaca_data_score_max_0.1_2600
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/godlikehhd__alpaca_data_score_max_0.1_2600-details.relabeled_alpacafarm_pythiasft_20K_preference_data_minlength
Dataset Card for "relabeled_alpacafarm_pythiasft_20K_preference_data_minlength"
More Information needed
alpaca_farm_gpt4SWE_TaskDecompositionfinance-alpaca-1k-testEpistemeAI__Alpaca-Llama3.1-8B-details
Dataset Card for Evaluation run of EpistemeAI/Alpaca-Llama3.1-8B
Dataset automatically created during the evaluation run of model EpistemeAI/Alpaca-Llama3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Alpaca-Llama3.1-8B-details.EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta-details.JP-AlpaCare-MedInstruct-52k
JP-AlpaCare-MedInstruct-52k
This dataset is a Japanese-translated and aligned version of AlpaCare-MedInstruct-52k.
The translation was performed automatically using gpt-4o-2024-05-13, preserving alignment between English and Japanese instructions, inputs, and outputs. Total data size is 51992.
Dataset Details
Original Dataset: AlpaCare-MedInstruct-52k
Translation Model: GPT-4o (gpt-4o-2024-05-13)
Fields:
id (ID)
instruction_ja, input_ja, output_ja (Japanese)
id_en… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/JP-AlpaCare-MedInstruct-52k.EpistemeAI2__Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.06-8B-Philos-dpo-details.EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama-3.1-8B-Philos-DPO-200-details.spoken-alpaca-gpt04alpaca_skewexp_minlength_merged
Dataset Card for "alpaca_skewexp_minlength_merged"
More Information needed
EpistemeAI2__Fireball-Alpaca-Llama3.1-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1-8B-Philos-details.EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R1-details.EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details
Dataset Card for Evaluation run of EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-Alpaca-Llama3.1.08-8B-Philos-C-R2-details.relabeled_alpacafarm_pythiasft_20K_preference_data
Dataset Card for "relabeled_alpacafarm_pythiasft_20K_preference_data"
More Information needed
relabeled_alpacafarm_pythiasft_20K_preference_data_modelength
Dataset Card for "relabeled_alpacafarm_pythiasft_20K_preference_data_modelength"
More Information needed
EpistemeAI2__Fireball-Alpaca-Llama3.1.03-8B-Philos-details
Dataset Card for Evaluation run of EpistemeAI2/Fireball-Alpaca-Llama3.1.03-8B-Philos
Dataset automatically created during the evaluation run of model EpistemeAI2/Fireball-Alpaca-Llama3.1.03-8B-Philos
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI2__Fireball-Alpaca-Llama3.1.03-8B-Philos-details.
