datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_sambanovasystems__SambaLingo-Arabic-Chat-70B
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Chat-70B
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Chat-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Chat-70B.x-self-instruct-seed-32
Dataset Card for xOA22 - Multilingual Prompts from OpenAssistant
Dataset Summary
x-self-instruct-seed-32 consists of 32 prompts chosen out of the 252 prompts in the self-instruct-seed dataset from the Self-Instruct paper. These 32 prompts were filtered out according to the following criteria:
Should be natural in a chat setting
Therefore, we filter out any prompts with "few-shot examples", as these are all instruction prompts that we consider unnatural in a chat setting… See the full description on the dataset page: https://huggingface.co/datasets/sambanovasystems/x-self-instruct-seed-32.details_sambanovasystems__SambaLingo-Arabic-Base
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Base
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Base.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Base.xOA22
Dataset Card for xOA22 - Multilingual Prompts from OpenAssistant
Dataset Summary
xOA22 consists of 22 prompts originally shown in Appendix E, page 25 of the OpenAssistant Conversations paper. These 22 prompts were then manually translated by volunteers into 5 languages: Arabic, Simplified Chinese, French, Hindi and Spanish.
These prompts were originally created for human evaluations of the multilingual abilities of BLOOMChat. Since not all prompts could be directly… See the full description on the dataset page: https://huggingface.co/datasets/sambanovasystems/xOA22.details_sambanovasystems__SambaLingo-Arabic-Chat
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Chat
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Chat.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Chat.details_sambanovasystems__SambaLingo-Arabic-Base-70B
Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Base-70B
Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Base-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Base-70B.distill_70b_infra_sambanova
