bel
Datasets
All datasets matching “bel”belebele
The Belebele Benchmark for Massively Multilingual NLU Evaluation
Belebele is a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. This dataset enables the evaluation of mono- and multi-lingual models in high-, medium-, and low-resource languages. Each question has four multiple-choice answers and is linked to a short passage from the FLORES-200 dataset. The human annotation procedure was carefully curated to create questions that discriminate… See the full description on the dataset page: https://huggingface.co/datasets/facebook/belebele.Belle_1.4M-SLAM-Omni
Belle_1.4M
This dataset is prepared for the reproduction of SLAM-Omni.
This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s)
🔧 Modifications
Data Filtering: We removed samples with excessively long data.
Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/Belle_1.4M-SLAM-Omni.bellhart_trainingtrain_1M_CN
内容
包含约100万条由BELLE项目生成的中文指令数据。
样例
{
"instruction": "给定一个文字输入,将其中的所有数字加1。\n“明天的会议在9点开始,记得准时到达。”\n",
"input": "",
"output": "“明天的会议在10点开始,记得准时到达。”"
}
字段:
instruction: 指令
input: 输入(本数据集均为空)
output: 输出
使用限制
仅允许将此数据集及使用此数据集生成的衍生物用于研究目的,不得用于商业,以及其他会对社会带来危害的用途。
本数据集不代表任何一方的立场、利益或想法,无关任何团体的任何类型的主张。因使用本数据集带来的任何损害、纠纷,本项目不承担任何责任。
scannetv2
ScanNet Instructions
To acquire the access to ScanNet dataset, Please refer to the ScanNet project page and follow the instructions there. You will get a download-scannet.py script after your request for the ScanNet dataset is approved. Note that only a subset of ScanNet is needed. Once you get download-scannet.py, please use the commands below to download the portion of ScanNet that is necessary for ScanRefer:
python2 download-scannet.py -o data/scannet --type _vh_clean_2.ply… See the full description on the dataset page: https://huggingface.co/datasets/Believe0029/scannetv2.train_0.5M_CN
内容
包含约50万条由BELLE项目生成的中文指令数据。
样例
{
"instruction": "给定一个文字输入,将其中的所有数字加1。\n“明天的会议在9点开始,记得准时到达。”\n",
"input": "",
"output": "“明天的会议在10点开始,记得准时到达。”"
}
字段:
instruction: 指令
input: 输入(本数据集均为空)
output: 输出
使用限制
仅允许将此数据集及使用此数据集生成的衍生物用于研究目的,不得用于商业,以及其他会对社会带来危害的用途。
本数据集不代表任何一方的立场、利益或想法,无关任何团体的任何类型的主张。因使用本数据集带来的任何损害、纠纷,本项目不承担任何责任。
