datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sea-commoncrawlsea-syntheticsailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B.
sea-commoncrawl-high-qualitysea-pdf-textsea-internetsailor2-pretrain-data-stage2The pre-training dataset (stage2) for the Sailor2 models, including 1B, 8B and 20B.
community-datasetsea-ultrafeedbacksailor2-sft-stage1Vietnamese_RAG
Dataset Card for Dataset Name
Vi's RAG is an comprehensive Vietnamese dataset optimized for RAG Evaluation, build by ZD AI lab and release under Apache license 2.0.
Dataset Details
There are four datasets in this card :
Vietnamese version of Expert QA that we utilize the strong translation ability of GPT-4 for translation task
RAG ViQuAD which was carefully chosen from UIT-ViQuAD2.0 with additional context column filtered by title
Legal RAG and BKAI_RAG are long form RAG… See the full description on the dataset page: https://huggingface.co/datasets/sailor2/Vietnamese_RAG.Flores-Plus-Evaluation-Log-Preview-Cleanedsailor2-sft-stage2Sakalti__Sailor-japanese-details
Dataset Card for Evaluation run of Sakalti/Sailor-japanese
Dataset automatically created during the evaluation run of model Sakalti/Sailor-japanese
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Sailor-japanese-details.
