CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sailor2 /sea-commoncrawltext100M<n<1B1 likes7.5k downloads2y agoHugging Face02sailor2 /sea-synthetictext10M<n<100M0 likes3.2k downloads2y agoHugging Face03sailor2 /sailor2-pretrain-data-stage1The pre-training dataset (stage1) for the Sailor2 models, including 1B, 8B and 20B. text100M<n<1B0 likes2.9k downloads2y agoHugging Face04sailor2 /sea-commoncrawl-high-qualitytext10M<n<100M0 likes2.1k downloads2y agoHugging Face05sailor2 /sea-pdf-texttext10M<n<100M1 likes1.6k downloads2y agoHugging Face06sailor2 /sea-internettext10M<n<100M1 likes1.6k downloads2y agoHugging Face07sailor2 /sailor2-pretrain-data-stage2The pre-training dataset (stage2) for the Sailor2 models, including 1B, 8B and 20B. text10M<n<100M0 likes489 downloads2y agoHugging Face08sailor2 /community-datasettext1M<n<10M1 likes380 downloads2y agoHugging Face09sailor2 /sea-ultrafeedbacktext10K<n<100K0 likes169 downloads2y agoHugging Face10sailor2 /sailor2-sft-stage1tabular1M<n<10M0 likes159 downloads2y agoHugging Face11sailor2 /Vietnamese_RAG Dataset Card for Dataset Name Vi's RAG is an comprehensive Vietnamese dataset optimized for RAG Evaluation, build by ZD AI lab and release under Apache license 2.0. Dataset Details There are four datasets in this card : Vietnamese version of Expert QA that we utilize the strong translation ability of GPT-4 for translation task RAG ViQuAD which was carefully chosen from UIT-ViQuAD2.0 with additional context column filtered by title Legal RAG and BKAI_RAG are long form RAG… See the full description on the dataset page: https://huggingface.co/datasets/sailor2/Vietnamese_RAG.text1K<n<10K10 likes108 downloads2y agoHugging Face12sailor2 /Flores-Plus-Evaluation-Log-Preview-Cleanedtext100K<n<1M0 likes22 downloads2y agoHugging Face13sailor2 /sailor2-sft-stage2tabular100K<n<1M0 likes19 downloads2y agoHugging Face14open-llm-leaderboard /Sakalti__Sailor-japanese-detailsgated Dataset Card for Evaluation run of Sakalti/Sailor-japanese Dataset automatically created during the evaluation run of model Sakalti/Sailor-japanese The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Sailor-japanese-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.