datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EagleX-WorldContinued
Dataset Card for EagleX v2 Dataset
This dataset was used to train RWKV Eagle 7B for continued pretrain of 1.1T tokens (approximately) (boosting it to 2.25T) with the final model being released as RWKV EagleX v2.
Dataset Details
Dataset Description
EagleX-WorldContinued is a pretraining dataset built from many of our datasets over at Recursal AI + a few others.
Curated by: M8than, KaraKaraWitch, Darok
Funded by [optional]: Recursal.ai
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/EagleX-WorldContinued.RWKV-World-Listing
RWKV World Corpus
(includes v3, v2.1 and v2 subsets)
This is an itemised and annotated list of the RWKV World corpus as described in the RWKV-7 paper
which is a multilingual dataset with about 3.1T tokens used to train the
"Goose" RWKV-7 World model series.
RWKV World v3 was crafted from public datasets spanning >100 world languages
(80% English, 10% multilang, and 10% code).
PREVIEW
Random subsampled subsets of the world v3 corpus are available in the… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/RWKV-World-Listing.
