datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-Research-GooseReason-0.7M
GooseReason-0.7M
Synthesized with Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
GooseReason-0.7M is a large-scale RLVR dataset with over 0.7 million tasks across mathematics, programming, and general scientific domains, synthesized by the Golden Goose pipeline. It is used to train GooseReason-4B-Instruct, which achieves new state-of-the-art results among 4B-Instruct models across 15 diverse benchmarks, spanning mathematics… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Research-GooseReason-0.7M.RWKV-World-v3
RWKV-7 (Goose) World v3 Corpus
Paper | Code
This is an itemised and annotated list of the RWKV World v3 corpus
which is a multilingual dataset with about 3.1T tokens used to train the
"Goose" RWKV-7 World model series.
RWKV World v3 was crafted from public datasets spanning >100 world languages
(80% English, 10% multilang, and 10% code). Also available as a HF Collection of Datasets.
Subsampled subsets (previews) of the corpus are available as 100k JSONL dataset and 1M JSONL dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goose-World/RWKV-World-v3.adaption-goose-governance-broad-seed-v1-augmented
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-goose_governance_broad_seed_v1 (augmented)
An English instruction-tuning dataset covering core aspects of governance, including political systems, public policy, international relations, and security. Prompts span multiple task types such as conceptual inquiries, comparative institutional analyses, policy trade-off assessments, and evidence synthesis. Completions provide neutral… See the full description on the dataset page: https://huggingface.co/datasets/darthludious/adaption-goose-governance-broad-seed-v1-augmented.rwkv-world-v3-subsample-100krwkv-world-v3-subsampleGive-Yourself-Goosebumps-ShareGPThttps://goosebumps.fandom.com/wiki/Give_Yourself_Goosebumps
