datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RWKU
Dataset Card for Real-World Knowledge Unlearning Benchmark (RWKU)
Dataset Summary
RWKU is a real-world knowledge unlearning benchmark specifically designed for large language models (LLMs).
This benchmark contains 200 real-world unlearning targets and 13,131 multi-level forget probes, including 3,268 fill-in-the-blank probes, 2,879 question-answer probes, and 6,984 adversarial-attack probes.
RWKU is designed based on the following three key factors:
For the task setting… See the full description on the dataset page: https://huggingface.co/datasets/jinzhuoran/RWKU.rwa-attest
RWA attestations
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Historical RWA attest tape (may include names outside XRPL reader-16). Live reader: /api/xrpl. Not a board writer.
Not a certificate. Hub is a printer of live GET, not a second engine. Fetch fail → UNCHECKABLE.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that… See the full description on the dataset page: https://huggingface.co/datasets/csoai/rwa-attest.RWKV-7-ArithmeticRWKV-7-Arithmetic-0.1B 加减法运算模型的训练和测试数据集。
该模型实现基础加减法运算和加减法方程求解功能,能够处理整数部分为 1-12 位、小数部分为 0-6 位的数值,支持中英文数字、全半角格式以及大小写字符的多种表示形式,可实现基础加减法运算和加减法方程求解功能。
训练数据集说明
以下是我们使用的加减法训练数据类型,共包含 30000587 33000147 条单轮加减法 QA 数据,约 1B(1014434168) token。
数据文件名
数据条数
数据说明
示例
ADD_4M
3997733
1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2 个随机空格3. 含简单自然语言描述/自然语言噪声
{"text": "User: 249476576 减 796580834 还剩多少?\n\nAssistant: -547104258"}
ADD_2M
1999673
1. 使用‘全角’、‘中文数字’、‘大写中文数字’随机替换整个数字2. 运算符附近有 1~2… See the full description on the dataset page: https://huggingface.co/datasets/shoumenchougou/RWKV-7-Arithmetic.blinkdl-rwkv-indonesiaEagleX-WorldContinued
Dataset Card for EagleX v2 Dataset
This dataset was used to train RWKV Eagle 7B for continued pretrain of 1.1T tokens (approximately) (boosting it to 2.25T) with the final model being released as RWKV EagleX v2.
Dataset Details
Dataset Description
EagleX-WorldContinued is a pretraining dataset built from many of our datasets over at Recursal AI + a few others.
Curated by: M8than, KaraKaraWitch, Darok
Funded by [optional]: Recursal.ai
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/EagleX-WorldContinued.rwa-testnet-unmeasured
RWA testnet — UNMEASURED, on purpose
Label: TESTNET · Register: UNMEASURED. No invented MEASURED scores; no fake AUM/TVL as grades.
Staging pack for NEXT_300 #186. Mirrors GET /api/rwa-attestation honesty (measured_score: null). Unsigned clean-play rows only until custody unlocks signed cards.
Related indices pack: csoai/labour-economy-unmeasured (#139 / #253) — labour/economy stay UNMEASURED
JMWH remains DEMO ONLY in the catalog fixture
Firewall: never fuse labour/economy into… See the full description on the dataset page: https://huggingface.co/datasets/csoai/rwa-testnet-unmeasured.RWKV-World-v3
RWKV-7 (Goose) World v3 Corpus
Paper | Code
This is an itemised and annotated list of the RWKV World v3 corpus
which is a multilingual dataset with about 3.1T tokens used to train the
"Goose" RWKV-7 World model series.
RWKV World v3 was crafted from public datasets spanning >100 world languages
(80% English, 10% multilang, and 10% code). Also available as a HF Collection of Datasets.
Subsampled subsets (previews) of the corpus are available as 100k JSONL dataset and 1M JSONL dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goose-World/RWKV-World-v3.RWKV-World-Listing
RWKV World Corpus
(includes v3, v2.1 and v2 subsets)
This is an itemised and annotated list of the RWKV World corpus as described in the RWKV-7 paper
which is a multilingual dataset with about 3.1T tokens used to train the
"Goose" RWKV-7 World model series.
RWKV World v3 was crafted from public datasets spanning >100 world languages
(80% English, 10% multilang, and 10% code).
PREVIEW
Random subsampled subsets of the world v3 corpus are available in the… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/RWKV-World-Listing.RWKV-notebook-assets
RWKV notebook assets
Various asset files, used in RWKV notebook tutorial examples and demos
tarot-rws-historical-meanings
StarTarot English RWS Historical Meanings and Golden Dawn Correspondences
Version 1.1.0 · English · 78 cards · DOI: https://doi.org/10.5281/zenodo.22917681 ·
all versions: https://doi.org/10.5281/zenodo.21381779
Project website ·
Dataset documentation ·
Tarot card guide ·
Tarot spreads
Summary
This dataset gives one structured record for each of the 78 cards of the
Rider–Waite–Smith (RWS) tarot deck. Each record combines:
A. E. Waite's divinatory meanings from… See the full description on the dataset page: https://huggingface.co/datasets/StarTarotOnline/tarot-rws-historical-meanings.rwkv-world-3-subsample-previewmod-rwkv-instruction-oldmod-rwkv-instruct-oigmoderationpizza_rwrrealworld-qa
Real World Records — Evaluation Dataset (QA pairs) + Coverage Summary
TL;DR
A publication-ready evaluation dataset for a Real World Studios ML system, delivered as JSONL (one JSON object per line with input, target, source). It covers Real World Records label/studio/WOMAD history, Peter Gabriel's full discography, and a broad cross-section of the label's artist catalogue with verified catalogue numbers, dates, producers, and collaborations.
Every answer is grounded in a named… See the full description on the dataset page: https://huggingface.co/datasets/rw-robai/realworld-qa.mod-rwkv-paper-civilcommentskinyarwanda-instruction-datasetTranslation-RwkvFormatmod-rwkv-paper-textrwkv-world-v3-subsample-100kRWKV__rwkv-raven-14b-details
Dataset Card for Evaluation run of RWKV/rwkv-raven-14b
Dataset automatically created during the evaluation run of model RWKV/rwkv-raven-14b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RWKV__rwkv-raven-14b-details.pizza_rwr_k10_iter1goodsmash-rwa-boardpizza_rwr_iter1RWKVmemorizationtestrwkvtest2instja5rwkv-world-v3-subsampledeepseek-v3-distill-RWA-1000
A Practical Exploration of Mixed-Style Response LLMs via Few-Shot LoRA Fine-Tuning
1. Research Background and Motivation
With the increasingly widespread application of Large Language Models (LLMs) today, how to make model outputs more transparent, natural, and understandable has become an important research direction. Traditional LLMs typically output the final answer directly, making their internal reasoning process a "black box" to the user. To enhance the… See the full description on the dataset page: https://huggingface.co/datasets/dianzinao/deepseek-v3-distill-RWA-1000.langchain-docsrwitz__go-bruins-v2-details
Dataset Card for Evaluation run of rwitz/go-bruins-v2
Dataset automatically created during the evaluation run of model rwitz/go-bruins-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/rwitz__go-bruins-v2-details.
