datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DepthBench-FineWeb-Edu-100BT-tokenized
DepthBench FineWeb-Edu 100BT Tokenized
This repository contains the tokenized FineWeb-Edu 100BT sample used by
DepthBench pretraining experiments.
Splits
train/: 139 shards, 99,585,913,529 tokens, and 97,045,608 documents.
eval/: 013_00008, containing 234,993,701 tokens and 225,078 documents.
All remaining source shards are assigned to training. Each source document is
terminated by an EOS token before documents are concatenated.
Format
Each shard… See the full description on the dataset page: https://huggingface.co/datasets/YangyiH/DepthBench-FineWeb-Edu-100BT-tokenized.DepthQA
Dataset Card for DepthQA
This dataset card is the official description of the DepthQA dataset used for Hierarchical Deconstruction of LLM Reasoning: A Graph-Based Framework for Analyzing Knowledge Utilization.
Dataset Details
Dataset Description
Language(s) (NLP): EnglishLicense: Creative Commons Attribution 4.0Point of Contact: Sue Hyun Park
Dataset Summary
DepthQA is a novel question-answering dataset designed to evaluate graph-based reasoning… See the full description on the dataset page: https://huggingface.co/datasets/kaist-ai/DepthQA.MATH-Composition-Depth3
Composition-RL
Paper | Code | Collection
Composition-RL is a data-efficient Reinforcement Learning with Verifiable Rewards (RLVR) approach that automatically composes multiple verifiable problems into a single, harder yet still-verifiable prompt. This method helps maintain informative training signals by combatting the growing number of "too-easy" prompts (pass-rate = 1) that occur during RL training.
Dataset Summary
This project introduces several compositional… See the full description on the dataset page: https://huggingface.co/datasets/xx18/MATH-Composition-Depth3.
