datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Demeter-LongCoT-6M
Demeter-LongCoT-6M
Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.Helios-R-6M
Helios-R-6M
Helios-R-6M is a high-quality, compact reasoning dataset designed to strengthen multi-step problem solving across mathematics, computer science, and scientific inquiry. While the dataset covers a range of disciplines, math constitutes the largest share of examples and drives the reasoning complexity.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset = load_dataset("prithivMLmods/Helios-R-6M"… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Helios-R-6M.fineweb-edu-dedup6m
Stage 1 (S1): General Knowledge Anchor — 6M FineWeb-Edu-Dedup
1. Project Overview
This dataset represents the General Knowledge Acquisition Phase (S1) for a research project focused on developing a Domain-Adaptive LLM for ISO 27001 Information Security Auditing.
S1 serves as the cognitive foundation. This corpus is designed to establish high-level linguistic proficiency and general reasoning before the introduction of specialized regulatory standards in Stage 2.… See the full description on the dataset page: https://huggingface.co/datasets/JoTeqtheFirstAI/fineweb-edu-dedup6m.SyntheticTexts6MThis synthetic dataset is generated with russian context-free grammar. It contains about ~88M tokens.
ruler-6m-niah-external
RULER-6M NIAH Eval — external handoff
50 × single_needle_uuid samples at ~6 M tokens per sample. One of four
length variants (1 M / 2 M / 6 M / 12 M) prepared for the external
long-context retrieval handoff.
field
value
samples
50
tasks
{single_needle_uuid: 50}
target tokens
6,000,000
seed
1344
negative_rate
0.0
eval/heldout/data.jsonl is the chat-templated form, ready for
model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-6m-niah-external.
