datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gargantua-R1-Compact
Gargantua-R1 Distribution
Gargantua-R1-Compact(experimental purpose)
Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Gargantua-R1-Compact.Gargantua-R1-Compact
Gargantua-R1 Distribution
Gargantua-R1-Compact(experimental purpose)
Gargantua-R1-Compact is a large-scale, high-quality reasoning dataset primarily designed for mathematical reasoning and STEM education. It contains approximately 6.67 million problems and solution traces, with a strong emphasis on mathematics (over 70%), as well as coverage of scientific domains, algorithmic challenges, and creative logic puzzles. The dataset is suitable for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Gargantua-R1-Compact.MedConclusion-Compact
MedConclusion-Compact
MedConclusion is a large-scale dataset of 5.7M PubMed structured abstracts for biomedical conclusion generation. Each instance pairs the non-conclusion sections of an abstract with the original author-written conclusion, providing naturally occurring supervision for evidence-to-conclusion reasoning. MedConclusion also includes journal-level metadata such as biomedical category and SJR, enabling subgroup analysis across biomedical domains.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/MedConclusion-Compact.lemonseed-compact-foundation-cogen
lemonseed-compact-foundation-cogen
LemonSeed — compact foundation teacher-co-gen training (v2).
Contents
intelligent_compact_foundation_train_v2.jsonl (88 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
PersonalFinance-CoTR-V2-Compact
Personal Finance Reasoning V2 - Compact Edition
This document describes a condensed version of the Personal Finance Reasoning dataset, specifically adapted for training smaller language models (e.g., 1B-4B parameters).
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge extraction. While valuable, these benchmarks often overlook the critical… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance-CoTR-V2-Compact.
