datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tsiolkovsky-papers
Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus
Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the
Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky
(1857–1935), who derived the rocket equation and described the multistage rocket
decades before anyone could test either.
The archive had been scanned and put online, but without a catalogue you could
query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.Alpha-Instruct
Alpha-Instruct
A synthetic instruction-tuning dataset for quantitative finance, covering formulaic alphas, technical indicators, and academic factor definitions. Designed to fine-tune language models on the vocabulary and reasoning patterns of quant researchers.
Dataset Summary
336 rows of instruction–response pairs in chat format, generated from three distinct quant finance source corpora and post-processed to remove noise and near-duplicates.
Each example is a messages… See the full description on the dataset page: https://huggingface.co/datasets/VladHong/Alpha-Instruct.Lewis_Instruct
Dataset Card for Lewis Carroll Conversational Dataset
Dataset Description
This dataset is a highly curated collection of conversational back-and-forths extracted from the classic, public-domain prose works of Lewis Carroll. It is designed for fine-tuning Large Language Models (LLMs) to adopt a whimsical, highly logical, and slightly absurd conversational tone, mirroring the unique banter found in the Alice in Wonderland universe.
Unlike standard dialogue datasets, the… See the full description on the dataset page: https://huggingface.co/datasets/VladHong/Lewis_Instruct.
