datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Know-Your-Sourcesfinancial-english-source-corpus
Financial English Source Corpus
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. This
version preserves the final pre-split source rows.
Derived 1280-token split versions are available separately:
financial-english-source-corpus-qwen35-1280
financial-english-source-corpus-gemma4-e2b-1280
Dataset
Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.source-classifications
NuBerea Source Gold Set
Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists.
This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.classifier_source
Dataset Card for Lapa High Quality Pretraining Dataset
Dataset Description
Dataset Summary
This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.financial-english-source-corpus-qwen35-1280
Financial English Source Corpus Qwen35 1280
This dataset is a filtered, fuzzy-deduplicated English source-text corpus for
financial-domain language-model training and translation-data generation. The
uploaded Parquet files are already prepared with the 1280-token source split
used by the downstream training pipeline.
This split version is derived from the pre-split
Financial English Source Corpus
by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source
Tool-Reasoning SFT — RLVR Retrieval Source Trajectories
156,381 multi-turn agentic retrieval trajectories across three document corpora, in a strict reasoning + tool-call format with validated FSM transitions. Each trajectory records a model searching a corpus, opening documents, and citing relevant passages to answer a question.
Author: Aman Priyanshu
Source Environments
Trajectories were collected against three RLVR retrieval environments from the FORMAT: Search -… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source.falsifyrl-source
FalsifyRL Reward-Hacking Falsification
FalsifyRL is a synthetic, executable benchmark for identifying and repairing proxy-reward failures
in embodied multi-agent reinforcement learning.
Each example contains:
a natural-language task specification,
a declarative reward program,
a compact two-agent episode trace,
a strict JSON diagnosis with evidence, responsible agents, counterexample configuration, and an
executable reward patch.
Dataset design
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-source.dclm-crossover-source
DCLM Cross-Over Source
Subset of DCLM-Baseline
selected for synthetic augmentation with format-aware prompt routing.
Selection
Picked every 3th shard (9313 of 27938 shards)
Word count filter: 50-8000
Per-site cap: 10,000
Format detection: skip prompts that duplicate native document format
Stats
Metric
Value
Source docs scanned
54,947,699
Selected
54,017,165
Total words
44,119,449,000
Avg words/doc
816
Length filtered
930,534… See the full description on the dataset page: https://huggingface.co/datasets/essobi/dclm-crossover-source.math-ai-bench-sources-latest
math-ai-bench-sources-latest
This dataset is an updated aggregated multi-trajectory benchmark built from the latest parallelthinking_benchmark files under /scratch/haowu/datasets/datasets/parallelthinking_benchmark_latest.
It follows the same high-level format as haowu89/math-ai-bench-sources, but it is a newer version with:
updated benchmark composition
updated model set
aligned question coverage across all included models
Included Models
Qwen2.5-1.5B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources-latest.math-ai-bench-sources
math-ai-bench-sources
This dataset contains math_ai_parallelthinking_benchmark.jsonl, built for comparing multiple reasoning trajectories across models on the same set of questions.
File
math_ai_parallelthinking_benchmark.jsonl
Data Construction
The benchmark is built from subsets of zechen-nlp/math-ai-bench (including gpqa) and distilled with the following 3 models:
Qwen_Qwen2.5-1.5B-Instruct
Qwen_Qwen3-4B-Nothinking
Qwen_Qwen3-4B-Thinking
For each model… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources.customer-transcript-source
Customer Transcript Source
Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads.
Motivation and intended use
This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token accounting… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-source.
