datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
auxiliary-views-knowledge-acquisition
Auxiliary Views Knowledge Acquisition
This repository contains the cleaned source documents and evaluation
probes used in Knowledge Acquisition During Pre-training? Large Language Models
Learn Better With Auxiliary Views (arXiv:2609.04180).
News
August 21, 2026: Our paper was accepted to Findings of EMNLP 2026.
Configurations
Configuration
Split
Rows
documents
train
30
factual_cloze
test
6,435
factual_mcqa_5shot
test
4,515… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/auxiliary-views-knowledge-acquisition.DOD-Acquisition-Transformation-Strategy
DoD Acquisition Transformation Strategy Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Acquisition Transformation Strategy: Rebuilding the Arsenal of Freedom.
The strategy presents a department-wide plan to transform the defense acquisition system into a Warfighting Acquisition System focused on speed, flexibility, risk-based… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Acquisition-Transformation-Strategy.LDM-CoT-Acq-SFT-16K
LDM-CoT-Acq-SFT-16K
Supervised fine-tuning (SFT) corpus for Large Discovery Models (LDM): a dataset that
distils an acquisition-guided, test-time search policy into a language-model proposer so
that a single forward pass emulates a full model-based optimization loop.
Dataset Summary
An LDM couples three components in a recurrent generate → select → evaluate → update
loop: an LLM that proposes candidate experiments, a probabilistic surrogate that maps
observations… See the full description on the dataset page: https://huggingface.co/datasets/Yangtze-ailab/LDM-CoT-Acq-SFT-16K.ptdbench-reward-design-reward-land-acquisition-008-dataset
PTDBench dataset snapshot: reward_land_acquisition_008
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-land-acquisition-008-dataset.
