data-selection
Hachipo_-_qwen2.5-0.5B_educational_instruct_selec10000_pythonblock_dataselection_enja-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec10000_pythonblock_dataselection_en-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec5000_pythonblock_dataselection_jaen-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec5000_pythonblock_dataselection_enja-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec5000_pythonblock_dataselection_ja-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec10000_pythonblock_dataselection_jaen-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec5000_pythonblock_dataselection_en-ggufHachipo_-_qwen2.5-0.5B_educational_instruct_selec10000_pythonblock_dataselection_ja-gguf
lg-longtail-data-selection-experiment-archive-20260829
LG Long-tail data-selection experiment archive
This public repository is the canonical, self-contained archive for the ConvFinQA
data-selection budget-scaling experiments started on 2026-08-29 and the preceding
diversity reproduction run started on 2026-08-26. It replaces the earlier split
model/subset repositories.
The archive preserves the local directory trees in full: selected subsets,
selection manifests, LoRA adapters, DeepSpeed optimizer states, trainer states,
logs, raw… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/lg-longtail-data-selection-experiment-archive-20260829.OpenR1_Math_220k_40k_samples_512_complen_128_block_128_diff_1_rollouts_copyOpenR1_Math_220k_40k_samples_512_complen_128_block_128_diff_1_rolloutslg-longtail-data-selection-artifacts
Frozen data-selection artifacts
Small, portable source-row index files live in indices/ and are committed to
Git. Large generated inputs live in the Hugging Face dataset repository
xxccho/lg-longtail-data-selection-artifacts,
recorded in manifest.json; every file is pinned by both Hub revision and
SHA256. The repository is public so collaborators can reproduce the selection
without account-specific filesystem access.
The bundle intentionally excludes diversity artifacts. That… See the full description on the dataset page: https://huggingface.co/datasets/xxccho/lg-longtail-data-selection-artifacts.KGQA_model_selection_dataLCB-Selection-Data-8192
LCB Selection Data 8192 (VERL format)
VERL-compatible parquet flatten of LCB selection data, used as the data source for current
GRPO training runs (4B / 8B / bs96-nrp variants) on both the legacy bilevel-reweighting
training stack and the VERL training stack.
Schema
Column
Type
Notes
data_source
str
constant: code_selection
prompt
list[dict]
judge prompt messages (system + user)
ability
str
constant: code_verification
reward_model
dict
{ground_truth:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/LCB-Selection-Data-8192.
LREC-2026-Data-Selection-EffectsAlgorithm-Selection-System-Based-On-Input-Datarepro-targate-target-aware-data-selection-via-token-attenuation-gatesrepro-is-data-shapley-not-better-than-random-in-data-selection-ask-nashrepro-single-rollout-hidden-state-dynamics-for-training-free-rlvr-data-selectionrepro-active-learning-with-low-rank-structure-for-data-selection
