datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lg-longtail-data-selection-experiment-archive-20260829
LG Long-tail data-selection experiment archive
This public repository is the canonical, self-contained archive for the ConvFinQA
data-selection budget-scaling experiments started on 2026-08-29 and the preceding
diversity reproduction run started on 2026-08-26. It replaces the earlier split
model/subset repositories.
The archive preserves the local directory trees in full: selected subsets,
selection manifests, LoRA adapters, DeepSpeed optimizer states, trainer states,
logs, raw… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/lg-longtail-data-selection-experiment-archive-20260829.OpenR1_Math_220k_40k_samples_512_complen_128_block_128_diff_1_rollouts_copyOpenR1_Math_220k_40k_samples_512_complen_128_block_128_diff_1_rolloutslg-longtail-data-selection-artifacts
Frozen data-selection artifacts
Small, portable source-row index files live in indices/ and are committed to
Git. Large generated inputs live in the Hugging Face dataset repository
xxccho/lg-longtail-data-selection-artifacts,
recorded in manifest.json; every file is pinned by both Hub revision and
SHA256. The repository is public so collaborators can reproduce the selection
without account-specific filesystem access.
The bundle intentionally excludes diversity artifacts. That… See the full description on the dataset page: https://huggingface.co/datasets/xxccho/lg-longtail-data-selection-artifacts.KGQA_model_selection_dataLCB-Selection-Data-8192
LCB Selection Data 8192 (VERL format)
VERL-compatible parquet flatten of LCB selection data, used as the data source for current
GRPO training runs (4B / 8B / bs96-nrp variants) on both the legacy bilevel-reweighting
training stack and the VERL training stack.
Schema
Column
Type
Notes
data_source
str
constant: code_selection
prompt
list[dict]
judge prompt messages (system + user)
ability
str
constant: code_verification
reward_model
dict
{ground_truth:… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/LCB-Selection-Data-8192.LLM-Data-Selectionrm-data-selection-random-dataset-12429ipda-sentence-selection-data
IPDA Sentence Selection Training Dataset
Training data for sentence-level claim selection in competitive debate. This dataset teaches models to select the most impactful claims to address during rebuttal speeches.
Dataset Structure
Files
File
Size
Description
sentence_selection_dataset.json
23MB
Full sentence selection dataset
sentence_dpo_format_consistent.json
4.4MB
DPO preference pairs (consistent format)
sentence_sft_train_v2.json
13MB
SFT… See the full description on the dataset page: https://huggingface.co/datasets/debaterhub/ipda-sentence-selection-data.gsm8k_std_variance_top20rm-data-selection-embeddingsrm-data-selection-random-dataset-6885rm-data-selection-random-dataset-5991ml-selection-benchmark-datarm-data-selection-random-dataset-14096BCB-Selection-Data-8192
BCB Selection Data 8192 (VERL format)
VERL-compatible parquet flatten of BigCodeBench selection data, used as the data source
for current GRPO training runs (BCB-specific yamls and cross-domain LCB+BCB yamls) on
both the legacy bilevel-reweighting training stack and the VERL training stack.
Schema
Column
Type
Notes
data_source
str
constant: code_selection
prompt
list[dict]
judge prompt messages (system + user)
ability
str
constant: code_verification… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/BCB-Selection-Data-8192.mats-selection-data-reviewrm-data-selection-random-dataset-23453rm-data-selection-promptsrm-data-selection-random-dataset-23205grpo-q-alignment-preference-data-bon-correct-selectionconseiller-numerique-liste-des-structures-validees-par-le-comite-de-selection-conum-et-conum-coo
Conseiller numérique - Liste des structures validées par le comité de sélection Conum et Conum Coordinateur
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Conseiller numérique - Liste des structures validées par le comité de sélection Conum et Conum Coordinateur qui est disponible à l'adresse https://www.data.gouv.fr/datasets/60bdf0ac3fb013906f98dfc6
Description
Au 5 avril 2023, les candidatures de 2889… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/conseiller-numerique-liste-des-structures-validees-par-le-comite-de-selection-conum-et-conum-coo.gsm8k_pass_rate_diff_top20rm-data-selection-random-dataset-25633rm-data-selection-random-dataset-25382gsm8k_random_control_top20gsm8k_everpass_variance_top20premiers-projets-hydrogene-francais-selectionnes-dans-le-cadre-du-projet-important-dinteret-euro
Premiers projets hydrogène français sélectionnés dans le cadre du Projet important d’intérêt européen commun (PIIEC)
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Premiers projets hydrogène français sélectionnés dans le cadre du Projet important d’intérêt européen commun (PIIEC) qui est disponible à l'adresse https://www.data.gouv.fr/datasets/6234040725644dbe3e927462
Description
Cette carte présente les 15… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/premiers-projets-hydrogene-francais-selectionnes-dans-le-cadre-du-projet-important-dinteret-euro.OpenR1_Math_220k_40k_samples_512_complen_128_block_128_diff_1_rollouts_301225_reward_1rm-data-selection-random-dataset-27772
