dataset-selection
Feature_Selection_Dataset
Feature Selection Benchmark Datasets (HRLFS)
This repository hosts the 21 benchmark datasets used in our ACM TKDD paper:
Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning
Weiliang Zhang, Xiaohan Huang, Yi Du, Ziyue Qiao, Qingqing Long, Zhen Meng, Yuanchun Zhou, Meng Xiao
ACM Transactions on Knowledge Discovery from Data (TKDD), 2026
📄 Paper code: https://github.com/coco11563/HARLFS
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Shaow/Feature_Selection_Dataset.dataset_selection_v2llm-selection-dataset
llm-selection-dataset
Selected training samples in Parquet format.
Each file contains dataset, id, and messages. Each message contains role and content.
The align_<dimension>_<threshold>.parquet files contain selections for the corresponding alignment dimension and minimum score threshold.
dataset-preferences-llm-course-model-selection
Dataset Card for dataset-preferences-llm-course-model-selection
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-model-selection/raw/main/pipeline.yaml"
or explore the configuration:
distilabel… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-model-selection.nhanes-dataset-prediction-selection-stage-1catalog-selection-dataset
