selection
osm-polygon-selection
osm-polygon-selection dataset
A curated set of OpenStreetMap polygons from 310
geographic units — sovereign countries plus sub-country regions
like Brazilian states, Chinese provinces, Indian zones, US states,
Canadian provinces, Japanese regions, and Indonesian islands —
classified by size bin (small / medium / large, area in
[0.1, 100] km²) and tagged by continent (Natural Earth admin0 lookup).
Size bins:
small — area in [0.1, 1) km² (10,000 m² to 1 km², roughly
100 m × 100 m… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-selection.MoE_expert_selection_trace
📖 Introduction
This repository serves as a supplement to our paper "Patterns behind Chaos: Forecasting Data Movement for Efficient Large-Scale MoE LLM Inference".
It contains expert selection profiling traces of four top-tier MoE LLMs ranging from 235B to 1T (DeepSeek-R1, Kimi-K2-Thinking, Llama4-Marverick, and Qwen3-235B) across multiple benchmarks. For each query or request, we log the activated expert ID of every model layer of every generated token.
We provide analyses and… See the full description on the dataset page: https://huggingface.co/datasets/core12345/MoE_expert_selection_trace.Feature_Selection_Dataset
Feature Selection Benchmark Datasets (HRLFS)
This repository hosts the 21 benchmark datasets used in our ACM TKDD paper:
Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning
Weiliang Zhang, Xiaohan Huang, Yi Du, Ziyue Qiao, Qingqing Long, Zhen Meng, Yuanchun Zhou, Meng Xiao
ACM Transactions on Knowledge Discovery from Data (TKDD), 2026
📄 Paper code: https://github.com/coco11563/HARLFS
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/Shaow/Feature_Selection_Dataset.2026-08-14-less-selection-difficult-advice
LESS data selection over the difficult-advice SFT pool
field
value
experiment
LESS (arXiv:2402.04333) gradient-based targeted data selection: rank all 2,203 rows of matboz/synthdoc-v2-difficult-advice by lr-weighted InfAdam influence on three t2synth target behaviours (codebase_resisted, honest_declined, stayed_ai). Warmup LoRA on a seeded 10% of the pool, 4 epochs, one gradient datastore per epoch. Top-220 trait enrichment vs a uniform pool: t6 35.9%, t3 33.6%, t9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-14-less-selection-difficult-advice.premise_selectionWhy do so many people come and download this dataset... the set is unready, the ac2 is incorrect, and it's unlicensed...
SCAND_traj_selection
