datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-oss-20b-moe-expert-power-traces-320k
GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer)
This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100.
What is recorded
Each trace corresponds to one capture trial where:
A fixed expert id is selected (expert_00 ... expert_31).
A random hidden-state tensor is generated once per trial.
The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.gpt-oss-20b-moe-expert-power-traces-320k-ds16k
GPT-OSS-20B MoE Expert Power Traces (Downsampled to 16k)
Downsampled variant of the 320k expert-trace capture set.
Source
Raw source dataset (same captures):
32 experts (expert_00..expert_31)
10,000 traces per expert
320,000 total traces
raw trace length ~195k samples per trace
Downsampling method
Each raw trace was resampled to exactly 16384 samples using linear interpolation (np.interp) matching the trainer resampling step.
No baseline normalization and no… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k-ds16k.gpt-oss-20b-lmsys-layer2-expert-traces-600k-10msps-tar-shards
GPT-OSS-20B LMSYS Layer-2 Expert Power Traces — 600k tokens at 10 MSPS, tar-sharded
This is the tar-sharded version of the 600k-token layer-2 expert power-trace capture. It contains the same source data as the raw run directory, but groups per-token trace/record files by shard to avoid a 1.2M-file Hugging Face repository.
Capture summary
Model: openai/gpt-oss-20b
Target layer: 2
Captured region: target_layer_experts_only
Prompts: LMSYS Chat, English-filtered… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-lmsys-layer2-expert-traces-600k-10msps-tar-shards.gpt-oss-20b-lmsys-layer2-expert-traces-600k-10msps
GPT-OSS-20B LMSYS Layer-2 Expert Power Traces — 600k tokens at 10 MSPS
This dataset contains real power/current traces captured from an NVIDIA H100 PCIe system while running GPT-OSS-20B decode on LMSYS prompts. The capture target is layer 2 MoE expert execution only; earlier/later transformer work is executed off-trace to avoid wasting acquisition time.
Capture summary
Model: openai/gpt-oss-20b
Target layer: 2
Captured region: target_layer_experts_only
Prompts:… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-lmsys-layer2-expert-traces-600k-10msps.bilibili-masterpieces
Dataset Card for Bilibili Masterpieces
Dataset Summary
The bilibili-masterpieces dataset is a curated collection of representative works from some of the early well-known content creators (up 主) on the Bilibili platform. This dataset captures key metadata from these videos, providing a snapshot of the creative output that has significantly influenced the Bilibili community.
Supported Tasks and Leaderboards
The dataset can be used for various tasks such as video… See the full description on the dataset page: https://huggingface.co/datasets/wencan2024/bilibili-masterpieces.marvel-masterpieces-with-3dmesh
Dataset Card for reconstructions
Wait! Before you go, ❤️ the dataset! Let's get this trending!
This is a FiftyOne dataset with 255 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces-with-3dmesh.marvel-masterpieces
Dataset Card for marvel_masterpieces
Wait! Before you go, ❤️ the dataset! Let's get this trending!
This is a FiftyOne dataset with 255 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/marvel-masterpieces")
#… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces.owes_public
OpenWebText Expert Selections (GPT-OSS 20B)
This dataset contains top-k expert selections for each token from the first
200,000,000 tokens of vietgpt/openwebtext_en, using the router logits from
openai/gpt-oss-20b. Sequences are chunked to a maximum length of 32 tokens
within each document (no cross-document continuity); shorter tail chunks are
included without padding.
Files
openwebtext_200m_idx.npy: uint16 indices, shape (200_000_000, 24, 4)
openwebtext_200m_val.npy: float16… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/owes_public.bertopic-conflictos-chile-v21-masterpiece
🏆 BERTopic v21 - THE MASTERPIECE
Mejoras sobre v20:
SmartDeduplicator (0.7): Más permisivo, conserva "Argentina", "transfronterizo"
Safety Net: Garantiza mínimo 6 keywords por topic
Stopwords SOTA: Sin meses, años, nombres personales
Visualizaciones: Barcharts y mapas HTML automáticos
Métricas:
K Natural: 50
Total docs: 3,268
Timestamp: 2026-01-13 17:07:41.278922
Archivos:
bertopic_topic_explanations_v21.xlsx: Keywords por K… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v21-masterpiece.Animenz-Piano-Masterpieces
Animenz Piano Masterpieces
Please visit official Animenz YouTube Channel for source audio and video files
Project Los Angeles
Tegridy Code 2025
router_dataset
