datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-1.7b-polaris-fp8-rollouts-20260912
Qwen3-1.7B Polaris FP8 rollouts
Snapshot of three FP8 rollout datasets taken on 2026-09-12. The model is Qwen3-1.7B-Base. Original Parquet files and attempt, file, and checkpoint-lineage metadata are preserved without rewriting.
Configuration
Training batches
Training responses
Validation responses
Total bytes
maxrl_strict
101
827392
144320
6654288117
maxrl_permissive
107
876544
144320
8168064932
dppo
126
1032192
173184
6733682534
Provenance and… See the full description on the dataset page: https://huggingface.co/datasets/steviel/qwen3-1.7b-polaris-fp8-rollouts-20260912.allenai_WildChat-1M-Full-neuralmagic_DeepSeek-Coder-V2-Instruct-FP8glm_5_2_fp8_gender_secret_female_rolloutsglm_5_2_fp8_gender_secret_male_rolloutsTopicAnnotations-Llama-3.1-405B-FP8
WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with topic labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/TopicClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page
top_choice_index: Index… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/TopicAnnotations-Llama-3.1-405B-FP8.imagenet_288_dcae_fp8
ImageNet-1k in 5GB
The full ImageNet-1k compressed to less than 5 GB
Compression procedure:
Resize shorter edge to 288 and crop longer edge to a multiple of 32
Analysis transform: DC-AE f32 c32
Quantization: 8 bit float (e4m3)
Entropy coding: TIFF (CMYK) with deflate
Example dataloader for training
import torch
import datasets
from types import SimpleNamespace
from diffusers import AutoencoderDC
from torchvision.transforms.v2 import ToPILImage, PILToTensor… See the full description on the dataset page: https://huggingface.co/datasets/danjacobellis/imagenet_288_dcae_fp8.FormatAnnotations-Llama-3.1-405B-FP8
WebOrganizer/FormatAnnotations-Llama-3.1-405B-FP8
[Paper] [Website] [GitHub]
This dataset contains 100K web pages annotated with format/type labels by the Llama-3.1-405B-FP8 model. The web pages are a sample of the DCLM RefinedWeb reproduction. It is used as second-stage training data for the WebOrganizer/FormatClassifier.
Dataset Structure
Each example contains the following fields:
text: The text content of the web page
url: The URL of the web page… See the full description on the dataset page: https://huggingface.co/datasets/WebOrganizer/FormatAnnotations-Llama-3.1-405B-FP8.glm_5_2_fp8_ab_contextual_optimism_rolloutsmagpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category
Dataset Card for magpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/gabrielmbmb/magpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category/raw/main/pipeline.yaml"
or… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/magpie-llama-3.1-405b-instruct-fp8-with-system-prompt-per-category.wildchat-4.8m_10k_seed1_Qwen_gemma_granite_FP8_metrics_extendedimagenet_288_dcae_fp8_captionsOriginal https://huggingface.co/datasets/danjacobellis/imagenet_288_dcae_fp8
Captions from https://huggingface.co/datasets/visual-layer/imagenet-1k-vl-enriched and https://huggingface.co/datasets/gmongaras/Imagenet21K_Recaption
Ive uploaded just the captions from both of those aswell. You can find them in the files of this repo
swebench_verified_random_100_folders_Qwen3_Coder_480B_A35B_Instruct_FP8_20260429_192404terminus-2__dev_set_71_tasks__together_ai_Qwen_Qwen3-Coder-480B-A35B-Instruct-FP8_20260211glm_5_2_fp8_ab_animal_welfare_rolloutsqwen3.5-122b-fp8-tpu-traces-consolidateddev_set_v2_rl__r2egym_deepswe_fp8_terminus_2_32b_20260406_202950Llama-4-Maverick-17B-128E-Instruct-FP8-OpenR1-Math-220kgaia_127_GLM_4_7_FP8_20260502_072129financeagent_terminal_Qwen3_Coder_480B_A35B_Instruct_FP8_20260505_192209glm_5_2_fp8_ab_hallucinates_citations_rolloutsterminal_bench_2_rl__r2egym_deepswe_fp8_terminus_2_32b_20260406_203014glm_5_2_fp8_ab_self_promotion_rolloutsglm_5_2_fp8_eval_sandbagger_rolloutsgranite_4.0_h_small_FP8_test_detoxificability_annotation_error_analysismtr-qwen35-fp8-12turn
MTR-Qwen3.5-FP8 12-Turn Multi-Turn Retrieval Dataset
Synthesized multi-turn retrieval dataset using Qwen3.5-FP8 (397B MoE) via SGLang.
Overview
Split
Rows
Description
train
107,928
8,994 conversations × 12 turns (each turn = 1 sample)
test
12,000
1,000 conversations × 12 turns
Each row is a single retrieval test case: a conversation truncated at a specific turn, with a ground-truth document the system should retrieve for the last user query.… See the full description on the dataset page: https://huggingface.co/datasets/OkayestProgrammer/mtr-qwen35-fp8-12turn.Llama-4-Maverick-17B-128E-Instruct-FP8-alpaca-cleanedterminal_bench_2_rl__nemotron_bash_fp8_terminus_2_32b_20260407_232112-tracesgaia_127_Qwen3_Coder_480B_A35B_Instruct_FP8_20260430_052937aider_polyglot_Qwen3_Coder_480B_A35B_Instruct_FP8_20260430_052720-tracesgenshin_woman_Formal_Outfit_flux1_kontext_fp8_extracted
