datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3.1-8B-BaldEagle3-UltrachatLlama3.1-8B-BaldEagle3-ShareGPTTransmem_ecsd_qwen3_8b_hotpotqa_n4_n8llama8b-eagle-sharegptdetails_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.openthoughts_18K_solutions_R1_distill_Llama_8Bllama-3.1-tulu-3-8b-preference-mixture
Tulu 3 8B Preference Mixture
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. Some portions of the dataset are non-commercial. We present the mixture as a research artifact.
This mix is made up from the following preference datasets:
https://huggingface.co/datasets/allenai/tulu-3-sft-reused-off-policy
https://huggingface.co/datasets/allenai/tulu-3-sft-reused-on-policy-8b… See the full description on the dataset page: https://huggingface.co/datasets/allenai/llama-3.1-tulu-3-8b-preference-mixture.DRT-SFT-8B-training-data
DRT-SFT-8B Training Data
Paper: DRT: Dense Reasoning Trace for Efficient and Grounded Multimodal ReasoningCode: https://github.com/HIT-leaderone/DRT
This dataset contains the SFT training parquet shards used for DRT-SFT-8B.
Contents
20 parquet shards: Vision-R1_part_0.parquet ... Vision-R1_part_19.parquet
Total rows: 194,719
Columns: problem_id, content, role, image
Downloaded size: about 30.4 GiB
Notes
The parquet files are uploaded without… See the full description on the dataset page: https://huggingface.co/datasets/leaderonehit/DRT-SFT-8B-training-data.Transmem_ecsd_llama3_1_8b_hotpotqa_n4_n8wikipedia_qwen_8b
Vector Database Dataset
Generated embeddings dataset for vector database training and evaluation with multiple format support.
Dataset Summary
This dataset contains 500,000 text samples with high-quality vector embeddings generated using Qwen/Qwen3-Embedding-8B from the wikimedia/wikipedia dataset. The dataset is designed for vector database training, similarity search, and retrieval tasks.
Dataset Structure
Base dataset: 500,000 samples with embeddings
Query… See the full description on the dataset page: https://huggingface.co/datasets/maknee/wikipedia_qwen_8b.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.llama3.1_8b_inst_as_ver_gemma27b_it_math158_32gen_asyncQwen3-8B-BaldEagle-ShareGPTqy_syn_mix_8b_new
Dataset: qy_syn_mix_8b
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/qy_syn_mix_8b/stage_1/tmp.
qwen3-8b-activations-l20-l36
Qwen3 8B Activations for Layers 20 and 36
This dataset contains assistant-token residual activations harvested from Qwen/Qwen3-8B over 980000 training conversations from lmsys/lmsys-chat-1m.
We only generated for Layer 20 and 36 because each one costs 2TB and we simply cannot afford to store more :)
You can use this dataset to train SAEs, linear probes, other mech interp models etc, for Qwen3 8B.
We picked Qwen3 8B because this is a small part of a larger experiment to use feature… See the full description on the dataset page: https://huggingface.co/datasets/sammyliu/qwen3-8b-activations-l20-l36.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructca498c8bsae-activations-llama-3.1-8b-layer19-lmsys-chat-1m
SAE Feature Activations — Llama 3.1 8B Instruct, Layer 19 (LMSYS-Chat-1M)
This dataset contains Sparse Autoencoder (SAE) feature activations extracted from layer 19 of Meta's Llama 3.1 8B Instruct on conversations from LMSYS-Chat-1M.
It also has natural language explainations of features generated by GPT OSS 120B. See subset 4 for details.
The SAE used is Goodfire/Llama-3.1-8B-Instruct-SAE-l19, which decomposes layer-19 residual stream activations into interpretable sparse features.… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/sae-activations-llama-3.1-8b-layer19-lmsys-chat-1m.details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct
Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.WebShop-Qwen3-8B-Adaptive-Pivot-evalWebShop-Qwen3-8B-GiGPO-eval-eddieWebShop-Qwen3-8B-Adaptive-Midpoint-evalwhest-p2-bakev2-d8b-sketch-g00
whest-p2-bakev2 — cumulant sketches of 16×1024 ReLU MLPs (round 1, 2026-09-13)
Monte-Carlo cumulants of the pre-/post-activations of 1024-wide, 16-layer ReLU MLPs under
standard-normal inputs, stored as Ω-sketches (every net) plus a few dense n×n blocks
(validation nets). Companion code: ap_p2_bakev2_schema.py (seeds, Ω generator,
conversions, loader), ap_p2_bakev2.py (bake), ap_p2_bakev2_check.py (validation).
Schema version bakev2-r1-2026-09-13.
bake
family / nets
tier… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/whest-p2-bakev2-d8b-sketch-g00.xprmt-llama-3.1-8b-instruct-multijailWebShop-Qwen3-8B-SDAR-evalhubble-8b-unlearning-resultsWebShop-Qwen3-8B-Adaptive-Random-evalMATH_train_Llama3.1-8B-instruct_1sample_temp0.7WebShop-Qwen3-8B-Adaptive-Pivot-Anywin-evalWebShop-Qwen3-8B-GRPO-eval
