datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma4-e4b-rl100-hf-bf16-sdpa-topk128-overlay
Gemma 4 E4B RL100 top-k-128 target overlay
Precomputed off-policy distillation targets for the E4B-RL-step-100 to E2B experiment.
Source traces: JWei05/gemma4-e4b-rl100-topk128-traces at revision 2b6e49a0a456ee9d67b16a1dc61785562bee90c9
Direction: Gemma 4 E4B RL step 100 teacher to Gemma 4 E2B base student
Target engine: Hugging Face BF16 SDPA full forward
Width: top-k 128
Stored target token IDs: int32
Stored target log-probabilities: float16
Causal alignment: response token… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e4b-rl100-hf-bf16-sdpa-topk128-overlay.PI0-BF16_100eps_Selection_Nuts_Bolt_Aug_10_inf-recordingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 60,
"total_frames": 77648,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/CatGoesMeow/PI0-BF16_100eps_Selection_Nuts_Bolt_Aug_10_inf-recording.details_Kquant03__CognitiveFusion2-4x7B-BF16
Dataset Card for Evaluation run of Kquant03/CognitiveFusion2-4x7B-BF16
Dataset automatically created during the evaluation run of model Kquant03/CognitiveFusion2-4x7B-BF16.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Kquant03__CognitiveFusion2-4x7B-BF16.lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Nanashi-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Nanashi-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private.lm-eval-results-Kquant03-Cognito-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Cognito-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Cognito-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Cognito-2x7B-bf16-private.BF16kEval_FinEval_16k_fulleval__3args_ours-eval_rllm-eval-results-CultriX-NeuralTrix-bf16-private
Dataset Card for Evaluation run of CultriX/NeuralTrix-bf16
Dataset automatically created during the evaluation run of model CultriX/NeuralTrix-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-CultriX-NeuralTrix-bf16-private.Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16
Qwen3.8 DSpark PerfectBlend 5M paired BF16 features
Private, checksum-closed paired feature corpus for the Qwen3.8 Flash / 27B
DSpark transplant project. Repository: MJPansa/Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16.
The Hugging Face DatasetDict rows are a compact index. Each row points into
three immutable SafeTensor files in tensors/shard-NNNNN/ using exact token
and anchor offsets. This keeps the ~180 GB dense BF16 corpus resumable and
memory-mappable instead of duplicating… See the full description on the dataset page: https://huggingface.co/datasets/MJPansa/Qwen3.8-DSpark-PerfectBlend-5M-Paired-BF16.BF16kEval_FinEval_16k_fulleval__3args_r1-eval_0BF16kEval_FinEval_16k_fulleval__3args_rlonly-eval_rllm-eval-results-CultriX-NeuralTrixlaser-bf16-private
Dataset Card for Evaluation run of CultriX/NeuralTrixlaser-bf16
Dataset automatically created during the evaluation run of model CultriX/NeuralTrixlaser-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-CultriX-NeuralTrixlaser-bf16-private.GLM-5.2-BF16-KLD-Reference-Logits-20260618
GLM-5.2 BF16 KLD Reference Logits
Reference logits for local GLM-5.2 KLD checks.
Contents:
prefill/logits_0.safetensors: BF16 prefill prompt logits generated from
zai-org/GLM-5.2 with context length 2048, stride 512, one window.
decode/decode_teacher_bf16_ref_ctx2048_t17_20260618.safetensors: BF16
teacher-forced decode logits for prompt length 2048 and 17 decode tokens.
decode/decode_teacher_bf16_ref_ctx2048_t17_20260618.safetensors.json:
metadata for the decode reference.… See the full description on the dataset page: https://huggingface.co/datasets/festr2/GLM-5.2-BF16-KLD-Reference-Logits-20260618.GLM-5.2-BF16-KLD-Reference-Logits-20260708
GLM-5.2 BF16 KLD Reference Logits 20260708
This dataset contains the current GLM-5.2 BF16 reference prompt logits used for
the July 2026 vLLM/Blackwell KLD checks.
The reference cache is intended for candidate-side KLD comparisons without
rerunning the expensive BF16 reference pass.
Files
reference-logits/logits_0.safetensors
reference-logits/manifest.json
generation-log/config.env
generation-log/scoremode_kld.log
Reference Generation
Field… See the full description on the dataset page: https://huggingface.co/datasets/festr2/GLM-5.2-BF16-KLD-Reference-Logits-20260708.eurospeech-latents-bf16summarize_from_feedback_tldr3_unlabelled_vllm_dpo_costa_2.8b_bf16.yml_6e799_newQwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
fblgit__una-cybertron-7b-v2-bf16-details
Dataset Card for Evaluation run of fblgit/una-cybertron-7b-v2-bf16
Dataset automatically created during the evaluation run of model fblgit/una-cybertron-7b-v2-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__una-cybertron-7b-v2-bf16-details.fp16-bf16-wfp16a16kvfp16-wbf16abf16kvbf16
FP16/BF16 W-FP16 / A-FP16 / KV-FP16 and W-BF16 / A-BF16 / KV-BF16 (Llama-3.1-8B-Instruct)
This dataset contains FP16 and BF16 reference artifacts for Llama-3.1-8B-Instruct, stored in the same layout as the FP8 W-FP8/A-FP16/KV-FP8 artifact.
w_of_wfp16a16kvfp16_llama_31_8b/ — FP16 weights
The FP16 weights of Llama-3.1-8B-Instruct.
Stored per layer: layer_0.safetensors ... layer_31.safetensors + embeddings.safetensors.
The 7 linears per layer are stored as FP16… See the full description on the dataset page: https://huggingface.co/datasets/taehyeonkim/fp16-bf16-wfp16a16kvfp16-wbf16abf16kvbf16.BF16kEval_FinEval_16k_fulleval__3args_ours-eval_rl_countdown_6argopentq-qwen36-bf16-sidecar
Qwen3.6-27B BF16 Sidecar Runs
This dataset publishes BF16 sidecar outputs used to compare OpenTQ GGUF artifacts against the base model Qwen/Qwen3.6-27B on pinned practical mini-subsets.
The Dataset Viewer uses flattened Parquet tables so columns have stable types. The original raw JSON files remain in runs/<job_id>/ for reproducibility.
Tables
Config
Grain
Purpose
results
one row per benchmark sample
prompts, task IDs, deterministic BF16 outputs, score fields… See the full description on the dataset page: https://huggingface.co/datasets/zlaabsi/opentq-qwen36-bf16-sidecar.summarize_from_feedback_tldr3_unlabelled_vllm_dpo_costa_2.8b_bf16.yml_6e799dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5
dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-None-relabel-v5.Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details
Dataset Card for Evaluation run of Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
Dataset automatically created during the evaluation run of model Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details.KernelBench-bf16
KernelBench-bf16
KernelBench, with bfloat16 data type. Generated from li-plus/KernelBench.
vram-4b-bf16vram-2b-bf16lm-eval-results-Kquant03-Samlagast-7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Samlagast-7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Samlagast-7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Samlagast-7B-bf16-private.BF16kEval_FinEval_RL_rlonly-eval_rl_countdown_4argdev-instructed-deception-NVIDIA-Nemotron-3-Super-120B-A12B-BF16-NoneBF16kEval_FinEval_16k_fulleval__3args_rlonly-eval_rl_commonsenseQA
