datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
navsim-metric-caches-from-a100p-data-a100-2p-data-a100-5p-data-a100-3Llama-3.2-1B-Instruct-best_of_n-completions-a100p-data-a100p-data-a100-4Llama-3.2-1B-Instruct-beam_search_16-completions-a100pallasbench-robust-gpu-a100
PallasBench: Robust Pallas GPU Kernel Benchmark (A100)
39/45 kernels passing on NVIDIA A100 80GB -- the first GPU-focused evaluation of JAX Pallas kernels.
What is this?
PallasBench is a suite of 45 JAX Pallas kernels across 3 difficulty levels. The original kernels were designed for TPU and failed on GPU because Pallas compiles to Triton on NVIDIA hardware, which has strict block size limits that TPU's Mosaic compiler does not.
We fixed all 45 kernels for GPU… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/pallasbench-robust-gpu-a100.Olmo-1B-Instruct-best_of_64-completions-a100dataset_A100
Disclaimer:
this dataset is curated for NeurIPS 2023 LLM efficiency challange, and currently work in progress. Please use at your own risk.
Data composition:
All data were derived from the training set portion of the open source dataset.
Data Sources:
-dolly: https://huggingface.co/datasets/databricks/databricks-dolly-15k
-cnn_dailymail: https://huggingface.co/datasets/cnn_dailymail
-mmlu: https://huggingface.co/datasets/cais/mmlu
-bbq:… See the full description on the dataset page: https://huggingface.co/datasets/zhongshupeng/dataset_A100.graded_deepseek_ai_deepseek_r1_distill_qwen_7b_a100cc0bOlmo-1B-Instruct-best_of_4-completions-a100Llama-3.2-1B-Instruct-beam_search_4-completions-a100Llama-3.2-1B-Instruct-dvts-4-completions-a100Olmo-1B-Instruct-beam_search_4-completions-a100tulu-3-70k-A100-fastOlmo-1B-Instruct-beam_search_16-completions-a100Olmo-1B-Instruct-dvts-4-completions-a100Olmo-1B-Instruct-dvts-16-completions-a100good_predicts_01_a100Llama-3.2-1B-Instruct-beam_search_64-completions-a100good_predicts_09_a100Olmo-1B-Instruct-dvts-64-completions-a100Olmo-1B-Instruct-best_of_16-completions-a100tmax-hosted-qwen35-4b-batch16x16-a100-ilc-job113825
Hosted TMAX trajectories — a100-ilc
This public, ungated dataset contains the validated output of run a100-ilc-113825
(Slurm job 113825) using Qwen/Qwen3.5-4B at revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, served as qwen3.5-4b.
Run configuration
Hardware: 2 × NVIDIA A100-SXM4-80GB on ampere1.stanford.edu
Parallelism: DP=2, TP=1
Tasks: 16; attempts per task: 16; rows: 256
Sampling: temperature=0.6, top_p=1.0
Limits: 8192 tokens/turn, 16 turns… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/tmax-hosted-qwen35-4b-batch16x16-a100-ilc-job113825.Llama-3.2-1B-Instruct-best_of_4-completions-a100Llama-3.2-1B-Instruct-dvts-64-completions-a100Llama-3.2-1B-Instruct-best_of_16-completions-a100Llama-3.2-1B-Instruct-dvts-completions-a100
