datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Nanashi-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Nanashi-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Nanashi-2x7B-bf16-private.lm-eval-results-Kquant03-Cognito-2x7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Cognito-2x7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Cognito-2x7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Cognito-2x7B-bf16-private.lm-eval-results-CultriX-NeuralTrix-bf16-private
Dataset Card for Evaluation run of CultriX/NeuralTrix-bf16
Dataset automatically created during the evaluation run of model CultriX/NeuralTrix-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-CultriX-NeuralTrix-bf16-private.lm-eval-results-CultriX-NeuralTrixlaser-bf16-private
Dataset Card for Evaluation run of CultriX/NeuralTrixlaser-bf16
Dataset automatically created during the evaluation run of model CultriX/NeuralTrixlaser-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-CultriX-NeuralTrixlaser-bf16-private.GLM-5.2-BF16-KLD-Reference-Logits-20260618
GLM-5.2 BF16 KLD Reference Logits
Reference logits for local GLM-5.2 KLD checks.
Contents:
prefill/logits_0.safetensors: BF16 prefill prompt logits generated from
zai-org/GLM-5.2 with context length 2048, stride 512, one window.
decode/decode_teacher_bf16_ref_ctx2048_t17_20260618.safetensors: BF16
teacher-forced decode logits for prompt length 2048 and 17 decode tokens.
decode/decode_teacher_bf16_ref_ctx2048_t17_20260618.safetensors.json:
metadata for the decode reference.… See the full description on the dataset page: https://huggingface.co/datasets/festr2/GLM-5.2-BF16-KLD-Reference-Logits-20260618.GLM-5.2-BF16-KLD-Reference-Logits-20260708
GLM-5.2 BF16 KLD Reference Logits 20260708
This dataset contains the current GLM-5.2 BF16 reference prompt logits used for
the July 2026 vLLM/Blackwell KLD checks.
The reference cache is intended for candidate-side KLD comparisons without
rerunning the expensive BF16 reference pass.
Files
reference-logits/logits_0.safetensors
reference-logits/manifest.json
generation-log/config.env
generation-log/scoremode_kld.log
Reference Generation
Field… See the full description on the dataset page: https://huggingface.co/datasets/festr2/GLM-5.2-BF16-KLD-Reference-Logits-20260708.Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmarkBenchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
692
Incorrect
295
Errors
13
Total samples
1000
Python tool calls
1508
Total completion tokens
3,806,045
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}
fblgit__una-cybertron-7b-v2-bf16-details
Dataset Card for Evaluation run of fblgit/una-cybertron-7b-v2-bf16
Dataset automatically created during the evaluation run of model fblgit/una-cybertron-7b-v2-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fblgit__una-cybertron-7b-v2-bf16-details.Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details
Dataset Card for Evaluation run of Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
Dataset automatically created during the evaluation run of model Edgerunners/meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Edgerunners__meta-llama-3-8b-instruct-hf-ortho-baukit-34fail-3000total-bf16-details.vram-4b-bf16vram-2b-bf16lm-eval-results-Kquant03-Samlagast-7B-bf16-private
Dataset Card for Evaluation run of Kquant03/Samlagast-7B-bf16
Dataset automatically created during the evaluation run of model Kquant03/Samlagast-7B-bf16
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kquant03-Samlagast-7B-bf16-private.vram-9b-bf16mlx-community__Mistral-Small-24B-Instruct-2501-bf16-details
Dataset Card for Evaluation run of mlx-community/Mistral-Small-24B-Instruct-2501-bf16
Dataset automatically created during the evaluation run of model mlx-community/Mistral-Small-24B-Instruct-2501-bf16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlx-community__Mistral-Small-24B-Instruct-2501-bf16-details.Kquant03__CognitiveFusion2-4x7B-BF16-details
Dataset Card for Evaluation run of Kquant03/CognitiveFusion2-4x7B-BF16
Dataset automatically created during the evaluation run of model Kquant03/CognitiveFusion2-4x7B-BF16
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Kquant03__CognitiveFusion2-4x7B-BF16-details.neopolita__jessi-v0.1-bf16-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.1-bf16-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.1-bf16-falcon3-7b-instruct
The dataset is composed of 77 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.1-bf16-falcon3-7b-instruct-details.
