datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers-merge-experimentslm-eval-results-automerger-Experiment28Yam-7B-private
Dataset Card for Evaluation run of automerger/Experiment28Yam-7B
Dataset automatically created during the evaluation run of model automerger/Experiment28Yam-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-automerger-Experiment28Yam-7B-private.zoya-image-1-experiments
ZOYA IMAGE-1 — Reproducible GGUF Experiments
Purpose
This dataset stores reproducible ZOYA IMAGE-1 image-generation
experiments together with the exact generation parameters,
model identities, SHA256 fingerprints, and validation reports.
The package is designed for controlled comparisons where the
tested variable is changed explicitly and all other relevant
variables remain fixed.
Current baseline
Experiment ID: ZOYA_PHASE0_BASELINE_00001… See the full description on the dataset page: https://huggingface.co/datasets/tigerking009/zoya-image-1-experiments.lm-eval-results-MaziyarPanahi-YamshadowInex12_Experiment26T3q-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Experiment26T3q
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Experiment26T3q
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowInex12_Experiment26T3q-private.lm-eval-results-MaziyarPanahi-MeliodasPercival_01_Experiment26T3q-private
Dataset Card for Evaluation run of MaziyarPanahi/MeliodasPercival_01_Experiment26T3q
Dataset automatically created during the evaluation run of model MaziyarPanahi/MeliodasPercival_01_Experiment26T3q
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-MeliodasPercival_01_Experiment26T3q-private.lm-eval-results-MaziyarPanahi-Experiment26Yamshadow_Ognoexperiment27Multi_verse_model-private
Dataset Card for Evaluation run of MaziyarPanahi/Experiment26Yamshadow_Ognoexperiment27Multi_verse_model
Dataset automatically created during the evaluation run of model MaziyarPanahi/Experiment26Yamshadow_Ognoexperiment27Multi_verse_model
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-Experiment26Yamshadow_Ognoexperiment27Multi_verse_model-private.emergent-misalignment-experiment-1-data
Emergent Misalignment Experiment 1 Data Artifacts
Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment.
This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.lm-eval-results-MaziyarPanahi-Experiment26Yam_Ognoexperiment27Multi_verse_model-private
Dataset Card for Evaluation run of MaziyarPanahi/Experiment26Yam_Ognoexperiment27Multi_verse_model
Dataset automatically created during the evaluation run of model MaziyarPanahi/Experiment26Yam_Ognoexperiment27Multi_verse_model
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-Experiment26Yam_Ognoexperiment27Multi_verse_model-private.lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Experiment26T3q-private
Dataset Card for Evaluation run of MaziyarPanahi/M7Yamshadowexperiment28_Experiment26T3q
Dataset automatically created during the evaluation run of model MaziyarPanahi/M7Yamshadowexperiment28_Experiment26T3q
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Experiment26T3q-private.lm-eval-results-MaziyarPanahi-YamshadowStrangemerges_32_Experiment24Ognoexperiment27-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowStrangemerges_32_Experiment24Ognoexperiment27
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowStrangemerges_32_Experiment24Ognoexperiment27
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowStrangemerges_32_Experiment24Ognoexperiment27-private.lm-eval-results-automerger-Experiment29Pastiche-7B-private
Dataset Card for Evaluation run of automerger/Experiment29Pastiche-7B
Dataset automatically created during the evaluation run of model automerger/Experiment29Pastiche-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-automerger-Experiment29Pastiche-7B-private.grok-1-ternary-quant-experiments
Grok-1 SAAQ quantization / route-preservation experiments
Dataset author: Raul Montoya Cardenas (rmems)
SAAQ stands for Spiking Adaptive Activity Quantization, a term coined by
the dataset author.
Attribution: Grok Build: Grok 4.5 (high) packaged the original 2026-08-10
dataset. Codex: GPT-5.6-Sol (OpenAI) implemented, executed, validated, and published the
canonical issue #85 v4 evidence added on 2026-08-24.
Personal research measuring route preservation when packing open… See the full description on the dataset page: https://huggingface.co/datasets/rmems/grok-1-ternary-quant-experiments.experiment-001-budget-boxed
KamiBench Experiment 001 — budget-boxed agents in Kamigotchi
Complete agentic traces from experiment 001, the KamiBench
calibration run: three LLM agents dropped into
Kamigotchi, a live, persistent, on-chain
world (Yominet), each with a $10 inference budget, a 7-day
wall-clock cap, and no further human contact. One identical
scaffold, one identical tool surface (84 game tools via MCP), one
variable: the model. The agents schedule their own wake-ups, keep
their own files, and act… See the full description on the dataset page: https://huggingface.co/datasets/KamiBench/experiment-001-budget-boxed.lm-eval-results-ChaoticNeutrals-Prima-LelantaclesV7-experimental-7b-private
Dataset Card for Evaluation run of ChaoticNeutrals/Prima-LelantaclesV7-experimental-7b
Dataset automatically created during the evaluation run of model ChaoticNeutrals/Prima-LelantaclesV7-experimental-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ChaoticNeutrals-Prima-LelantaclesV7-experimental-7b-private.sometimesanotion__IF-reasoning-experiment-80-details
Dataset Card for Evaluation run of sometimesanotion/IF-reasoning-experiment-80
Dataset automatically created during the evaluation run of model sometimesanotion/IF-reasoning-experiment-80
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__IF-reasoning-experiment-80-details.demo-experiment-two
Demo Experiment
A demo experiment to test the builder flow with various question types.
Dataset Overview
Property
Value
Run ID
036df3a6-6f8c-45fb-979d-be4f990bd0bf
Status
completed
Created
12/23/2025, 9:11:08 PM
Generator
LocalBench v0.1.0
Statistics
Metric
Value
Total Generations
10
Successful
10 (100.0%)
Failed
0
Average Latency
2803ms
Total Duration
25.2s
Configuration
Models… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/demo-experiment-two.Pinkstack__Superthoughts-lite-1.8B-experimental-o1-details
Dataset Card for Evaluation run of Pinkstack/Superthoughts-lite-1.8B-experimental-o1
Dataset automatically created during the evaluation run of model Pinkstack/Superthoughts-lite-1.8B-experimental-o1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Pinkstack__Superthoughts-lite-1.8B-experimental-o1-details.experiment-llm__exp-3-q-r-details
Dataset Card for Evaluation run of experiment-llm/exp-3-q-r
Dataset automatically created during the evaluation run of model experiment-llm/exp-3-q-r
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/experiment-llm__exp-3-q-r-details.DoppelReflEx__MN-12B-Mimicore-Orochi-v3-Experiment-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Orochi-v3-Experiment
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Orochi-v3-Experiment
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Orochi-v3-Experiment-details.DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-4-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-4
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-4-details.DoppelReflEx__MN-12B-LilithFrame-Experiment-2-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame-Experiment-2
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame-Experiment-2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-Experiment-2-details.DoppelReflEx__MN-12B-LilithFrame-Experiment-3-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame-Experiment-3
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame-Experiment-3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-Experiment-3-details.experiment_real_bitsandbytes
/hub_data4/seohyun/saves/ecva_instruct/full/sft/checkpoint-350 · happy8825/valid_ecva_clean results
Model: /hub_data4/seohyun/saves/ecva_instruct/full/sft/checkpoint-350
Dataset: happy8825/valid_ecva_clean
Generated: 2026-01-08 12:14:39Z
Metrics
Metric
Value
Total samples
924
With GT
0
Parsed answers
0
Top-1 accuracy
0
Recall@5
0
MRR
0
The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/experiment_real_bitsandbytes.sethuiyer__LlamaZero-3.1-8B-Experimental-1208-details
Dataset Card for Evaluation run of sethuiyer/LlamaZero-3.1-8B-Experimental-1208
Dataset automatically created during the evaluation run of model sethuiyer/LlamaZero-3.1-8B-Experimental-1208
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sethuiyer__LlamaZero-3.1-8B-Experimental-1208-details.DoppelReflEx__MN-12B-Mimicore-Orochi-v4-Experiment-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Orochi-v4-Experiment
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Orochi-v4-Experiment
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Orochi-v4-Experiment-details.DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-3-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-3
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-3-details.DoppelReflEx__MN-12B-LilithFrame-Experiment-4-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-LilithFrame-Experiment-4
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-LilithFrame-Experiment-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-LilithFrame-Experiment-4-details.DoppelReflEx__MN-12B-Mimicore-Orochi-v2-Experiment-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Orochi-v2-Experiment
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Orochi-v2-Experiment
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Orochi-v2-Experiment-details.DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-1-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-1
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-1-details.DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-2-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-2
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-2-details.
