datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test
mazhdrak/test — Mixed Instruction Dataset
A personal mixed-domain instruction dataset compiled from JSON files, CSV tables, Word documents, hardware reports, chatbot histories, and production manuals.
Languages: English + Bulgarian. Built for fine-tuning, RAG, and LLM evaluation.
Dataset Stats
Subset
File
Records
Description
Master (all)
train.jsonl
811
Complete unified dataset
Chat Exports
chat_exports.jsonl
235
Tabular Data
tabular_data.jsonl
233… See the full description on the dataset page: https://huggingface.co/datasets/mazhdrak/test.lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Strangemerges_30Experiment26-private
Dataset Card for Evaluation run of MaziyarPanahi/M7Yamshadowexperiment28_Strangemerges_30Experiment26
Dataset automatically created during the evaluation run of model MaziyarPanahi/M7Yamshadowexperiment28_Strangemerges_30Experiment26
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Strangemerges_30Experiment26-private.maze-17x17-maxrl-rolloutsazure_docs_fulllm-eval-results-MaziyarPanahi-Topxtral-4x7B-v0.1-private
Dataset Card for Evaluation run of MaziyarPanahi/Topxtral-4x7B-v0.1
Dataset automatically created during the evaluation run of model MaziyarPanahi/Topxtral-4x7B-v0.1
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-Topxtral-4x7B-v0.1-private.lm-eval-results-MaziyarPanahi-YamshadowInex12_Experiment26T3q-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Experiment26T3q
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Experiment26T3q
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowInex12_Experiment26T3q-private.lm-eval-results-MaziyarPanahi-MeliodasPercival_01_Experiment26T3q-private
Dataset Card for Evaluation run of MaziyarPanahi/MeliodasPercival_01_Experiment26T3q
Dataset automatically created during the evaluation run of model MaziyarPanahi/MeliodasPercival_01_Experiment26T3q
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-MeliodasPercival_01_Experiment26T3q-private.gpt_oss_maze_acts_120_m11_v1lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowInex12_Multi_verse_modelExperiment28-private.lm-eval-results-MaziyarPanahi-Calme-7B-Instruct-v0.9-private
Dataset Card for Evaluation run of MaziyarPanahi/Calme-7B-Instruct-v0.9
Dataset automatically created during the evaluation run of model MaziyarPanahi/Calme-7B-Instruct-v0.9
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-Calme-7B-Instruct-v0.9-private.lm-eval-results-MaziyarPanahi-Experiment26Yamshadow_Ognoexperiment27Multi_verse_model-private
Dataset Card for Evaluation run of MaziyarPanahi/Experiment26Yamshadow_Ognoexperiment27Multi_verse_model
Dataset automatically created during the evaluation run of model MaziyarPanahi/Experiment26Yamshadow_Ognoexperiment27Multi_verse_model
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-Experiment26Yamshadow_Ognoexperiment27Multi_verse_model-private.maze2d_easy_plain_ordered
maze2d_easy_plain_ordered
BAGEL VLM-Gym world-model dataset (maze2d / plain).
Maze2D easy native-256, random start/goal, stop-required; non-CoT; ordered with a held-out test split.
layout: Gzipped-JSONL shards under training/ (train) and testing/ (held-out); base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching maze2d checkpoint(s) under the companion model org; CoT and non-CoT
variants share the same… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/maze2d_easy_plain_ordered.lite_SFT_train_7lanlite_SFT_train_7lan
This dataset is designed for training and fine-tuning language models with multilingual question-answer pairs in seven languages: English (en), Arabic (ar), Urdu (ur), Persian/Farsi (fa), Indonesian (id), Turkish (tr), and Bengali (bn). The database contains over 1,300 high-quality Q&A entries, each fully translated across all seven languages. Each entry includes:
Original question-answer pairs in English
Translated versions for Arabic, Urdu, Persian, Indonesian, Turkish… See the full description on the dataset page: https://huggingface.co/datasets/mazrba/lite_SFT_train_7lan.lm-eval-results-MaziyarPanahi-TheTop-5x7B-Instruct-S5-v0.1-private
Dataset Card for Evaluation run of MaziyarPanahi/TheTop-5x7B-Instruct-S5-v0.1
Dataset automatically created during the evaluation run of model MaziyarPanahi/TheTop-5x7B-Instruct-S5-v0.1
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-TheTop-5x7B-Instruct-S5-v0.1-private.lm-eval-results-MaziyarPanahi-Experiment26Yam_Ognoexperiment27Multi_verse_model-private
Dataset Card for Evaluation run of MaziyarPanahi/Experiment26Yam_Ognoexperiment27Multi_verse_model
Dataset automatically created during the evaluation run of model MaziyarPanahi/Experiment26Yam_Ognoexperiment27Multi_verse_model
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-Experiment26Yam_Ognoexperiment27Multi_verse_model-private.lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Experiment26T3q-private
Dataset Card for Evaluation run of MaziyarPanahi/M7Yamshadowexperiment28_Experiment26T3q
Dataset automatically created during the evaluation run of model MaziyarPanahi/M7Yamshadowexperiment28_Experiment26T3q
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Experiment26T3q-private.lm-eval-results-MaziyarPanahi-YamshadowStrangemerges_32_Experiment24Ognoexperiment27-private
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowStrangemerges_32_Experiment24Ognoexperiment27
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowStrangemerges_32_Experiment24Ognoexperiment27
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-YamshadowStrangemerges_32_Experiment24Ognoexperiment27-private.Tiny-Maze-Mock-GRPOMaziyarPanahi__calme-2.1-qwen2.5-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.1-qwen2.5-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.1-qwen2.5-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.1-qwen2.5-72b-details.MaziyarPanahi__calme-2.2-rys-78b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.2-rys-78b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.2-rys-78b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.2-rys-78b-details.MaziyarPanahi__calme-2.2-qwen2.5-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.2-qwen2.5-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.2-qwen2.5-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.2-qwen2.5-72b-details.MaziyarPanahi__calme-2.4-qwen2-7b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.4-qwen2-7b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.4-qwen2-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.4-qwen2-7b-details.MaziyarPanahi__calme-2.3-llama3.1-70b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3.1-70b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3.1-70b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.3-llama3.1-70b-details.maze-solver-benchmark
Maze Solver Benchmark — BFS Shortest Path by Maze Size
Search effort and path length for the LK Forge Maze Solver,
which generates a perfect maze with a recursive-backtracker and finds the unique shortest
path with breadth-first search. 200 seeded mazes per size.
How the solver works: https://lkforge.com/tools/puzzles/maze-solver/how-the-maze-solver-works.html
Try it: https://lkforge.com/tools/puzzles/maze-solver/
Producer: LK Forge — client-side AI games, solvers, and tools.… See the full description on the dataset page: https://huggingface.co/datasets/LKForge/maze-solver-benchmark.MaziyarPanahi__calme-2.1-qwen2-72b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.1-qwen2-72b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.1-qwen2-72b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.1-qwen2-72b-details.MaziyarPanahi__calme-2.1-rys-78b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.1-rys-78b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.1-rys-78b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.1-rys-78b-details.MaziyarPanahi__calme-3.3-baguette-3b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-3.3-baguette-3b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-3.3-baguette-3b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-3.3-baguette-3b-details.maze2d_easy_native256_stepmsg_fixedstart_plain
maze2d_easy_native256_stopreq_stepmsg_fixedstart_ordered / maze2d_easy_native256_stopreq_stepmsg_fixedstart_cot_ordered
Built by build_maze2d_native256_ordered_pair.py at 20260606_stepmsg_fixedstart_v2.
Manifest version: maze2d_native256_stopreq_stepmsg_fixedstart_plain_cot_ordered100k_v2. CoT policy: maze2d_native256_stopreq_stepmsg_fixedstart_cot_branch_v2.
Plain and CoT rows share the same retained episode manifests and order for each split.
MaziyarPanahi__Qwen2-7B-Instruct-v0.8-details
Dataset Card for Evaluation run of MaziyarPanahi/Qwen2-7B-Instruct-v0.8
Dataset automatically created during the evaluation run of model MaziyarPanahi/Qwen2-7B-Instruct-v0.8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__Qwen2-7B-Instruct-v0.8-details.MaziyarPanahi__calme-2.4-rys-78b-details
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.4-rys-78b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.4-rys-78b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/MaziyarPanahi__calme-2.4-rys-78b-details.
