datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers-merge-experimentsprompt-swap-mixed12-5xlr-e1-mxfp4-mergedprompt-swap-mixed12-5xlr-e2-mxfp4-mergedsat-vl-sft-postprocessed-merged-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-postprocessed-merged-v1.prompt-swap-medium12-e2-mxfp4-mergedgrad_clip0.28_mergedraw-mergedguanaco_belle_merge_v1.0Thanks for Guanaco Dataset and Belle Dataset
This dataset was created by merging the above two datasets in a certain format so that they can be used for training our code Chinese-Vicuna
lm-eval-results-alnrg2arg-blockchainlabs_7B_merged_test2_4-private
Dataset Card for Evaluation run of alnrg2arg/blockchainlabs_7B_merged_test2_4
Dataset automatically created during the evaluation run of model alnrg2arg/blockchainlabs_7B_merged_test2_4
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-alnrg2arg-blockchainlabs_7B_merged_test2_4-private.merged-trecdlmerge-accuracy
merge-accuracy — does aligning a merge improve DOWNSTREAM ACCURACY?
The mergeability line of work measures merge obstruction in nats/token. This dataset supplies
the missing axis: task accuracy, on real released models, for the merge recipe practitioners
actually run — the chat-vector recipe
theta_new = theta_fork + lambda * ( theta_instruct - theta_base )
with meta-llama/Llama-3.1-8B, its official Instruct release, and three community
continued-pretrained language forks… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/merge-accuracy.LoRA-Merge-Imagestextbook-openstax-yawp-mergecache_stage1_mergedlm-eval-results-chlee10-T3Q-Merge-Mistral7B-private
Dataset Card for Evaluation run of chlee10/T3Q-Merge-Mistral7B
Dataset automatically created during the evaluation run of model chlee10/T3Q-Merge-Mistral7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Merge-Mistral7B-private.lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-private.fable_5_distillation_merged_cleaned_25k
Claude Fable 5 Distillation Dataset
25,719 high-quality distilled examples for training LLMs to mimic Claude Fable 5's reasoning style — featuring multi-step chain-of-thought with <think> tags across 23+ technical domains.
This dataset captures the distinctive reasoning patterns of Claude Fable 5 (Anthropic's Mythos-class model released June 2026): systematic decomposition, first-principles analysis, self-verification, alternative consideration, and synthesis.… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/fable_5_distillation_merged_cleaned_25k.claudesidian-behaviors-merged
Claudesidian Merged Behavioral Dataset
Dataset Description
This dataset contains 1,852 synthetic training examples demonstrating 8 different behavioral patterns for training language models to use the Claudesidian-MCP toolset effectively with Obsidian vaults.
The dataset is specifically formatted for KTO (Kahneman-Tversky Optimization) preference learning with properly interleaved positive and negative examples.
Behavioral Categories
This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/claudesidian-behaviors-merged.lm-eval-results-MiniMoog-Mergerix-7b-v0.5-private
Dataset Card for Evaluation run of MiniMoog/Mergerix-7b-v0.5
Dataset automatically created during the evaluation run of model MiniMoog/Mergerix-7b-v0.5
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MiniMoog-Mergerix-7b-v0.5-private.lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v4-test-private.lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-v2-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-v2
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-v2-private.lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-private.lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v2-private.Matplotlib_Seaborn_merged_prompt_completion_10klm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3-private
Dataset Card for Evaluation run of eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3
Dataset automatically created during the evaluation run of model eren23/ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-eren23-ogno-monarch-jaskier-merge-7b-OH-PREF-DPO-v3-private.Evol-instruct-mergejust merging a few evol instruct datasets the pros made
numina-0.02B-drop-merge-llama
Dataset: numina-0.02B-drop-merge-llama
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_llama3/numina-0.02B-drop-merge/stage_1.
numina-0.2B-drop-entropy50-merge
Dataset: numina-0.2B-drop-entropy50-merge
This dataset was uploaded from /mnt/yulan_pretrain/mount/data_final_train_qwen3/numina-0.2B-drop-entropy50-merge/stage_1.
math_merged_cot_sol_pair_mixedPair Type Breakdown:
Correct-Incorrect (C-I) Pairs: 5500
C-I with correct first ('[1]'): 2750
C-I with correct second ('[2]'): 2750
Correct-Correct (C-C) Pairs (Target: 2750, Max Diff: 150): 2750
C-C pairs from 'all_correct' problems: 906
Incorrect-Incorrect (I-I) Pairs (Target: 2750, Max Diff: 150): 2750
I-I pairs from 'all_incorrect' problems: 1156
PATH-VQA
