datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mechanism-merging-exp089-handoff
Mechanism Merging Exp089 analysis handoff
这是 weekly_experiment_report_20260812.md 对应版本的公开研究交付仓库。它保存报告第 8、10、11 章
所用的模型、训练数据、逐样本评测轨迹和分析实物;代码、可直接浏览的报告与图片位于私有 GitHub
仓库 xlxcomputer/mechanism-merging-exp089-repro。
四个分析输入模型
目录
角色
冻结点
训练成本口径
checkpoints/anchor_mixed_stage1_step10014/huggingface/
三种方法共同且唯一的 mixed Stage-1 anchor
step 10014
共同成本,不进入方法横轴
checkpoints/mix-rl_step1036/huggingface/
Mix-RL final
step 1036
online FLOPs 100210702010088587264… See the full description on the dataset page: https://huggingface.co/datasets/2041Xu/mechanism-merging-exp089-handoff.livecodebench-merging-leaderboard
LiveCodeBench v6 Evaluation Leaderboard
Evaluation results for cross-capability merging of OLMo-3 and OLMo-3.1 RL-Zero models on 454 coding problems.
Evaluation
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
Leaderboard
Model
pass@4
pass@1
Loop Rate
Qwen/Qwen3-4B-Thinking-2507
54.6%
45.4%
0.4%
pmahdavi/Olmo-3-7B-Think-Math-Code… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/livecodebench-merging-leaderboard.AI-Helps-Finding-Best-Merging-LLMs
Dataset Card for AI Helps Finding Best Merging LLMs
Dataset Summary
AI Helps Finding Best Merging LLMs is a prompt-response comparison dataset created by manually submitting the same user-written evaluation template to multiple LLM applications and collecting their responses.
The creator and founder of WithIn Us Ai (Guy Edward DuGan II) known as gss1147 wrote a structured ranking template and fed it to each LLM individually in its own app environment. The… See the full description on the dataset page: https://huggingface.co/datasets/11-47/AI-Helps-Finding-Best-Merging-LLMs.aime2025-merging-leaderboard
AIME 2025 Evaluation Leaderboard
Evaluation results for 8 models on AIME 2025 with 30 problems and 32 rollouts per problem.
Evaluation
We followed the evaluation guidelines and prompts from OLMo 3. Best effort was made to ensure reported numbers are as accurate as possible.
Code: pmahdavi/modal-eval
Leaderboard
Model
pass@32
avg@32
Loop Rate
allenai/Olmo-3.1-7B-RL-Zero-Math
80.0%
38.0%
2.9%
pmahdavi/Olmo-3.1-7B-Math-Code
73.3%
36.4%
2.4%… See the full description on the dataset page: https://huggingface.co/datasets/pmahdavi/aime2025-merging-leaderboard.model-merging-papers
Model Merging Papers — FineSet
A research-paper dataset on Model Merging Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Model Merging Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score float… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/model-merging-papers.
