datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
in1k_clip_qwen25vl_3b_448res_256tokens_new_merged_ptbeetle-merge-eval
Beetle merged models — benchmark evaluation against their parents
Minimal-pair benchmark accuracy for the Beetle merged models published in the
Mergeability org, scored against their own parent models and, where one
exists, the jointly-trained ceiling on the same harness.
The existing sweep datasets (Mergeability/merge-sweep-results,
Mergeability-2/mergeability-results) record merge quality in nats (NLL,
delta_floor, rel_damage, barrier, geometry). They contain no downstream… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/beetle-merge-eval.aim-activation-informed-merging
AIM: does activation-informed merging change what makes a merge work?
Headline
AIM does exactly what it claims, the targeting is what makes it work — and it changes
nothing about what predicts a good merge.
AIM is exactly what it says on the tin, and that is verifiable from public artefacts alone.
The published with-AIM checkpoints are recovered, to R² = 0.9992, as a closed-form
per-input-channel shrinkage of their baseline twins toward the base model, with ω̂ =… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/aim-activation-informed-merging.Merged-CWA
CWA Benchmark: A Seismic Dataset from Taiwan for Seismic Research
Dataset Description
This dataset includes a larger number of seismic events, especially high-magnitude. A comprehensive set of events collected by the
Central Weather Bureau in Taiwan. The CWA benchmark features over 40 attributes and ∼500,000 seismograms, providing
valuable data labels for various seismology-related tasks. In the future, we will keep updating the dataset to ensure its relevance and… See the full description on the dataset page: https://huggingface.co/datasets/NLPLabNTUST/Merged-CWA.mergebench-property-sweep
MergeBench property sweep
Pre-merge pairwise properties for every mergeable pair in the
MergeBench suite (40 checkpoints, 8 base families, 5 domains),
computed with the metric panel behind Figure F8 of the Heterogeneous Mergeability project.
Read this before you read a number
MergeBench publishes no pairwise merge. Every merge score in their release
(arXiv:2505.10833, Tables 8-17, and the two eval dumps in their
GitHub repo) is for a merge of all five domain… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/mergebench-property-sweep.gazeta.ru-mergedThis dataset is identical to the well-known Russian-language news dataset (Gazeta.ru)[https://huggingface.co/datasets/IlyaGusev/gazeta] by Ilya Gusev.
Therefore, details of the dataset should be sought at this link.
The main differences are two - a single dataset (training, testing and validation have been merged) and column names were changed for the convenience of building a news aggregator.
extrinsic-evaluations
Extrinsic evaluations — the union view
One tidy long-format table of every extrinsic (downstream, task-level) evaluation
produced across the 2026-08-26 mergeability workstreams, so that a single file answers
"how did model X score on benchmark Y" regardless of which experiment produced it.
The per-experiment datasets remain the authoritative record of their own methods,
figures and caveats. This is the union view, not a replacement, and it deliberately
carries no analysis of its… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/extrinsic-evaluations.sentiment_merged
Dataset Card for Sentiment Merged (SST-3, DynaSent R1, R2)
This is a dataset for 3-way sentiment classification of reviews (negative, neutral, positive). It is a merge of Stanford Sentiment Treebank (SST-3) and DynaSent Rounds 1 and 2, licensed under Apache 2.0 and Creative Commons Attribution 4.0 respectively.
Dataset Details
The SST-3, DynaSent R1, and DynaSent R2 datasets were randomly mixed to form a new dataset with 102,097 Train examples, 5,421 Validation… See the full description on the dataset page: https://huggingface.co/datasets/jbeno/sentiment_merged.yelp2018_merged_coredmerge-sweep-results
Heterogeneous Mergeability — merge sweep results
Per-merge outcomes from the sweep behind Heterogeneous Mergeability: A Quotient-Space Theory of
When Neural Networks Compose Across Tokenizers and Architectures. Every row is one merge that was
actually executed and scored: a pair of independently trained LMs, a merge operator, and an
alignment arm (merge naively, or align the two models into a shared frame first).
This dataset is a snapshot of a sweep that is still running.… See the full description on the dataset page: https://huggingface.co/datasets/Mergeability/merge-sweep-results.Olist-preprocessed-data-mergedmerged-dataMerged_QAs
Merged_QAs Dataset
Description
The Merged_QAs dataset combines Q&A pairs from two primary sources: StackOverflow Q&A related to various projects within the CNCF (Cloud Native Computing Foundation) landscape and the cncf-qa-dataset-for-llm-tuning designed for fine-tuning large language models (LLMs).
StackOverflow Q&A Dataset for Various Projects
This dataset includes questions and their corresponding answers sourced from StackOverflow discussions pertaining to… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/Merged_QAs.merged-data-v2
Info
This dataset is a merge of the following datasets:
flpelerin/openorca-alpaca-50k
sam-liu-lmi/databricks-dolly-15k-alpaca-style
TokenBender/roleplay_alpaca
vicgalle/alpaca-gpt4
CreitinGameplays/chat-assistant
CreitinGameplays/filter
wikitext-2-raw-v1-merged-45ktombench_merged
TomBench Merged Dataset (Exact Matching)
This dataset contains the merged results of TomBench evaluation with the original TomBench dataset, using exact string matching.
Dataset Statistics
Total records: 2860
Exact matches: 2860
Manual matches: 0
Average model score: 0.5066
Matching Strategy
This version uses exact string matching after text normalization:
Remove extra whitespace and normalize formatting
Match stories exactly between datasets
Report any… See the full description on the dataset page: https://huggingface.co/datasets/ycfNTU/tombench_merged.merged_yellow_tripdata
Taxi-Demand-Fare-Prediction-Dataset
About Dataset
This dataset contains records of taxi trips from New York City, including both yellow and green taxi trip data. The data was provided to the NYC Taxi and Limousine Commission (TLC) by technology providers authorized under the Taxicab & Livery Passenger Enhancement Programs (TPEP/LPEP). Please note that TLC did not create this data and makes no representations regarding its accuracy.
Key Information:
TLC Trip… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/merged_yellow_tripdata.fma-merged-metadata-and-featuresid-hoax-report-merge-v3We do not maintain this repository further. For accessing the most recent Indonesian Fake News dataset that we created, please visit BRIN's dataverse: https://data.brin.go.id/dataset.xhtml?persistentId=hdl:20.500.12690/RIN/7QBRKQ
The dataset is taken from nlp-brin-id/id-hoax-report-merge-v2 by filtering out null samples.
re-merged-pf-2merge-textclassifymerged2merged113-go-emotions-mergereal-estate-data-mergedmerged_characters_tinyllamawiki_mergeddata_mergedVenusX_Res_Epi_MP_Merged_Subseqsa_merge_v1
