diagnostics
Linguistic-Diagnostics-Pragmatics
LINDSEA Pragmatics
LINDSEA Pragmatics is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, pragmatics in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Pragmatics is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
LINDSEA Pragmatics only has an Indonesian (id) split, with additional splits containing… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Pragmatics.Linguistic-Diagnostics-Syntax
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
LINDSEA Syntax only has an Indonesian (id) split, with additional splits containing fewshot examples. Below… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax.grokking-diagnostics-runs
Grokking Diagnostics Runs
Per-run training records and aggregate fits backing:
Weight Decay Regimes in Grokking Transformers: Cheap Online Diagnostics
Lucky Verma. Independent Researcher. 2026.
Paper ·
DOI ·
PDF ·
Code
Contents
The paper provenance indexes 1,792 paper-run records: 1,442 records from the
main paper-integrated run tree plus 350 cross-architecture scope-probe records.
This dataset repository also includes convenience subset mirrors, so the… See the full description on the dataset page: https://huggingface.co/datasets/lucky-verma/grokking-diagnostics-runs.Linguistic-Diagnostics-Syntax-Judge
LINDSEA Syntax
LINDSEA Syntax is a linguistic diagnostic from BHASA that evaluates a model's understanding of linguistic phenomena, syntax in particular, for Indonesian.
Supported Tasks and Leaderboards
LINDSEA Syntax is designed for evaluating chat or instruction-tuned large language models (LLMs).
Languages
Indonesian (id)
Dataset Details
Data Sources
Data Source
License
Language/s
Split/s
CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Linguistic-Diagnostics-Syntax-Judge.crossarch-1b-diagnostics
Cross-architecture mergeability diagnostics for five ~1B monolingual LMs
A third model family for the mergeability project, alongside Goldfish and Beetle/MergeBench.
Read this first: what "merging" means for these five models
Five independently trained ~1B monolingual models were requested: Pythia-1.4B (EN),
Zh-Pythia-1.4B (ZH), Tucano-1b1 (PT), Bielik-1.5B-v3 (PL), Minerva-1B (IT). They differ in
architecture family, hidden dimension (1536 vs 2048), depth (16 /… See the full description on the dataset page: https://huggingface.co/datasets/Cross-Mergeability/crossarch-1b-diagnostics.painting-restoration-eval-diagnostics
