datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-bobofrut-ladybird-base-7B-v8-private
Dataset Card for Evaluation run of bobofrut/ladybird-base-7B-v8
Dataset automatically created during the evaluation run of model bobofrut/ladybird-base-7B-v8
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bobofrut-ladybird-base-7B-v8-private.LadderBench
LadderBench
A graded capability ladder for evaluating (large) language models: 150
multiple-choice questions organized in 5 difficulty levels (30 each),
strictly ordered from easy to expert. Because every level is scored
separately, LadderBench shows where a model's abilities break down, not
just an average score - two models with the same overall accuracy can have
very different level profiles.
Motivation
Averaged benchmark scores hide useful structure: a model… See the full description on the dataset page: https://huggingface.co/datasets/luigicfilho/LadderBench.LADBench
LADBench: A Benchmark for Logical Anomaly Detection in Images
Large Vision Language Models (VLMs) excel at visual question answering and semantic grounding, but their capacity for autonomous logical reasoning remains underexplored. Existing anomaly benchmarks emphasize visual errors or direct prompting rather than the physical and social common sense needed for open-world deployment. To address this, we introduce LAD-Bench, a benchmark of more than 1,000 curated synthetic images… See the full description on the dataset page: https://huggingface.co/datasets/SahasraK/LADBench.insight-ladder-imo2024
Insight Ladder - IMO 2024 Hint-Annotated Diagnostic Substrate
Supplementary dataset for "The Insight Ladder: Quantifying the Search-Execution Gap in LLM Mathematical Reasoning" (NeurIPS 2026 Evaluations & Datasets Track, double-blind submission).
Overview
A high-density diagnostic substrate for studying search failure vs execution failure in LLM mathematical proof generation. Covers 31 IMO 2024 Shortlist problems with:
4-level hint hierarchy (L1 domain, L2 first step, L3… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-insightladder-2026/insight-ladder-imo2024.LA_dataset_blyc
Dataset Card: Learning Analytics Dataset
Overview
This dataset has been carefully curated to support fine-tuning of large language models (LLMs) with a specific focus on Learning Analytics. It is structured into three JSON files, each representing a different source or collection strategy. The dataset is particularly suited for applications in education, learning analytics, and academic research.
Dataset Description
Purpose
The dataset is designed… See the full description on the dataset page: https://huggingface.co/datasets/ibrahimBlyc/LA_dataset_blyc.eh-correctness-geometry-scale-ladder
correctness-geometry-scale-ladder -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-correctness-geometry-scale-ladder
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-correctness-geometry-scale-ladder.tt639c-rebalanced-assistant-ladder-v1
TT639C Rebalanced Assistant Ladder v1
Goal:
Repair TT639B's new simple-QA/assistant/explanation slips while preserving:
TT638D dense + dyadic proof lineage
TT639B regression repair
TT639B rule/evidence behavior
This is still a tiny exact assistant ladder, not a broad chatbot/general-knowledge claim.
Evidence:
TT639B passed regression + rule, but failed seen/heldout in simple groups such as assistant_role, tests_matter, yes/no reasoning, return_statement, and capital_france.… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt639c-rebalanced-assistant-ladder-v1.lad-multiturn-adversarial
LAD Multi-Turn Adversarial Dataset
Synthetic multi-turn conversations with three-phase turn-level labels (benign/pivoting/adversarial) for training adversarial intent detection probes on LLM activations.
Overview
Split
Conversations
Turns
Adversarial
Benign
Train
1,125
13,528
885
240
Test
797
9,142
597
200
Extended Pivoting
329
6,287
297
32
Total
2,251
28,957
1,779
472
Splits
train.json / test.json — Core Synthetic Dataset
6… See the full description on the dataset page: https://huggingface.co/datasets/pskulkarni/lad-multiturn-adversarial.tt639-tiny-assistant-ladder-v1
TT639 Tiny Assistant Ladder v1
Goal:
Prove a small assistant behavior ladder after TT638D.
This is not a broad chatbot or general knowledge claim.
It tests exact tiny behaviors:
evidence-missing / inspect-first behavior
simple Q/A
tiny explanations
four-function code answers
rewrites
classification
yes/no with reason
I don't know behavior
tiny context/multi-turn inside one prompt
Gate:
TT639 only starts after TT638D dense + dyadic/Mercy reconstruction proof.
If heldout fails:… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt639-tiny-assistant-ladder-v1.ladakhi
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: AAYAN MATEEN
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): ENGLISH,LADAKHI
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/himsec/ladakhi.Discott639f-task-routing-ladder-v1
TT639F Task Routing Ladder v1
Purpose:
Prove task routing separately after TT639E2 context-copy repair.
Task families:
classify_code / classify_prose
sentiment_positive / sentiment_negative
rewrite_short_late / rewrite_professional_attend
idk_favorite_movie / idk_food_today
Blocking gates:
seen_pass
tt639d_task_regression_pass
heldout_pass
anti_collision_pass
Do not run combined TT639G until TT639F passes.
mcqa_ladin_italian_manual
Italian-Ladin MCQA Dataset (Golden)
This is a manually created multiple-choice question answering dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_italian', 'choices_ladin', 'answer (correct choice)', 'max_choices (number of answer options)'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/mcqa_ladin_italian_manual.sentiment_analysis_ladin_italian_manual
Italian-Ladin Sentiment Analysis Dataset (Golden)
This is a manually created sentiment analysis dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'italian', 'ladin', 'label'
Labels: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
@inproceedings{nlp-ladin-2026,
title = "Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language"… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/sentiment_analysis_ladin_italian_manual.mcqa_ladin_italian
Italian-Ladin MCQA Dataset
This is a translated multiple-choice question answering dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_all_italian', 'choices_all_ladin', 'max_choices (number of answer options)', 'answer (correct choice)'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/mcqa_ladin_italian.ladydaina__ECE-FDF-details
Dataset Card for Evaluation run of ladydaina/ECE-FDF
Dataset automatically created during the evaluation run of model ladydaina/ECE-FDF
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ladydaina__ECE-FDF-details.tt639e-context-copy-ladder-v1
TT639E Context Copy Ladder v1
This rung isolates context extraction/copy behavior.
Blocking gates:
seen_pass
tt639d_seen_value_context_regression_pass
heldout_seen_value_pass
Non-blocking measured challenge:
open_copy_challenge_pass
Why:
TT639D defaulted to memorized answers such as Riley/silver.
This script balances answer values and tests unseen templates with values that appeared as answers during training.
Open-copy with entirely unseen answer values is a harder… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt639e-context-copy-ladder-v1.sentiment_analysis_ladin_italian
Italian-Ladin Sentiment Analysis Dataset
This is a translated sentiment analysis dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'italian', 'ladin', 'label'
Label: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
@inproceedings{nlp-ladin-2026,
title = "Towards the First NLP Benchmark for Ladin - an Extremely Low-Resource Language",
author = "Nuha, Ulin… See the full description on the dataset page: https://huggingface.co/datasets/ulinnuha/sentiment_analysis_ladin_italian.sentiment_analysis_ladin_italian
Italian-Ladin Sentiment Analysis Dataset
This is a translated sentiment analysis dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'italian', 'ladin', 'label'
Label: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
mcqa_ladin_italian
Italian-Ladin MCQA Dataset
This is a translated multiple-choice question answering dataset in the Ladin language (Val Badia Variant) from the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_all_italian', 'choices_all_ladin', 'max_choices', 'answer'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
lad-extended-pivoting
LAD Extended Pivoting Dataset
Synthetic multi-turn conversations with deliberately extended pivoting phases (4+ pivoting turns) for validating early adversarial detection. Companion to lad-multiturn-adversarial.
Why This Dataset?
The core LAD dataset has a mean of 3.3 pivoting turns per adversarial conversation, yielding 22-26% early detection. This dataset tests whether longer pivoting phases improve early detection — they do dramatically:
Metric
Core… See the full description on the dataset page: https://huggingface.co/datasets/pskulkarni/lad-extended-pivoting.sentiment_analysis_ladin_italian_manual
Italian-Ladin Sentiment Analysis Dataset (Golden)
This is a manually created sentiment analysis dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'italian', 'ladin', 'label'
Labels: 'pos', 'neg'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
mcqa_ladin_italian_manual
Italian-Ladin MCQA Dataset (Golden)
This is a manually created multiple-choice question answering dataset in the Ladin language (Val Badia Variant) paired with the Italian language.
Columns: 'question_italian', 'question_ladin', 'choices_italian', 'choices_ladin', 'answer', 'max_choices'
Max_choices: '3', '4', '5'
License: CC BY-NC 4.0
Citation
If this repository is helpful for your research, please cite our paper:
To be announced.
ladakmvninatec_dataset
