datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
granite-decisions-synthetic
Granite Decisions synthetic datasets
Original, deterministic English fixtures for Adam Pippert's personal
Granite Decisions project.
The original default config has 162 examples: 54 train, 54 calibration, and 54 test.
These exercise the pipeline; they are not a representative quality benchmark.
Source and license
The source is the project's original template generator, published here as
make_smoke_data.py, from
release v0.1.0,
commit… See the full description on the dataset page: https://huggingface.co/datasets/adampippert/granite-decisions-synthetic.traces.claude-code.mlx-lm-granitemoehybridibm-granite__granite-3.0-2b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-2b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-2b-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-2b-base-details.for-train-granite-4.0datasets:
5CD-AI/Vietnamese-Multi-turn-Chat-Alpaca
HuggingFaceH4/no_robots
teknium/OpenHermes-2.5
open-thoughts/OpenThoughts-114k
ty
ibm-granite__granite-3.2-8b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.2-8b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.2-8b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.2-8b-instruct-details.ibm-granite__granite-3.0-2b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-2b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-2b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-2b-instruct-details.ibm-granite__granite-3.0-1b-a400m-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-1b-a400m-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-1b-a400m-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-1b-a400m-instruct-details.ibm-granite__granite-3.0-1b-a400m-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-1b-a400m-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-1b-a400m-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-1b-a400m-base-details.ibm-granite__granite-3.1-3b-a800m-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-3b-a800m-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-3b-a800m-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-3b-a800m-base-details.ibm-granite__granite-3.0-8b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-8b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-8b-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-8b-base-details.ibm-granite__granite-3.0-3b-a800m-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-3b-a800m-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-3b-a800m-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-3b-a800m-base-details.Granite-v4.1-Distilled-15K
⛰️ Granite-v4.1-Distilled-15K
Dataset Summary
Granite-v4.1-Distilled-15k is a supervised fine-tuning dataset for logic-oriented distillation. The prompts to the questions come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated using the only granite-4.1-8b teaching model.
Dataset Details
Dataset
constructai/Granite-v4.1-Distilled-15K
Source questions
Jackrong/GLM-5.1-Reasoning-1M-Cleaned
Teacher model
Granite-4.1-8b… See the full description on the dataset page: https://huggingface.co/datasets/constructai/Granite-v4.1-Distilled-15K.ibm-granite__granite-3.1-8b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-8b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-8b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-8b-instruct-details.ibm-granite__granite-7b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-7b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-7b-base
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-7b-base-details.ibm-granite__granite-3.2-2b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.2-2b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.2-2b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.2-2b-instruct-details.ibm-granite__granite-3.0-8b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-8b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-8b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-8b-instruct-details.ibm-granite__granite-7b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-7b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-7b-instruct-details.ibm-granite__granite-3.0-3b-a800m-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.0-3b-a800m-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-3b-a800m-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-3b-a800m-instruct-details.ibm-granite__granite-3.1-2b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-2b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-2b-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-2b-base-details.ibm-granite__granite-3.1-8b-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-8b-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-8b-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-8b-base-details.ibm-granite__granite-3.1-3b-a800m-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-3b-a800m-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-3b-a800m-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-3b-a800m-instruct-details.ibm-granite__granite-3.1-2b-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-2b-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-2b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-2b-instruct-details.ibm-granite__granite-3.1-1b-a400m-base-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-1b-a400m-base
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-1b-a400m-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-1b-a400m-base-details.ibm-granite__granite-3.1-1b-a400m-instruct-details
Dataset Card for Evaluation run of ibm-granite/granite-3.1-1b-a400m-instruct
Dataset automatically created during the evaluation run of model ibm-granite/granite-3.1-1b-a400m-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.1-1b-a400m-instruct-details.m2v_c4-featurized-granite-125mrag-docs-granitegranite_hotpotqa_grounded_resultgranite-asciiblindspots-frontier-models-granite-4-0-1b-base
Blind Spots of Frontier Models (IBM Granite 4.0 1B Base)
Model tested: ibm-granite/granite-4.0-1b-baseModel card: https://huggingface.co/ibm-granite/granite-4.0-1b-base
For inference, I ran this model locally, though I also experimented with free models from OpenRouter.
This dataset contains 10 evaluation rows with:
input
expected_output
model_output
notes
is_correct
I loaded the model with transformers and evaluated it using strict concise-answer prompts.
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/Tomodovodoo/blindspots-frontier-models-granite-4-0-1b-base.granite3.3_quality_baseline_result
