datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fruit-fidelity-quant-siq-v1
GLM-5.2-SIQ-Fruit — quantized fidelity dataset (hidden form, reconstructed weights)
The candidate half of a three-step fidelity measurement on
malaiwah/GLM-5.2-SIQ-Fruit
@ c1798e3676fa16b4a874381171adab1e3033fbd5, captured on the same panel, the
same lane and the same engine as its root,
malaiwah/fruit-fidelity-root-v1.
step 1 capture reference weights + panel -> fruit-fidelity-root-v1
step 2 capture quantized weights + panel -> THIS REPOSITORY
step 3 compare root… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/fruit-fidelity-quant-siq-v1.SIQA_es
Dataset Card for SIQA (Spanish Version)
Dataset summary
This dataset provides the Spanish translation and adaptation of the SIQA (Social Interaction Question Answering) validation set. The original dataset was designed to evaluate social commonsense reasoning in LLMs by presenting a collection of questions based on everyday social situations, with the ultimate goal of challenging models to infer motivations, reactions and social implications behind human actions.
This… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/SIQA_es.siqa
SIQA — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror. OpenCompass-native eval data for SIQA, laid out at siqa/ exactly as the OpenCompass dataset loader expects. Sourced from the OpenCompass data distribution; license CC-BY-4.0, unchanged; all rights remain with the original authors.
siqi00__Mistral-7B-DFT2-details
Dataset Card for Evaluation run of siqi00/Mistral-7B-DFT2
Dataset automatically created during the evaluation run of model siqi00/Mistral-7B-DFT2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/siqi00__Mistral-7B-DFT2-details.siqa-mt-pt
SIQA-PT
Portuguese machine translation of Social IQA, a benchmark for social commonsense reasoning about everyday situations.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/social_i_qa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/siqa-mt-pt.siqasiqi00__Mistral-7B-DFT-details
Dataset Card for Evaluation run of siqi00/Mistral-7B-DFT
Dataset automatically created during the evaluation run of model siqi00/Mistral-7B-DFT
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/siqi00__Mistral-7B-DFT-details.
