datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
siqa
Dataset Card for "siqa"
More Information needed
siqaTrainSet
SIQA TrainSet
A standardized TrainSet for the Scientific Image Quality Assessment (SIQA), designed to train multimodal models on two core tasks SIQA-U & SIQA-S.
NOTICE:Due to Hugging Face Datasets' automatic caching mechanism, the local disk usage can be up to ~20× larger than the actual dataset size.
You only need to download the 'images/' folder to run inference or evaluation, please ignore the .parquet in 'SIQA-S/' and 'SIQA-U'
To Quick Start, you download by git:
git clone… See the full description on the dataset page: https://huggingface.co/datasets/SIQA/TrainSet.llama2_7b_chat-siqa-resultssiqa-multilingualqwen2.5_openthoughts_0.7_1.0_-1_8192SIQA
SIQA
An MTEB dataset
Massive Text Embedding Benchmark
Measuring the ability to retrieve the groundtruth answers to reasoning task queries on SIQA.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://leaderboard.allenai.org/socialiqa/submissions/get-started
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("SIQA")
evaluator = mteb.MTEB([task])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/SIQA.siqa_ca
Dataset Card for siqa_ca
siqa_ca is a multiple choice question answering dataset in Catalan that has been professionally translated from the SIQA
validation set in English.
Dataset Details
Dataset Description
siqa_ca (Social Interaction Question Answering - Catalan) is designed to evaluate social commonsense intelligence using multiple choice question-answer instances based on reasoning about people’s actions and their
social implications. It includes… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/siqa_ca.siqaalignment-honesty-multisample_p1-7b-1kfruit-fidelity-quant-siq-v1
GLM-5.2-SIQ-Fruit — quantized fidelity dataset (hidden form, reconstructed weights)
The candidate half of a three-step fidelity measurement on
malaiwah/GLM-5.2-SIQ-Fruit
@ c1798e3676fa16b4a874381171adab1e3033fbd5, captured on the same panel, the
same lane and the same engine as its root,
malaiwah/fruit-fidelity-root-v1.
step 1 capture reference weights + panel -> fruit-fidelity-root-v1
step 2 capture quantized weights + panel -> THIS REPOSITORY
step 3 compare root… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/fruit-fidelity-quant-siq-v1.qwen_ultrafeedbackSIQA_es
Dataset Card for SIQA (Spanish Version)
Dataset summary
This dataset provides the Spanish translation and adaptation of the SIQA (Social Interaction Question Answering) validation set. The original dataset was designed to evaluate social commonsense reasoning in LLMs by presenting a collection of questions based on everyday social situations, with the ultimate goal of challenging models to infer motivations, reactions and social implications behind human actions.
This… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/SIQA_es.tulu_siqasiqa
SIQA — evaluation data (OpenCompass format)
Bud Ecosystem eval mirror. OpenCompass-native eval data for SIQA, laid out at siqa/ exactly as the OpenCompass dataset loader expects. Sourced from the OpenCompass data distribution; license CC-BY-4.0, unchanged; all rights remain with the original authors.
siqi00__Mistral-7B-DFT2-details
Dataset Card for Evaluation run of siqi00/Mistral-7B-DFT2
Dataset automatically created during the evaluation run of model siqi00/Mistral-7B-DFT2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/siqi00__Mistral-7B-DFT2-details.mistral_metamath_question_0.7_1.0_50_256gemma_ultrafeedbacksiqa-mt-pt
SIQA-PT
Portuguese machine translation of Social IQA, a benchmark for social commonsense reasoning about everyday situations.
Translated using a Finetuned GemmaX2-9B for pt-PT.
Original Dataset: https://huggingface.co/datasets/social_i_qa
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/siqa-mt-pt.llama2_7b_chat-siqaqwen_openr1mathsiqa_ca_old
Dataset Card: SIQA_CA (Pre-revision version)
Description
SIQA_CA (Pre-revision) is an earlier Catalan translation of the Social IQa (SIQA) dataset, a benchmark designed to evaluate commonsense reasoning about social interactions. This version consists of manually translated instances from the original English dataset into Catalan. It is used as a baseline for comparison against a revised and improved version of the dataset (SIQA_CA v2).
Motivation and Use Case… See the full description on the dataset page: https://huggingface.co/datasets/langtech-languagemodeling/siqa_ca_old.siqa-ko
Dataset Card for "siqa-ko"
More Information needed
ultrafeedback_binarizedalignment-honesty-absolute_p1-7b-fullsiqatestchatsiqamistral_ultrafeedback_unhelpful_chatprompt_0.7_1.0_50_320This dataset is used in the paper Discriminative Finetuning of Generative Large Language Models without Reward Models and Preference Data. It contains paired examples of "real" conversation turns and multiple generated alternatives. The goal is to train a model to discriminate between high-quality and low-quality generations.
The dataset is structured as follows: Each example contains a "real" conversation turn and 16 generated alternatives. Each turn is represented as a list of… See the full description on the dataset page: https://huggingface.co/datasets/siqi00/mistral_ultrafeedback_unhelpful_chatprompt_0.7_1.0_50_320.qwen2.5_7b_ultrafeedback_0.7_1.0_50_320
