datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lsat_logic_games-analytical_reasoningNovel annotated evaluation dataset of LSAT logic games associated with paper:
Lost in the Logic: An Evaluation of Large Language Models’ Reasoning Capabilities on LSAT Logic Games
Arxiv: http://arxiv.org/pdf/2409.19012
If you find this dataset useful, please cite the paper!
@misc{malik2024lostlogicevaluationlarge,
title={Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games},
author={Saumya Malik},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saumyamalik/lsat_logic_games-analytical_reasoning.shipping_news_articles_lsamulti_genome_species_2kQESC
QESC — Quranic Emotional Situation Corpus
Dataset Description
QESC is the first structured corpus of Quranic prophet situational scenes annotated for emotional intensity and narrative resolution type, designed for child-facing emotional support retrieval in multilingual contexts (English, French, Moroccan Darija).
The corpus supports the CASS (Context-Aware Semantic Similarity) framework — a lightweight safety-aware retrieval system that matches a child's… See the full description on the dataset page: https://huggingface.co/datasets/lsadouk1111/QESC.PhonEx
PhonEx — French Phonics Exercise Corpus for Dyslexia
Dataset Description
PhonEx is a structured corpus of French-language phonics exercises for children aged 6–9 with reading difficulties including dyslexia, annotated for phonological difficulty level and primary error type. It is designed for constraint-aware exercise retrieval — matching a parent's or teacher's free-text description of a child's reading difficulty to the most appropriate remediation exercise.… See the full description on the dataset page: https://huggingface.co/datasets/lsadouk1111/PhonEx.5utr_af5utr_pat_ben_classmulti_genome_species_1kMyTestDSgenomic_regions_annotated
