datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AraLingBench
AraLingBench
📄 Paper: arXiv:2511.14295💻 GitHub: hammoudhasan/AraLingBench
AraLingBench is a 150-question Arabic multiple-choice benchmark that tests core linguistic competence of language models across five pillars:
النحو (Grammar)
الصرف (Morphology)
الإملاء (Spelling & Orthography)
فهم اللغة (Reading Comprehension)
التركيب اللغوي والأسلوبي (Syntax & Stylistics)
All questions are human-authored and validated, with a single correct answer and a difficulty label: Easy, Medium, or… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/AraLingBench.Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1
Persian Civil Procedure QA Dataset
Dataset Description
این مجموعهداده شامل پرسشوپاسخهای حقوقی به زبان فارسی در حوزه آیین دادرسی مدنی است.
هر نمونه شامل سه فیلد اصلی است:
question: پرسش حقوقی
answer: پاسخ پرسش
evidence_quote: عبارت دقیق و مستند از دادهٔ منبع که پاسخ بر اساس آن استخراج شده است
هدف مجموعهداده، فراهمکردن دادهای ساختاریافته برای آموزش، ارزیابی و توسعه مدلهای زبانی فارسی در زمینه پرسشوپاسخ حقوقی است.
Dataset Structure
نمونهای… See the full description on the dataset page: https://huggingface.co/datasets/hamidsalimi/Persian-Civil-Procedure1-QA-Dataset-AYIN-DADRESI-MADANI-1.Marathon
Dataset Card for Marathon
Release
[2024/05/15] 🔥 Marathon is accepted by ACL 2024 Main Conference.
Dataset Summary
Marathon benchmark is a new long-context multiple-choice benchmark, mainly based on LooGLE, with some original data from LongBench. The context length can reach up to 200K+. Marathon benchmark comprises six tasks: Comprehension and Reasoning, Multiple Information Retrieval, Timeline Reorder, Computation, Passage Retrieval, and Short Dependency… See the full description on the dataset page: https://huggingface.co/datasets/Hambaobao/Marathon.multiple-choice-questions
Questões de Múltipla Escolha - Base de dados (PT-BR)
Contextualização
Este repositório contém uma base de dados (data.json) com questões de múltipla escolha, a qual foi utilizada principalmente no desenvolvimento de modelos de recuperação de informação.
Descrição do conjunto de dados
O conjunto de dados é composto por questões de múltipla escolha, abrangendo uma variedade de temas dentro da área da Ciência da Computação. Cada questão é estruturada em formato… See the full description on the dataset page: https://huggingface.co/datasets/mateus-hamade/multiple-choice-questions.fitness-qa
Fitness-QA
This is a synthetic dataset for fitness content based on "neuml/txtai-wikipedia" embedding index.
The generation of statements from context uses txtinstruct.
This dataset contains questions generated from contexts using the statement generator "flan-t5-base" trained on SQuAD dataset.
Each context includes generated questions with coherent relevant answers, and the irrelevant questions with (I don't have data on that).
Fitness data is pulled from wikipedia data stored… See the full description on the dataset page: https://huggingface.co/datasets/hammamwahab/fitness-qa.MoroccanMedMCQA-FR
MoroccanMedMCQA-FR: A French-Language Moroccan Medical Multiple-Choice QA Benchmark
Dataset Description
MoroccanMedMCQA-FR is the first French-language medical multiple-choice question answering (MCQ) benchmark grounded in the Moroccan medical faculty curriculum. It comprises 6,771 officially sourced MCQs drawn from past examinations of the Faculty of Medicine and Pharmacy of Fès (FMPF), Sidi Mohammed Ben Abdellah University, Morocco… See the full description on the dataset page: https://huggingface.co/datasets/hamzaaouadi/MoroccanMedMCQA-FR.hammurabis-code
Hammurabi's Code: A Dataset for Evaluating Harmfulness of Code-Generating LLMs
Dataset Description
This dataset, named Hammurabi's Code, is designed to evaluate the potential harmfulness of Large Language Models (LLMs) when applied to code generation and software engineering tasks. It provides prompts crafted to elicit responses related to potentially harmful scenarios, allowing for a systematic assessment of model alignment and safety. The dataset focuses on the risks… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/hammurabis-code.urdu-emergency-calls
Urdu Emergency Call Conversations Dataset (Pakistan)
Overview
This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems.
The conversations simulate real-world emergency scenarios such as:
Floods
Medical emergencies
Accidents
Crimes
Natural disasters
Public safety incidents
The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.qwen2.5-3b-civil-engineering-blind-spots
Blind Spots of Qwen2.5-3B on Civil Engineering Domain Knowledge
Model Tested
Qwen/Qwen2.5-3B — a 3-billion parameter base language model released by Alibaba Cloud in September 2024. This is the base (not instruction-tuned) variant, selected because it represents the raw pre-training knowledge without task-specific fine-tuning.
How the Model Was Loaded
The model was loaded in a Google Colab notebook using a free T4 GPU:
from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/HammamAkrami/qwen2.5-3b-civil-engineering-blind-spots.StethoBench
StethoBench
StethoBench is a comprehensive benchmark for cardiopulmonary auscultation, comprising 77,027 instruction–response pairs synthesized from 16,125 labeled recordings across 11 public datasets. It is the training and evaluation benchmark for StethoLM, published in the Transactions on Machine Learning Research (TMLR).
Dataset Description
StethoBench was constructed by synthesizing instruction–response pairs from existing labeled cardiopulmonary audio datasets… See the full description on the dataset page: https://huggingface.co/datasets/hamedfrogh/StethoBench.hamma-data
HAMMA DevOps Failure States Dataset ⬛⬜
This dataset was created to fine-tune the core intelligence engine for HAMMA — a local-first, zero-cloud SSH client built for the Gemma 4 Good Hackathon. It contains 3,701 curated problem-solution pairs covering the full surface area of Linux server administration: permissions, systemd, networking, Docker, Kubernetes, databases, storage, security, and CI/CD.
🎯 The "Zero-Fluff" Philosophy
Standard instruction-tuned LLMs respond… See the full description on the dataset page: https://huggingface.co/datasets/xayrullonematov/hamma-data.qa-wikipedia-sudan
Wikipedia-Sudan
This is a synthetic dataset based on "neuml/txtai-wikipedia" embedding index.
The generation of statements from context uses txtinstruct.
pakistan-political-leaders-chatml-dataset
🇵🇰 Pakistan Political Leaders ChatML Dataset
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Pakistani political history and leadership.
This dataset contains approximately 2500 curated question-answer pairs in ChatML format, enabling models to understand and respond to queries about major political figures in Pakistan.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a knowledgeable political… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/pakistan-political-leaders-chatml-dataset.data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/hamesh05/TruthfulQA.
