datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BertaQA
Dataset Card for BertaQA
BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.good-thinking-corpus
Good Thinking Corpus v1.0
A 27,252-record training corpus for teaching critical thinking, logic, decision theory, game theory, and related reasoning skills through interactive NPC-driven scenarios. Organized around a 182-code taxonomy spanning six tracks.
Overview
Metric
Value
Total records
27,252
Taxonomy codes
182
Tracks
6 (Logic & Critical Thinking, Decision Theory, Game Theory, Cognitive Biases, Microeconomics, Dennett's Thinking Tools)
Source types… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/good-thinking-corpus.bert-dataset
Road Traffic Act QA Dataset
This dataset is automatically generated question-answer pairs based on the official Road Traffic Act (Republic of Korea). The dataset is designed to support RAG (Retrieval-Augmented Generation) and legal NLP tasks.
Dataset Summary
Source: Road Traffic Act (English version)
Task: Question Answering (QA)
Type: Automatically generated by GPT-4o with custom multi-QA prompt
Size: 2,000+ QA pairs
Language: English
Format: CSV (Question, Answer)… See the full description on the dataset page: https://huggingface.co/datasets/YeahOuts/bert-dataset.safety-qa-bert-dataset
Safety QA Dataset
Dataset Description
There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.
