datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sharegpt-quizz-generation-json-output
ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.supervised-finetuning_quiz_student_responsesnrk_quiz_qa
Dataset Card for NRK-Quiz-QA
Dataset Details
Dataset Description
NRK-Quiz-QA is a multiple-choice question answering (QA) dataset designed for zero-shot evaluation of language models' Norwegian-specific and world knowledge. It comprises 4.9k examples from over 500 quizzes on Norwegian language and culture, spanning both written standards of Norwegian: Bokmål and Nynorsk (the minority variant). These quizzes are sourced from NRK, the national public broadcaster… See the full description on the dataset page: https://huggingface.co/datasets/ltg/nrk_quiz_qa.music-quiz-scoresBEAR
BEAR Dataset
Dataset to evaluate common factual knowledge in language models. This dataset was created as part of the paper "BEAR: A Unified Framework for Evaluating Relational Knowledge in Causal and Masked Language Models".
For more information visit the LM Pub Quiz website.
Dataset Structure
Data Instances
For each instance, there is a template string, a subject to be filled in, and a list of object strings (one of which is correct), as well as a template… See the full description on the dataset page: https://huggingface.co/datasets/lm-pub-quiz/BEAR.HarryPotter-Quiz
Dataset Card for "HarryPotter-Quiz"
Harry Potter Quiz (or HP-Quiz in short) is a manually collected dataset based on information about characters and magic spells from the Harry Potter Wiki pages. For characters, we gather attributes such as gender, hair color, house, and relationships with other characters. For magic spells, we collect data on the types of spells they belong to. We then create multiple-choice questions and answers based on this information.
The dataset includes 300… See the full description on the dataset page: https://huggingface.co/datasets/cross-ling-know/HarryPotter-Quiz.supervised-finetuning_quiz_student_responsesmy-quiz-evalQuizQuestionBankquiz-no-mori
mahiyama/quiz-no-mori
クイズの森 (https://quiz-schedule.info/quiz_no_mori/) のクイズ Q&A を日本語版 Wikipedia パッセージで根拠付けした
Open-domain QA データセットに、本リポ内で Hard Negative Mining と Cross-Encoder
蒸留スコア付与を実施した日本語 retrieval / 蒸留学習用データセット。
本リポに供給される (query, positive) ペアは、クイズの森 のクイズ問題文を query、
その短い正答 (例: 「青年海外協力隊」) を 日本語版 Wikipedia 全 passage コレクションへ
引き当てて取得した 根拠 passage 本文 を positive としたもの。短答テキストから passage
本文への引き当て (hydration) は別パイプラインで済んだ状態で入力されており、本リポでは
それを起点に Hard Negative Mining と Cross-Encoder… See the full description on the dataset page: https://huggingface.co/datasets/mahiyama/quiz-no-mori.quiz-works
mahiyama/quiz-works
クイズワークス (https://quiz-works.com/) のクイズ Q&A を日本語版 Wikipedia パッセージで根拠付けした
Open-domain QA データセットに、本リポ内で Hard Negative Mining と Cross-Encoder
蒸留スコア付与を実施した日本語 retrieval / 蒸留学習用データセット。
本リポに供給される (query, positive) ペアは、クイズワークス のクイズ問題文を query、
その短い正答 (例: 「青年海外協力隊」) を 日本語版 Wikipedia 全 passage コレクションへ
引き当てて取得した 根拠 passage 本文 を positive としたもの。短答テキストから passage
本文への引き当て (hydration) は別パイプラインで済んだ状態で入力されており、本リポでは
それを起点に Hard Negative Mining と Cross-Encoder 蒸留スコア付け、quality_score… See the full description on the dataset page: https://huggingface.co/datasets/mahiyama/quiz-works.quiz_olympus_dataset
Quiz Olympus Questions Dataset
This is a public dataset of questions from the game Quiz Olympus (https://www.quizolympus.com). It is intended for research, language model development, trivia games, demos, and educational applications.
Format
Available in JSON as quiz_olympus_questions.json.
Can also be exported to CSV for analysis or integration into other tools.
Example JSON entry:
{
"question": "Texto de la pregunta",
"uuid": "uuid-de-la-pregunta"… See the full description on the dataset page: https://huggingface.co/datasets/Mariano72/quiz_olympus_dataset.quizgen_langfusebrightai-quiz-v0.2-datasetquiz_militareQUIZBOT_AI_DATASET
Dataset Card for "QUIZBOT_AI_DATASET"
More Information needed
quiz-no-moriクイズの杜様に掲載のクイズのうち、2024年8月5日時点において取得可能だったクイズのうち「二次利用許諾レベル」が「フリー」であったものを収載したデータセットです。
検索拡張生成(RAG)や文書検索システムの構築などへの利用に適した高品質なデータです。
クイズの杜様ホームページ上の問題の二次利用についてに記載されている「二次利用許諾レベル」に関する説明のうち
フリー改変した問題、前フリをつけた問題、インスピレーションを受けて作成した類題等を、商用目的で利用する(例・クイズ番組での出題、利益を上げているオープン大会での出題、販売)コピー行い、不特定多数に配布したり、www上に公開したりする。フォーマットを変換し、不特定多数に公開する。他、全ての二次利用は自由。
との記述に基づき、本データセットも同様、自由に二次利用可能なデータセットとして公開しております。
ただし、クイズの杜様およびその他関係者の皆様へ迷惑のかかる利用方法はお断りさせていただきますことをご了承ください。
Contact… See the full description on the dataset page: https://huggingface.co/datasets/hpprc/quiz-no-mori.scraped_quizzesquizgen-chat-mdgenerate-quiz-datasetvikwiki-quiz
VIK Wiki quiz
Leírás
VIK wikiről scrapelt kikérdező parsolva, .jsonl formátumban.
A scrapelést/parsolást végző kód a itt a repo-ban megtalálható.
Cél
Én elsősorban LLM-ek tanításra/kiértékelésre gondoltam.
Adatforrás
Minden VIK wiki kikérdező.
Mezők
title: a két == közötti szövegrész. Szinte mindig meg van adva. Általában ez tartalmazza magát a kérdést.
question: a title és a {{Kvízkérdés...}} közti rész. Általában üres. Több kontextust… See the full description on the dataset page: https://huggingface.co/datasets/boapps/vikwiki-quiz.24problems_quiz24problems_quiz-eval-n4-1-10-24pdf_to_quizz_mistralROI-Math-QuizAlignedquiz-worksQuiz Works様に掲載のクイズのうち、2024年8月4日~8月5日時点において取得可能だったクイズを収載したデータセットです。
検索拡張生成(RAG)や文書検索システムの構築などへの利用に適した高品質なデータです。
Quiz Works様ホームページ上のこのサイトについてに記載されている
このサイトに掲載されるクイズはすべて自由に二次利用可能です。
との記述に基づき、本データセットも同様、自由に二次利用可能なデータセットとして公開しております。
ただし、Quiz Works様およびその他関係者の皆様へ迷惑のかかる利用方法はお断りさせていただきますことをご了承ください。
Contact
本ページ上に公開されているデータについて何らかの問題・お気づきの点があれば本ページ作成者までお問い合わせくださいませ。
Acknowledgement
本データの掲載元であるQuiz Works様に感謝申し上げます。
my-quiz-evalgpt5-attention-is-all-you-need-quiz
Dataset Card for "Attention is all you need" quiz
GPT5-Thinking generated synthetic dataset of 15 questions, choices and answers.
Prompt used for generation
https://arxiv.org/pdf/1706.03762 Create a quiz of 15 questions with answers in the form of a hf dataset.
Dataset Card Contact
IDK
scraped_quizzesru-alpaca-quiz
Это переработка в alpaca-friendly формат датасетов от:
MERA-evaluation[MERA]
Из датасета взяты и переработаны только subsets (chegea, multiq)
