CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01naklecha /minecraft-question-answer-700k minecraft-question-answer-700k Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline. about the dataset rows - 694,814 tokens - 47,133,624 source - https://minecraft.wiki/ Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.textquestion-answering100K<n<1M46 likes116 downloads2y agoHugging Face02BoltMonkey /psychology-question-answerA JSON formatted dataset comprising 197,180 question and answer pairs covering a wide range of topics encountered in a Bachelor level psychology course. I have included a broad range of question types, topics, and answer styles. The dataset was created using personal notes and several LLMs (such as GPT4) and manually assessed for veracity and completeness of response. Despite this, the size of the dataset prohibits me from ensuring every single answer is 100% accurate and up-to-date. As such… See the full description on the dataset page: https://huggingface.co/datasets/BoltMonkey/psychology-question-answer.textquestion-answering100K<n<1M11 likes62 downloads2y agoHugging Face03Aixr /Math-Question-Answertexttext-generation1K<n<10K3 likes44 downloads2y agoHugging Face04LLaMAX /BenchMAX_Question_Answering Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios. The subtasks are similar to the subtasks in RULER. The data is sourcing from UN Parallel Corpus and xquad. The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.texttext-generationn<1K0 likes38 downloads2y agoHugging Face05kurumikz /Question-Answering_Kazakh 🇰🇿 Question-Answering_Kazakh A comprehensive Kazakh-language question-answer dataset for fine-tuning and training language models.Created and maintained by Kurumikz. Free to use with attribution. 📌 Overview Question-Answering_Kazakh is an open-domain QA dataset written entirely in the Kazakh language (kk). It covers a wide range of topics — from the history and geography of Kazakhstan to Kazakh grammar, culture, economy, and language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.textquestion-answering1K<n<10K1 likes36 downloads6mo agoHugging Face06kurumikz /Question-answeringsmall-ru Dataset Card for Question Answering Russian Dataset 🧠 Quick Summary Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях. 📚 Dataset Details Curated by: @kurumikz Language(s): Russian (ru) License: CC-BY 4.0 — свободное использование с обязательным указанием автора Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.textquestion-answering1K<n<10K1 likes31 downloads11mo agoHugging Face07Bytte-AI /Pidgin_Question-English_Answer_Dataset Pidgin Question - English Answer Dataset (Sample) Data Card v1.0 Dataset Name: Pidgin Question - English Answer Dataset (Sample)Dataset Type: Sample DatasetVersion: 1.0Release Date: 2026Organization: Bytte AILicense: CC-BY-4.0Contact: contact@bytteai.xyzWebsite: https://www.bytte.xyz/ Note: This is a sample dataset containing 331 cross-lingual question-answer pairs (Pidgin questions → English answers). Generated through AI chatbot interactions with human validation… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin_Question-English_Answer_Dataset.texttext-classificationn<1K0 likes26 downloads7mo agoHugging Face08kimleang123 /khmer_question_answergatedThe data collected from https://www.khsearch.com/ related to the general question-answering examination. It used to train fine-tuned models from many LLMs, including LlaMa, Qwen, Mistral, and Gemma. Under the research title "Fine-tuning for Question Answering in Low-Resource Languages: A Case Study on Khmer" conducted at ViLa Lab, Institute of Technology of Cambodia, Phnom Penh. Lab Info: https://www.facebook.com/vilalabitc Paper:… See the full description on the dataset page: https://huggingface.co/datasets/kimleang123/khmer_question_answer.textquestion-answering10K<n<100K3 likes10 downloads1y agoHugging Face09RomainPct /steve-jobs-question-and-answerstexttext-generationn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.