datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MRCL
MRCL: Multimodal Reasoning Continual Learning
MRCL is a five-stage benchmark for studying catastrophic forgetting during continual post-training of vision-language models. It brings together recent, challenging, and reasoning-intensive multimodal datasets spanning medical understanding, navigation and planning, geometry, visual-spatial reasoning, and financial chart analysis.
The benchmark is introduced in RL Forgets! Towards Continual Policy Optimization.
Training and… See the full description on the dataset page: https://huggingface.co/datasets/MaolinLuo/MRCL.idk-mrc
Dataset Card for IDK-MRC
Dataset Summary
I(n)dontKnow-MRC (IDK-MRC) is an Indonesian Machine Reading Comprehension dataset that covers answerable and unanswerable questions. Based on the combination of the existing answerable questions in TyDiQA, the new unanswerable question in IDK-MRC is generated using a question generation model and human-written question. Each paragraph in the dataset has a set of answerable and unanswerable questions with the corresponding answer.… See the full description on the dataset page: https://huggingface.co/datasets/rifkiaputri/idk-mrc.agri-vet-multilingual-dataset
Agri-Vet Multilingual Dataset
This repository contains JSON and JSONL resources for multilingual agricultural and veterinary language tasks. It is intended to support conversational, retrieval, classification, or instruction-tuning experiments spanning crop, animal, and related user questions.
Working with the files
Inspect each JSON/JSONL record and preserve its language, domain, prompt, response, label, and provenance fields when creating a derived dataset.… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/agri-vet-multilingual-dataset.CTF-Instructopenai_mrcr_specdecNSFW-Stories-JsonLConverted to JsonL from: bluuwhale/nsfwstory2
HumanLike-Casual-2Kklue-mrc-messages
KLUE-MRC 전체 messages 데이터
KLUE-MRC의 원본 train 17,554행과 validation 5,841행을 id와
messages(system/user/assistant)로 변환한 수정본이다. 세 문제 유형과
원본 순서·ID·지문·질문을 모두 유지했으며 필터링하거나 자르지 않았다.
문맥과 질문을 주면 짧은 정답을 답하는 독해 학습용이다. 답 있는 문항은
첫 원본 정답을 assistant target으로 사용한다. 답 없는 문항은
지문에서 답을 찾을 수 없습니다.로 변환하며 그럴듯한 오답 후보를 쓰지 않는다.
평가할 때는 원본의 모든 허용 정답을 사용해야 한다.
원본 train과 validation에 같은 지문이 일부 존재한다. 원본 split을 그대로
공개하므로 미학습 문서 일반화 평가용 분리라고 해석하지 않는다.
validation은 평가 자료이며 학습에 섞지 않는다. 별도의 로컬 연구용
search/final… See the full description on the dataset page: https://huggingface.co/datasets/YoungjaeDev/klue-mrc-messages.climate-change-MRCThe Climate Change MRC dataset, also known as CCMRC, is a part of the work "Climate Bot: A Machine Reading Comprehension System for Climate Change Question Answering", accepted at IJCAI-ECAI 2022. The paper was accepted in the special system demo track "AI for Good".
If you use the dataset, cite the following paper:
@inproceedings{rony2022climatemrc,
title={Climate Bot: A Machine Reading Comprehension System for Climate Change Question Answering.},
author={Rony, Md Rashad Al Hasan and Zuo… See the full description on the dataset page: https://huggingface.co/datasets/rony/climate-change-MRC.wave-ui-1kgpt4-instruct-similarity-0.9-dataset_yourgpt
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mrcuddle/gpt4-instruct-similarity-0.9-dataset_yourgpt.Nifty-Gay-Category
Siterip of articles from the LGBT literature digital archive nifty.org (excluding any beastiality).
Datasets is currently Work-In-Progress.
Nifty-Authoritarian-ScrapeData Scrape from LGBT Literature Archive Nifty.Org
-Category: Authoritarian
KLUE-MRC-exerciseosint-data-trainingPUGG_MRC
PUGG: KBQA, MRC, IR Dataset for Polish
Description
This repository contains the PUGG dataset designed for three NLP tasks in the Polish language:
KBQA (Knowledge Base Question Answering)
MRC (Machine Reading Comprehension)
IR (Information Retrieval)
Paper
For more detailed information, please refer to our research paper titled:
"Developing PUGG for Polish: A Modern Approach to KBQA, MRC, and IR Dataset Construction"
Authored by:
Albert Sawczyn
Katsiaryna… See the full description on the dataset page: https://huggingface.co/datasets/clarin-pl/PUGG_MRC.KGLQA-KnowledgeBant-CCLUE-MRCKGLQA-KeySentenceSelect-CCLUE-MRCSynthetic-JailBreak-RPQandAmrc_imageability_ratingsnemotron-regen-qwen3-1.7bVietnamese-Chatting-Dataset
Vietnamese Dialogue Dataset
Dữ liệu hội thoại tiếng Việt ngắn gọn, tự nhiên, dùng để fine-tune các mô hình ngôn ngữ như Mamba, LLaMA, Gemma. Nội dung bao gồm giao tiếp hàng ngày, câu hỏi thường gặp, phản hồi cảm xúc...
✅ Chuẩn định dạng JSONL
✅ Sẵn sàng dùng cho huấn luyện instruction tuning
✅ Không chứa nội dung nhạy cảm hay vi phạm
📌 Tạo bởi @hoanghai2110 để phục vụ cộng đồng mã nguồn mở AI Việt Nam.
airoboros-uncensoredDPO_Pairs_Roleplay-Alpaca
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mrcuddle/DPO_Pairs_Roleplay-Alpaca.airoboros-uncensored-conversationSD-Prompt-DPOmrcHuman-Like-AlpacaQandA_Mixtral
