QED
Datasets
All datasets matching “QED”qedQED, is a linguistically informed, extensible framework for explanations in question answering. A QED explanation specifies the relationship between a question and answer according to formal semantic notions such as referential equality, sentencehood, and entailment. It is an expertannotated dataset of QED explanations built upon a subset of the Google Natural Questions dataset.qed_amaraThe QCRI Educational Domain Corpus (formerly QCRI AMARA Corpus) is an open multilingual collection of subtitles for educational videos and lectures collaboratively transcribed and translated over the AMARA web-based platform.
Developed by: Qatar Computing Research Institute, Arabic Language Technologies Group
The QED Corpus is made public for RESEARCH purpose only.
The corpus is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. Copyright Qatar Computing Research Institute. All rights reserved.
225 languages, 9,291 bitexts
total number of files: 271,558
total number of tokens: 371.76M
total number of sentence fragments: 30.93Mtask1690_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1690_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1690_qed_amara_translation.qed-vie-bitextmining
qed-vie-bitextmining
Deduplicated copy of kornwtp/qed-vie-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-vie-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-vie-bitextmining.task769_qed_summarization
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task769_qed_summarization
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task769_qed_summarization.qed-tet-bitextmining
qed-tet-bitextmining
Deduplicated copy of kornwtp/qed-tet-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tet-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tet-bitextmining.
