datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qedQED, is a linguistically informed, extensible framework for explanations in question answering. A QED explanation specifies the relationship between a question and answer according to formal semantic notions such as referential equality, sentencehood, and entailment. It is an expertannotated dataset of QED explanations built upon a subset of the Google Natural Questions dataset.qed_amaraThe QCRI Educational Domain Corpus (formerly QCRI AMARA Corpus) is an open multilingual collection of subtitles for educational videos and lectures collaboratively transcribed and translated over the AMARA web-based platform.
Developed by: Qatar Computing Research Institute, Arabic Language Technologies Group
The QED Corpus is made public for RESEARCH purpose only.
The corpus is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. Copyright Qatar Computing Research Institute. All rights reserved.
225 languages, 9,291 bitexts
total number of files: 271,558
total number of tokens: 371.76M
total number of sentence fragments: 30.93Mtask1690_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1690_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1690_qed_amara_translation.qed-vie-bitextmining
qed-vie-bitextmining
Deduplicated copy of kornwtp/qed-vie-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-vie-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-vie-bitextmining.task769_qed_summarization
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task769_qed_summarization
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task769_qed_summarization.qed-zsm-bitextmining
qed-zsm-bitextmining
Deduplicated copy of kornwtp/qed-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-zsm-bitextmining.qed-ind-bitextmining
qed-ind-bitextmining
Deduplicated copy of kornwtp/qed-ind-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-ind-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-ind-bitextmining.qed-tet-bitextmining
qed-tet-bitextmining
Deduplicated copy of kornwtp/qed-tet-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tet-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tet-bitextmining.qed-tam-bitextmining
qed-tam-bitextmining
Deduplicated copy of kornwtp/qed-tam-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tam-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tam-bitextmining.qed-mya-bitextmining
qed-mya-bitextmining
Deduplicated copy of kornwtp/qed-mya-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-mya-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-mya-bitextmining.qed-tha-bitextmining
qed-tha-bitextmining
Deduplicated copy of kornwtp/qed-tha-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tha-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tha-bitextmining.qed-khm-bitextmining
qed-khm-bitextmining
Deduplicated copy of kornwtp/qed-khm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-khm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-khm-bitextmining.qed_fil_bitextmining
qed_fil_bitextmining
Deduplicated copy of kornwtp/qed_fil_bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed_fil_bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed_fil_bitextmining.qed-lao-bitextmining
qed-lao-bitextmining
Deduplicated copy of kornwtp/qed-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-lao-bitextmining.safe_druglike_QED_33m
Druglike QED 36M - SAFE Dataset
This dataset is derived from Druglike molecule datasets for drug discovery. SAFE (sequential attachment-based fragment embedding) representations were generated using safe-mol (0.1.13).
Dataset Overview:
Source:
SAFE representation: safe-mol (0.1.13)
Total Entries: 33M SMILES
QE-DA-datasetsaugmented_canonical_druglike_QED_43m
Druglike QED 43M - Augmented SMILES Dataset
This dataset is derived from the Druglike molecule datasets for drug discovery dataset and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency.
To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA's molmim method for SMILES augmentation.
Dataset Overview:
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_druglike_QED_43m.augmented_canonical_druglike_QED_Pfizer_15m
Druglike QED Pfizer 15M - Augmented SMILES Dataset
This dataset is derived from the Druglike molecule datasets for drug discovery dataset and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency.
To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA's molmim method for SMILES augmentation.
Dataset Overview:
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_druglike_QED_Pfizer_15m.qe_data_yanymqedQED - The QCRI Educational Domain Corpus (formerly QCRI AMARA Corpus) is an open multilingual collection of subtitles for educational videos and lectures collaboratively transcribed and translated over the AMARA web-based platform.
It's developed by Qatar Computing Research Institute, Arabic Language Technologies Group. Along with English, it covers multiple SEA languages, such as vie (Vietnamese), mya (Burnmese), jav (Javanese), id (Indonesia), tha (Thai), tl (Tagalog),
ms (Malaysia).qedqed-ind-bitextminingref: https://opus.nlpl.eu
task1689_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1689_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1689_qed_amara_translation.qed-khm-bitextminingref: https://opus.nlpl.eu
task1691_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1691_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1691_qed_amara_translation.QEDBench
QEDBench: A Benchmark for LLM Mathematical Proof Evaluation
QEDBench is a challenging mathematical reasoning benchmark consisting of 272 proof-based problems spanning 10 mathematical domains, designed to evaluate LLMs on formal proof generation and evaluation.
🌐 Project Website
For an interactive overview of the benchmark, pass-rate visualizations, and error analysis, please see the supplementary material.
👥 Team
Principal Investigators: Quanquan C. Liu… See the full description on the dataset page: https://huggingface.co/datasets/qqggez/QEDBench.qed-tet-bitextminingqedqa_dataset_with_qed_logp_mw_grpoqed-lao-bitextmining
