datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task1690_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1690_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1690_qed_amara_translation.qed-vie-bitextmining
qed-vie-bitextmining
Deduplicated copy of kornwtp/qed-vie-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-vie-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-vie-bitextmining.task769_qed_summarization
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task769_qed_summarization
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task769_qed_summarization.qed-zsm-bitextmining
qed-zsm-bitextmining
Deduplicated copy of kornwtp/qed-zsm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-zsm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-zsm-bitextmining.qed-ind-bitextmining
qed-ind-bitextmining
Deduplicated copy of kornwtp/qed-ind-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-ind-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-ind-bitextmining.qed-tet-bitextmining
qed-tet-bitextmining
Deduplicated copy of kornwtp/qed-tet-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tet-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tet-bitextmining.qed-tam-bitextmining
qed-tam-bitextmining
Deduplicated copy of kornwtp/qed-tam-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tam-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tam-bitextmining.qed-mya-bitextmining
qed-mya-bitextmining
Deduplicated copy of kornwtp/qed-mya-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-mya-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-mya-bitextmining.qed-tha-bitextmining
qed-tha-bitextmining
Deduplicated copy of kornwtp/qed-tha-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-tha-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-tha-bitextmining.qed-khm-bitextmining
qed-khm-bitextmining
Deduplicated copy of kornwtp/qed-khm-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-khm-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-khm-bitextmining.qed_fil_bitextmining
qed_fil_bitextmining
Deduplicated copy of kornwtp/qed_fil_bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed_fil_bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed_fil_bitextmining.qed-lao-bitextmining
qed-lao-bitextmining
Deduplicated copy of kornwtp/qed-lao-bitextmining,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/qed-lao-bitextmining
Deduplicated on: 2026-09-04
Task type: bitext_mining
Splits: train
What changed
Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every removed row was an exact duplicate after normalization; no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/qed-lao-bitextmining.safe_druglike_QED_33m
Druglike QED 36M - SAFE Dataset
This dataset is derived from Druglike molecule datasets for drug discovery. SAFE (sequential attachment-based fragment embedding) representations were generated using safe-mol (0.1.13).
Dataset Overview:
Source:
SAFE representation: safe-mol (0.1.13)
Total Entries: 33M SMILES
augmented_canonical_druglike_QED_43m
Druglike QED 43M - Augmented SMILES Dataset
This dataset is derived from the Druglike molecule datasets for drug discovery dataset and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency.
To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA's molmim method for SMILES augmentation.
Dataset Overview:
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_druglike_QED_43m.augmented_canonical_druglike_QED_Pfizer_15m
Druglike QED Pfizer 15M - Augmented SMILES Dataset
This dataset is derived from the Druglike molecule datasets for drug discovery dataset and has been canonicalized using RDKit (2024.9.4) to ensure structural consistency.
To enhance molecular diversity, 33% of the dataset was randomly sampled and augmented using RDKit’s Chem.MolToRandomSmilesVect function, following an approach similar to NVIDIA's molmim method for SMILES augmentation.
Dataset Overview:
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Derify/augmented_canonical_druglike_QED_Pfizer_15m.qedtask1689_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1689_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1689_qed_amara_translation.task1691_qed_amara_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1691_qed_amara_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1691_qed_amara_translation.qed-tet-bitextminingqedqa_dataset_with_qed_logp_mw_grpoqed-lao-bitextmininguld_loss_Llama-2-7b-chat-hf-qed
Dataset Card for "uld_loss_Llama-2-7b-chat-hf-qed"
More Information needed
safe_druglike_QED_Pfizer_11m
Drug-like QED Pfizer 11M — SAFE Dataset
This dataset is derived from the Drug-like Molecule Datasets for Drug Discovery collection. Molecular structures were converted to SAFE (Sequential Attachment-based Fragment Embedding) representations using safe-mol v0.1.14.
Source
qa_dataset_modification_qed_bin_full_jsonDCAgent2_swebench-verified-random-100-folders_DCAgent_nl2bash-nl2bash-bugsseq_Qed333699QED-en-aruld_loss_Mistral-7B-Instruct-v0.2-qed
Dataset Card for "uld_loss_Mistral-7B-Instruct-v0.2-qed"
More Information needed
flan_combined_task1691_qed_amara_translationQED_br_fr
Description
Paires breton/français du jeu de données QED disponible sur OPUS.
⚠ Attention ⚠ : il y a des problèmes d'alignement. Ce jeu de données n'est donc pas utilisbale tel quel et un post-processing serait à effectuer.
Citations
QED
@inproceedings{abdelali-etal-2014-amara,
title = "The {AMARA} Corpus: Building Parallel Language Resources for the Educational Domain",
author = "Abdelali, Ahmed and Guzman, Francisco and Sajjad, Hassan and Vogel… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/QED_br_fr.
