pairs
Datasets
All datasets matching “pairs”distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO.
Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0
distilabel-intel-orca-dpo-pairs
distilabel Orca Pairs for DPO
The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved.
Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.Pexels-Pairs-Masklets-330Korca_dpo_pairs
Dataset Card for Orca DPO Pair
Dataset Description
This is a pre-processed version of the OpenOrca dataset.
The original OpenOrca dataset is a collection of augmented FLAN data that aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing preference-tuned model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Dataset Summary
The OrcaDPO Pair… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/orca_dpo_pairs.orca_dpo_pairsThe dataset contains 12k examples from Orca style dataset Open-Orca/OpenOrca.
MoCa-CL-Pairs
MoCa Contrastive Learning Data
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
This dataset contains datasets used for the supervised finetuning of MoCa (MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings):
MMEB (with hard negative)
InfoSeek (from M-BEIR)
TAT-DQA
ArxivQAVisRAG
ViDoRe
ColPali
E5 text pairs (can not release due to restrictions of Microsoft)
Image Preparation
First, you… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MoCa-CL-Pairs.
