datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MEDI2
Dataset Card for MEDI2
Citation
@misc{muennighoff2024generative,
title={Generative Representational Instruction Tuning},
author={Niklas Muennighoff and Hongjin Su and Liang Wang and Nan Yang and Furu Wei and Tao Yu and Amanpreet Singh and Douwe Kiela},
year={2024},
eprint={2402.09906},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
MEDI2BGEGRIT
GRIT: Large-Scale Training Corpus of Grounded Image-Text Pairs
Dataset Summary
We introduce GRIT, a large-scale dataset of Grounded Image-Text pairs, which is created based on image-text pairs from COYO-700M and LAION-2B. We construct a pipeline to extract and link text spans (i.e., noun phrases, and referring expressions) in the caption to their corresponding image regions. More details can be found in the paper.
Supported Tasks
During the construction, we… See the full description on the dataset page: https://huggingface.co/datasets/zzliang/GRIT.grit_2mMEDI
Dataset Card for MEDI
Citation
@misc{muennighoff2024generative,
title={Generative Representational Instruction Tuning},
author={Niklas Muennighoff and Hongjin Su and Liang Wang and Nan Yang and Furu Wei and Tao Yu and Amanpreet Singh and Douwe Kiela},
year={2024},
eprint={2402.09906},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
wikitext-2-raw-v1-preprocessedopen-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokensopen-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokens-outputs-v1AOPS_Full_Verified_gritlm_7btuluv2-expanded-150k-samples-distill-insert-ret-tokenstuluv2-expanded-full-distill-insert-ret-tokensGRIT_data
GRIT Dataset
This repository contains data used for training and evaluating the GRIT (Grounded Reasoning with Images and Texts) model as presented in our GRIT project.
Dataset Structure
GRIT_data/
├── gqa/
│ └── images/ (not included, see download instructions)
├── ovd_position/
│ └── images/ (not included, see download instructions)
├── vsr/
│ └── images/ (not included, see download instructions)
├── mathvista_mini/
│ └── images/ (not included, see download… See the full description on the dataset page: https://huggingface.co/datasets/yfan1997/GRIT_data.open-hermes-2.5-sft-active-retrieval-sample-300k-v1tuluv2-expanded-fulltulu-v2-sft-mixture-first-stage-classifierasqa_evaltulu-v2-sft-mixture-first-stage-classifier-v1grit-webdAOPS_Full_Verified_gritlm_7b_5678910AOPS_Full_Verified_gritlm_7b_1234tulu2tuluv2-expanded-150k-samplestuluv2-expanded-150k-part_v0-chat-format-syn-knowledgetulu-v2-sft-mixture-first-stage-classifier-outputstulu-v2-sft-short-instructPILE_Wikipedia_Pretraining_subset_100k-distill-insert-ret-tokenstulu-v2-sft-mixture-active-retrieval-v1open-hermes-2.5-sft-active-retrieval-v1PILE_Wikipedia_Pretraining_subset_100k-distill-insert-ret-tokens-outputsGritLM__GritLM-7B-KTO-details
Dataset Card for Evaluation run of GritLM/GritLM-7B-KTO
Dataset automatically created during the evaluation run of model GritLM/GritLM-7B-KTO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/GritLM__GritLM-7B-KTO-details.
