datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msmarco-v2.1-snowflake-arctic-embed-m-v1.5
Snowflake Arctic Embed M V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed M v1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
It's worth noting that Snowflake's Arctic Embed M v1.5 is optimized for efficient embeddings and thus supports embedding truncation and quantization. More… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-m-v1.5.msmarco-v2.1-gte-large-en-v1.5
Alibaba GTE-Large-V1.5 Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using GTE Large V1.5 and are intended to serve as a simple baseline for dense retrieval-based methods.
Note, that the embeddings are not normalized so you will need to normalize them before usage.
Retrieval Performance
Retrieval performance for the TREC DL21-23… See the full description on the dataset page: https://huggingface.co/datasets/spacemanidol/msmarco-v2.1-gte-large-en-v1.5.LLaVA-v1.5-Instruct-620K-JA
Dataset Details
Dataset Type:Japanese LLaVA v1.5 Instruct 620K is a localized version of part of the original LLaVA v1.5 Visual Instruct 655K dataset. This version is translated into Japanese using DeepL API and is aimed at serving similar purposes in the context of Japanese language.
Resources for More Information:For information on the original dataset: LLaVA
License:Attribution-NonCommercial 4.0 International (CC BY-NC-4.0)The dataset should abide by the policy of OpenAI: OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/LLaVA-v1.5-Instruct-620K-JA.gsm8k-train-nomic-text-v1.5
Overview
Dataset containing embeddings / classification information for GSM8K
Vietnamese-BAAI-SVIT-llava-v1.5-format-gg-translated
