bge-m3
Datasets
All datasets matching “bge-m3”wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.wikipedia-2024-06-bge-m3
Wikipedia Embeddings with BGE-M3
This dataset contains embeddings from the
June 2024 Wikipedia dump
for the 11 most popular languages.
The embeddings are generated with the multilingual
BGE-M3 model.
The dataset consists of Wikipedia articles split into paragraphs,
and embedded with the aforementioned model.
To enhance search quality, the paragraphs are prefixed with their
respective article titles before embedding.
Additionally, paragraphs containing fewer than 100 characters… See the full description on the dataset page: https://huggingface.co/datasets/Upstash/wikipedia-2024-06-bge-m3.bge-m3-data
Dataset Summary
This depository contains all the fine-tuning data for the bge-m3 model, including:
Dataset
Language
MS MARCO
English
NQ
English
HotpotQA
English
TriviaQA
English
SQuAD
English
COLIEE
English
PubMedQA
English
NLI from SimCSE
English
DuReader
Chinese
mMARCO-zh
Chinese
T2Ranking
Chinese
Law-GPT
Chinese
cMedQAv2
Chinese
NLI-zh
Chinese
LeCaRDv2
Chinese
Mr.TyDi
11 languages
MIRACL
16 languages
MLDR
13 languages
Note: The… See the full description on the dataset page: https://huggingface.co/datasets/Shitao/bge-m3-data.wikipedia_20240801_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240801_10-languages_bge-m3_qdrant_index.cwicr-vector-db-bgem3-v3
CWICR Vector Database — BGE-M3 V3 Snapshots
Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search.
These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.tokenlearn-c4-multilingual-bge-m3
minishlab/tokenlearn-c4-multilingual-bge-m3 Dataset Card
This dataset was created with Tokenlearn for training Model2Vec models. It contains mean token embeddings produced by a sentence transformer, used as training targets for static embedding distillation.
Dataset Details
Field
Value
Source dataset
allenai/c4
Source split
train
Embedding model
BAAI/bge-m3
Embedding dimension
1024
Rows
9999954
Dataset Structure
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/minishlab/tokenlearn-c4-multilingual-bge-m3.
