Pika4028/low-resource-rag-indexes
Low-Resource RAG: Wikipedia FAISS Indexes (BGE-M3) FAISS indexes over Wikipedia 2023 passages for four languages, embedded with BAAI/bge-m3. Source corpus: CohereLabs/wikipedia-2023-11-embed-multilingual-v3. Each language folder has two files aligned row-by-row: index.faiss — FAISS IndexFlatIP, dim=1024, L2-normalized BGE-M3 vectors dataset/ — Arrow dataset with _id, title, text per passage Languages Language Code Passages Size Hindi hi TBD 2.6G… See the full description on the dataset page: https://huggingface.co/datasets/Pika4028/low-resource-rag-indexes.
Add BGE-M3 indexes for Bengali, Gujarati, Urdu, Santali, Tamil, Telugu, Kannada, Malayalam, Punjabi, and Assamese
Add BGE-M3 index for as
Add BGE-M3 index for pa
Add BGE-M3 index for ml
Add BGE-M3 index for kn
Add BGE-M3 index for te
Add BGE-M3 index for ta
Add BGE-M3 index for sat
Add BGE-M3 index for ur
Add BGE-M3 index for gu
Add BGE-M3 index for bn
Rename dataset_card.md to README.md
Upload dataset_card.md with huggingface_hub
Add files using upload-large-folder tool
initial commit
