CoolFace
20 results

languages

stanford-oval /wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3 This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.text-retrieval100M<n<1B0 likes3.8k downloads2y agoHugging FaceAletheia-ng /low_resource_languages_pretrain_data5text100M<n<1B0 likes1.2k downloads11mo agoHugging FaceAletheia-ng /low_resource_languages_pretrain_data2text100M<n<1B0 likes1.1k downloads1y agoHugging FaceBenji-fish /ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes1k downloads1mo agoHugging Facestanford-oval /wikipedia_20240801_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3 This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240801_10-languages_bge-m3_qdrant_index.text-retrieval100M<n<1B3 likes975 downloads2y agoHugging FaceBeardedMonster /low_resource_languages_pretrain_data8text100M<n<1B0 likes891 downloads9mo agoHugging Face