CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01v-bible /catholic-resources Vietnamese Catholic resources by v-bible Data Structure calendar: Generated Liturgical calendars using v-bible/js-sdk. misc/proper-names.json: Name translation from ktcgkpv.org, generated by v-bible/bible-scraper. liturgical: Liturgical data from The Lectionary for Mass (1998/2002 USA Edition), compiled by Felix Just, S.J., Ph.D., and generated by v-bible/bible-scraper. books/bible: Generated Bible markdown data. books/catechism-books: Official catechism… See the full description on the dataset page: https://huggingface.co/datasets/v-bible/catholic-resources.image10K<n<100K1 likes29k downloads18d agoHugging Face02JDRJ /kjv-bibletext10K<n<100K0 likes18k downloads2y agoHugging Face03LetsChurch /bible-embeddings Bible Embeddings A comprehensive tool for generating and evaluating Bible verse embeddings using various state-of-the-art embedding models. This project supports both commercial APIs (OpenAI, Google Gemini, Voyage AI) and open-source models (HuggingFace sentence-transformers) for semantic search across biblical texts. Setup This project is managed with uv. Make sure you have uv installed, then set up the project: # Install dependencies uv sync # Install specific… See the full description on the dataset page: https://huggingface.co/datasets/LetsChurch/bible-embeddings.5 likes9.4k downloads7mo agoHugging Face04multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.7k downloads3mo agoHugging Face05davidguzmanr /open-bible-resourcesaudio1M<n<10M0 likes2.6k downloads4mo agoHugging Face06AfriSpeech /open-bible-speech-african Open Bible Resources — African Languages Spoken-audio Bible recordings aligned to verse-level text for 19 African languages — roughly 1,741 hours of audio across ~552,907 audio–text pairs (~357 GB). This dataset is the African-language subset of davidguzmanr/open-bible-resources, re-hosted here by AfriSpeech to make the African languages easy to find and use on their own. The audio and text are unchanged from the source; only the non-African configurations have been removed. All… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/open-bible-speech-african.audioautomatic-speech-recognition100K<n<1M3 likes2.5k downloads2mo agoHugging Face07davidstap /biblenlp-corpus-mmtebThis dataset pre-computes all English-centric directions from bible-nlp/biblenlp-corpus, and as a result loading is significantly faster. Loading example: >>> from datasets import load_dataset >>> dataset = load_dataset("davidstap/biblenlp-corpus-mmteb", "eng-arb", trust_remote_code=True) >>> dataset DatasetDict({ train: Dataset({ features: ['eng', 'arb'], num_rows: 28723 }) validation: Dataset({ features: ['eng', 'arb'], num_rows: 1578 })… See the full description on the dataset page: https://huggingface.co/datasets/davidstap/biblenlp-corpus-mmteb.text1M<n<10M3 likes2.5k downloads2y agoHugging Face08mteb /biblenlp-corpus-mmteb BibleNLPBitextMining An MTEB dataset Massive Text Embedding Benchmark Partial Bible translations in 829 languages, aligned by verse. Task category t2t Domains Religious, Written Reference https://arxiv.org/abs/2304.09919 Source datasets: davidstap/biblenlp-corpus-mmteb How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_task("BibleNLPBitextMining") evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biblenlp-corpus-mmteb.texttranslation1M<n<10M2 likes1.3k downloads10mo agoHugging Face09ghanaopenai /asante-twi-bible-speech-text This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. audio10K<n<100K1 likes1.1k downloads3mo agoHugging Face10bible-nlp /sign-bibles bible-nlp/sign-bibles This dataset is still being generated and currently includes only test files This dataset contains sign language videos from the Digital Bible Library (DBL), processed for machine learning applications. The dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0). Dataset Structure Each sample contains: ["mp4"] the original video ["json"] Metadata, including bible reference, copyright information… See the full description on the dataset page: https://huggingface.co/datasets/bible-nlp/sign-bibles.textn<1K2 likes1k downloads10mo agoHugging Face11mteb /biblenlp-corpusThis dataset pre-computes all English-centric directions from bible-nlp/biblenlp-corpus, and as a result loading is significantly faster. Loading example: >>> from datasets import load_dataset >>> dataset = load_dataset("davidstap/biblenlp-corpus-mmteb", "eng-arb", trust_remote_code=True) >>> dataset DatasetDict({ train: Dataset({ features: ['eng', 'arb'], num_rows: 28723 }) validation: Dataset({ features: ['eng', 'arb'], num_rows: 1578 })… See the full description on the dataset page: https://huggingface.co/datasets/mteb/biblenlp-corpus.text1M<n<10M1 likes1k downloads7mo agoHugging Face12flagship-ai /cameroon_bibles Cameroon Bibles — verse-aligned scripture corpus The text corpus behind Lingo / NativeAI: verse-aligned scripture across 60 Cameroonian languages (64 translation versions). Scripture is one of the few sources of sentence-aligned parallel text for these low-resource languages — the aligned backbone of our corpus (see the research log). Layout <Language>/<BOOK>.<chapter>.txt e.g. Ngi/MAT.2.txt Each file is one chapter; lines are verse-numbered, alignable across… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/cameroon_bibles.texttranslation100K<n<1M0 likes883 downloads4mo agoHugging Face13ghanaopenai /ewe-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes792 downloads3mo agoHugging Face14roycvn /suvartha-bible-audio-irv Suvartha Bible audio — CC BY-SA 4.0 Chapter-by-chapter readings of the Bible, one MP3 per chapter, as played in the Bible reader at https://suvartha.in. Folder Language Credit Source Licence hi-irv Hindi Indian Revised Version — text and audio © Bridge Connectivity Solutions / Davar Partners International, CC BY-SA 4.0 (via Snow Mountain dataset) source CC BY-SA 4.0 kn-irv Kannada Indian Revised Version — text and audio © Bridge Connectivity Solutions / Davar… See the full description on the dataset page: https://huggingface.co/datasets/roycvn/suvartha-bible-audio-irv.audio1K<n<10K0 likes725 downloads8d agoHugging Face15Emmet-Allen /The-Bible-KJVtabular10K<n<100K0 likes696 downloads1y agoHugging Face16bible-nlp /biblenlp-corpus Dataset Card for BibleNLP Corpus Dataset Summary Partial and complete Bible translations in 833 languages, aligned by verse. Languages aai, aak, aau, aaz, abt, abx, aby, acf, acr, acu, adz, aer, aey, agd, agg, agm, agn, agr, agt, agu, aia, aii, aka, ake, alp, alq, als, aly, ame, amf, amk, amm, amn, amo, amp, amr, amu, amx, anh, anv, aoi, aoj, aom, aon, apb, ape, apn, apr, apu, apw, apz, arb, are, arl, arn, arp, asm, aso, ata, atb, atd, atg, att, auc, aui, auy… See the full description on the dataset page: https://huggingface.co/datasets/bible-nlp/biblenlp-corpus.translation1M<n<10M42 likes551 downloads2y agoHugging Face17k-mktr /bible_king_james_version_en King James Version (1611) Description The most influential English Bible translation in history, commissioned by King James I of England and first published in 1611. The translation was prepared by 47 scholars organized into six committees, working from the original Hebrew, Aramaic, and Greek texts, as well as consulting earlier English translations (Tyndale, Coverdale, Geneva Bible) and the Latin Vulgate. The KJV is renowned for the majesty of its prose and its… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/bible_king_james_version_en.tabular10K<n<100K0 likes551 downloads2mo agoHugging Face18vpetukhov /bible_tts_hausa Dataset Card for BibleTTS Hausa Dataset Summary BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings. This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603. Languages Hausa Dataset Structure Data Fields audio: audio path sentence: transcription of the audio locale: always set to ha book: 3-char book encoding verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.textautomatic-speech-recognition10K<n<100K7 likes532 downloads4y agoHugging Face19ghanaopenai /dagbani-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes528 downloads3mo agoHugging Face20Flux9665 /BibleMMSThe Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina Meyer, Lyonel Behringer, Frank Zalkow, Phat Do, Matt Coler, Emanuël A. P. Habets and Ngoc Thang Vu (Interspeech 2024). We generate 2000 spoken utterances per language using the subsets of the eBible dataset [1] that are under free licenses as the text input to the MMS TTS models [2]. The languages associated with the following ISO-639-3 codes are represented in this… See the full description on the dataset page: https://huggingface.co/datasets/Flux9665/BibleMMS.audiotext-to-speech100K<n<1M82 likes522 downloads2y agoHugging Face21quintic /bible-text-fixing-v00 likes497 downloads3y agoHugging Face22hmar-heritage-org /zo-biblegated zo-bible A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng) Family: Zo Languages Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.tabulartranslation10K<n<100K4 likes495 downloads8d agoHugging Face23StephenZao /bible-sphere-statstabularn<1K3 likes484 downloads8h agoHugging Face24nordpolemil /biblenlp-corpus BibleNLP Corpus This is a conversion of BibleNLP corpus to the Parquet format, Dataset Summary The dataset contains partial and complete Bible translations in 835 languages, aligned by verse. Each language is stored as a separate Parquet file (eng.parquet, fra.parquet, …). This format is derived from the eBible corpus corpus.json and is intended for fast columnar loading with Hugging Face datasets, Polars, Pandas, or DuckDB. Languages 835 ISO 639-3… See the full description on the dataset page: https://huggingface.co/datasets/nordpolemil/biblenlp-corpus.tabulartranslation10M<n<100M0 likes360 downloads3mo agoHugging Face25Svngoku /kikongo-bible-asr Kikongo Bible ASR This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/kikongo-bible-asr.audioautomatic-speech-recognition1K<n<10K3 likes347 downloads2y agoHugging Face26Helsinki-NLP /bible_paraThis is a multilingual parallel corpus created from translations of the Bible compiled by Christos Christodoulopoulos and Mark Steedman. 102 languages, 5,148 bitexts total number of files: 107 total number of tokens: 56.43M total number of sentence fragments: 2.84Mtranslation10K<n<100K22 likes268 downloads3y agoHugging Face27meowmeowcatskill /bible-new-testamentaudion<1K0 likes251 downloads2y agoHugging Face28manassehzw /shona-bible-bdsc-aligned Shona Bible Speech Alignment Dataset Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC source audio made available by Biblica, Inc. through Open.Bible. This release contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech segments covering approximately 75.55 hours. Dataset summary Language: Shona (sna) Speaker: narrator 1 Speaker sex: male Books: 66 Clips: 31,284 Audio: approximately 75.55 hours Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.audioautomatic-speech-recognition10K<n<100K3 likes237 downloads14d agoHugging Face29versae /biblesMultilingual Biblestext1M<n<10M4 likes233 downloads4y agoHugging Face30sanjeevafk /biblelm BibleLM Dataset A high-performance, stateless Bible dataset optimized for edge-first RAG (Retrieval-Augmented Generation). This dataset contains the processed Bible text, morphological data, and search indices used by the BibleLM project. 📚 What's inside? Combined Bible Index: Cleaned and tokenized text for BSB (Berean Standard Bible), KJV, WEB, and ASV. Search Engine State: Pre-computed BM25 term frequencies (bm25-state.json) allowing for <10ms search engine… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/biblelm.tabularquestion-answering10K<n<100K0 likes233 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.