CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes295k downloads1y agoHugging Face02ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes135k downloads1y agoHugging Face03ieasybooks-org /shamela-waqfeya-library Shamela Waqfeya Library 📖 Overview Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.tabularimage-to-text1K<n<10K4 likes92k downloads1y agoHugging Face04prasadonly /webtoepub-library8 likes29k downloads3d agoHugging Face05thuml /Time-Series-Library Time-Series-Library (TSLib) TSLib is an open-source library for deep learning researchers, especially for deep time series analysis. We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification. This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/thuml/Time-Series-Library.tabulartime-series-forecasting1M<n<10M9 likes22k downloads11mo agoHugging Face06nvidia /video-to-data-robot-dexterity-task-library-and-dataset Video to Data: Robot Dexterity Task Library and Dataset Dataset Description This dataset contains samples of human demonstrations on manipulation tasks retargeted to bimanual Sharpa robot hands and episodes of robot executions that mimic the original human demonstrations. The former allows a Video to Data user to easily experiment with the Video to Data grounding pipeline, and the latter is an example of the grounded robot data that can be generated with the Video… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-to-data-robot-dexterity-task-library-and-dataset.robotics12 likes16k downloads1mo agoHugging Face07pollen-robotics /reachy-mini-emotions-library Reachy Mini Emotions Library Curated emotion recordings for the Reachy Mini robot, maintained by Pollen Robotics. Each move is a JSON trajectory (head pose, antennas, body yaw, sampled over time) paired with an Opus audio track. Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves non-.wav audio sidecars). File layout Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.audioroboticsn<1K18 likes12k downloads3mo agoHugging Face08VoiceOfML /MLMRL-Library 仓库信息 电报地址:https://t.me/vomebook ,有问题请在:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/discussions 提出。 此仓库存储重要书库备份:https://huggingface.co/datasets/VoiceOfML/MLMRL-Library/tree/main 。 请使用:https://voiceofml-search.hf.space/MLMRL-Library 进行文件检索(备用搜索站:https://vomebook.github.io/search/#/MLMRL-Library )。 可使用:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 进行仓库内容查看(备用站:https://voiceofml-search.hf.space/MLMRL-Library?wide=1 )。 你可以仅下载指针(只有文件名的信息) If you want to clone… See the full description on the dataset page: https://huggingface.co/datasets/VoiceOfML/MLMRL-Library.1 likes6.4k downloads12d agoHugging Face09biglam /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.imageimage-classification1M<n<10M64 likes6.2k downloads1mo agoHugging Face10open-agreements /legal-practice-library legal-practice-library A clean, source-cited snapshot of the OpenAgreements practice-guide corpus: plain-English explainers of US state (and select international) law, currently covering non-compete / restrictive-covenant law and consumer data-privacy law. Published and maintained by openagreements.org. Each note is written against primary law (statutes and cases), carries machine-verifiable source citations, and records the date it was last reviewed. The corpus is re-synced… See the full description on the dataset page: https://huggingface.co/datasets/open-agreements/legal-practice-library.textn<1K0 likes2.9k downloads1d agoHugging Face11pollen-robotics /reachy-mini-dances-library Reachy Mini Dances Library Curated dance moves for the Reachy Mini robot, maintained by Pollen Robotics. Each move is a JSON trajectory (head pose, antennas, body yaw, sampled over time). Motion-only — no audio tracks in this set. File layout Files live at the root of the dataset, named <dance>.json. How to use Python — via the reachy_mini package: from reachy_mini import ReachyMini from reachy_mini.motion.recorded_move import RecordedMoves library… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-dances-library.textroboticsn<1K12 likes2.8k downloads3mo agoHugging Face12Faizaniqbal /british-library-book-images British Library Book Images 1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published between c. 1510 and c. 1900, digitised by the British Library in partnership with Microsoft and released by British Library Labs on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography, philosophy, history, poetry and literature, in several languages. The four image types British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.imageimage-classification1M<n<10M0 likes2.6k downloads1mo agoHugging Face13ieasybooks-org /prophet-mosque-library-compressed Prophet's Mosque Library - Compressed 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.textimage-to-text10K<n<100K0 likes2.3k downloads1y agoHugging Face14ieasybooks-org /waqfeya-library-compressed Waqfeya Library - Compressed 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/waqfeya-library, with one key difference: the contents… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed.tabularimage-to-text10K<n<100K6 likes2k downloads1y agoHugging Face15ahmedforat /spe-petrowiki-raw-librarydocumentn<1K1 likes1.4k downloads2mo agoHugging Face16ziggylott /agent-tts-libraryaudion<1K0 likes1.1k downloads2mo agoHugging Face17common-pile /library_of_congress_filtered Library of Congress Description The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs. Dataset Statistics Documents UTF-8 GB 129,052 35.6 License Issues While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress_filtered.texttext-generation100K<n<1M2 likes990 downloads1y agoHugging Face18remiai3 /synthetic-captchas-library 🌍 Synthetic Multilingual CAPTCHA Library Repository: remiai3/synthetic-captchas-libraryA multilingual dataset of synthetic 4-character CAPTCHA images designed for OCR, multilingual vision models, and script recognition research. This dataset spans 44 world writing systems and is especially useful for low-resource script OCR training. 📌 Dataset Summary Each script includes 100,000 unique CAPTCHA images.The dataset is provided in two parallel formats: CSV version… See the full description on the dataset page: https://huggingface.co/datasets/remiai3/synthetic-captchas-library.image1M<n<10M1 likes985 downloads8mo agoHugging Face19Iuda /sescape-library Sescape sound library, browsing copies Low-bitrate browsing previews (AAC, 80 kbps, previews/) and the catalogue JSON (site/) behind the Sescape library viewer. 81,780 recordings, every one under CC0, a public-domain statement, or CC BY, with the licence and the required credit on each sound's record (site/sounds/<shard>.json, keyed by id; the shard is the top 12 bits of FNV-1a over the id, in hex). Captions, verified labels, timed annotations and similar-sound links come with… See the full description on the dataset page: https://huggingface.co/datasets/Iuda/sescape-library.audioaudio-classification10K<n<100K0 likes915 downloads8d agoHugging Face20biglam /harvard-library-bibliographic-datasettext10M<n<100M2 likes815 downloads1y agoHugging Face21prasaduser /webtoepub-library LinkToEpub Community Library EPUBs converted on https://linktoepub.com. Older books live in prasadonly/webtoepub-library. 0 likes813 downloads3m agoHugging Face22RootCauseAnalytics /Healthcare-Library-Sample RCA Medical Library - 25-Document Review Pack Public sample from the RCA Medical Library. 25 synthetic Australian medical documents across 18 document types with ground truth and bounding boxes. Fully synthetic. No real patient, clinician, hospital, Medicare, MRN, provider or customer data. NSW conventions applied throughout: AU postcodes, valid Medicare number format, AU clinician postnominals (FRACGP, FRACP, FRCPA, FRANZCR, FACEM, FRACS), TRN-PROV provider numbers. SYNTHETIC… See the full description on the dataset page: https://huggingface.co/datasets/RootCauseAnalytics/Healthcare-Library-Sample.documentn<1K1 likes810 downloads4mo agoHugging Face23AtharvImmverse /tiny-captcha-library 🌍 Tiny Multilingual CAPTCHA Library Repository: remiai3/tiny-captcha-libraryA multilingual dataset of synthetic 4-character CAPTCHA images designed for OCR, multilingual vision models, and script recognition research. This dataset spans 44 world writing systems and is especially useful for low-resource script OCR training. 📌 Dataset Summary Each script includes 10,000 unique CAPTCHA images.The dataset is provided in two parallel formats: CSV version (standard… See the full description on the dataset page: https://huggingface.co/datasets/AtharvImmverse/tiny-captcha-library.100K<n<1M0 likes787 downloads2mo agoHugging Face24common-pile /biodiversity_heritage_library Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 42 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.texttext-generation10M<n<100M2 likes706 downloads1y agoHugging Face25common-pile /library_of_congress Library of Congress (subset of Common Pile) Description The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs. This dataset is a subset of the Common Pile v0.1. For more information, see The Common Pile v0.1 paper. Dataset Statistics Documents UTF-8 GB 135,500… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress.texttext-generation100K<n<1M1 likes607 downloads10mo agoHugging Face26remiai3 /tiny-captcha-library 🌍 Tiny Multilingual CAPTCHA Library Repository: remiai3/tiny-captcha-libraryA multilingual dataset of synthetic 4-character CAPTCHA images designed for OCR, multilingual vision models, and script recognition research. This dataset spans 44 world writing systems and is especially useful for low-resource script OCR training. 📌 Dataset Summary Each script includes 10,000 unique CAPTCHA images.The dataset is provided in two parallel formats: CSV version (standard ML… See the full description on the dataset page: https://huggingface.co/datasets/remiai3/tiny-captcha-library.100K<n<1M0 likes582 downloads8mo agoHugging Face27fraug-library /synonyms_dictionnaries Description Apache OpenOffice dictionnaries text1M<n<10M3 likes578 downloads2y agoHugging Face28open-index /open-library Open Library The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links. What is it? Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.tabulartext-generation100M<n<1B9 likes490 downloads6mo agoHugging Face29common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes468 downloads1y agoHugging Face30Amono5667 /webtoepub-librarytext1 likes440 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.