CoolFace
20 results

Books

applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes377k downloads2y agoHugging FaceMohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes33k downloads2y agoHugging Facehozifa1 /noor-platform-books0 likes19k downloads12d agoHugging Facechcaa /kb-books open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish Library published between 1750 and 1930. The dataset has each page of each document in image and text format. The text was extracted with OCR. The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.image1M<n<10M4 likes19k downloads10mo agoHugging FaceHelsinki-NLP /opus_books Dataset Card for OPUS Books Dataset Summary This is a collection of copyright free books aligned by Andras Farkas, which are available from http://www.farkastranslations.com/bilingual_books.php Note that the texts are rather dated due to copyright issues and that some of them are manually reviewed (check the meta-data at the top of the corpus files in XML). The source is multilingually aligned, which is available from http://www.farkastranslations.com/bilingual_books.php.… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/opus_books.texttranslation1M<n<10M96 likes10k downloads2y agoHugging Facetiendung /vi-books_tve-4u.orgTODOs dtv-ebook.com bóc tách và convert dtv ebooks 3543 files | 4.5G (without PDF) dtv_categories_whitelist.jsonl cho vào ../zinz/30-vien_book_epub_van-hoc/ phần còn lại để vào ../zinz/40-vi_character_entertain/ xếp dữ liệu theo từng categories tve-4u.org bóc tách và convert thành text ~11k files | 7.2G (without PDF) download_links_with_post_url liệt kê toàn bộ download links để download ebooks và truy xuất ngược về post có chưa download link tương ứng posts_with_content là các… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/vi-books_tve-4u.org.0 likes9.1k downloads3y agoHugging Face