CoolFace
29 results

MOS

ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes291k downloads1y agoHugging FaceAISE-TUDelft /MOSAIC-Refactoring Agentic Pull Request Dataset Dataset Overview The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below. Cohort Pull Requests Merged Pull Requests Repositories Sum of Additions Sum of Deletions Humans 517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.tabular10M<n<100M4 likes29k downloads3mo agoHugging Faceoscar-corpus /mOSCARMore info can be found here: https://oscar-project.github.io/documentation/versions/mOSCAR/ Paper link: https://arxiv.org/abs/2406.08707 New features: Additional filtering steps were applied to remove toxic content (more details in the next version of the paper, coming soon). Spanish split is now complete. Face detection in images to blur them once downloaded (coordinates are reported on images of size 256 respecting aspect ratio). Additional language identification of the documents to… See the full description on the dataset page: https://huggingface.co/datasets/oscar-corpus/mOSCAR.text100M<n<1B18 likes9.5k downloads2y agoHugging Facemalaysia-ai /mosaic-combine-all Mosaic format for combine all dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all load it, from streaming import LocalDataset import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.textn<1K0 likes8.2k downloads3y agoHugging Facelogantl /mos0 likes6.3k downloads1y agoHugging Facemostafabehroozi /Transformation0 likes5.1k downloads3d agoHugging Face

People

Projects