CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes291k downloads1y agoHugging Face02ieasybooks-org /prophet-mosque-library-compressed Prophet's Mosque Library - Compressed 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.textimage-to-text10K<n<100K0 likes1.3k downloads1y agoHugging Face03mospira /sp500-daily-candles-2025 S&P 500 Daily Candles (2025) This dataset provides daily OHLCV (Open, High, Low, Close, Volume) candles for all S&P 500 tickers between 01-01-2025 and 10-04-2025. Dataset Summary Date range: 2025-01-01 → 2025-10-04 Frequency: 1 day Fields: ticker, date, open, high, low, close, volume File format: CSV (sp500-daily-tickers-2025.csv) Example Schema Column Type Description ticker string Stock symbol (e.g., AAPL, MSFT, AMZN) date datetime… See the full description on the dataset page: https://huggingface.co/datasets/mospira/sp500-daily-candles-2025.tabulartime-series-forecasting100K<n<1M1 likes205 downloads9mo agoHugging Face04Thoria /mandarin-most-common-words-tr-en Mandarin Most Common Words (TR-EN) Overview The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis. This dataset was created by Stephanie Liu and Kamil Murat Yilmaz. Dataset Content The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.text1K<n<10K2 likes179 downloads5mo agoHugging Face05mospira /sp500-daily-candles-2024 SPY Daily Candles 2024 This dataset contains daily OHLCV (Open, High, Low, Close, Volume) candlestick data for all tickers listed on the S&P 500 from 01-01-2024 to 01-01-2025. Columns ticker, date, open, high, low, close, volume tabulartime-series-forecasting100K<n<1M1 likes156 downloads1y agoHugging Face06Growing-Moss-Data /automotive-service-intelligence-sample 🚗 Automotive Service Intelligence Sample Dataset Connected • Longitudinal • Feature-Engineered • Commercially Available This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development. Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.tabulartabular-classification1K<n<10K2 likes122 downloads3mo agoHugging Face07vintp /CMU-Mosei-texttabular10K<n<100K2 likes113 downloads2y agoHugging Face08katielink /moses Molecular Sets (MOSES): A benchmarking platform for molecular generation models Deep generative models are rapidly becoming popular for the discovery of new molecules and materials. Such models learn on a large collection of molecular structures and produce novel compounds. In this work, we introduce Molecular Sets (MOSES), a benchmarking platform to support research on machine learning for drug discovery. MOSES implements several popular molecular generation models and provides a… See the full description on the dataset page: https://huggingface.co/datasets/katielink/moses.text1M<n<10M5 likes109 downloads3y agoHugging Face09MosaicBenchmark /mosaic-bench MOSAIC 199 compositional attack chains across 10 real-world web applications, used to benchmark whether AI coding agents will compose individually-routine tickets into a deployable vulnerability. Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark Datasheet: DATASHEET.md · Croissant 1.1: croissant.json What's in this release Artifact Contents mosaic-bench.xlsx Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.tabulartext-generationn<1K0 likes82 downloads5mo agoHugging Face10Mosab-Rezaei /19th-century-novelists19th-century novelists' sentences We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.tabulartext-generation100K<n<1M2 likes69 downloads7mo agoHugging Face11antoinebcx /smiles-molecules-moses MOSES Molecule Generation Dataset Dataset Description Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset. Task Description For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.text1M<n<10M2 likes67 downloads2y agoHugging Face12SOTAagi2030 /MosaicHarbor-Intake MosaicHarbor intake register Clearance state: verified Intake register shipment_id port review_state right_refs HA-10 north pier queued RH-1 HA-11 blue quay cleared RH-1 HA-12 east basin cleared RH-2 HA-13 south dock cleared RH-3 HA-14 old harbor held RH-4 HA-15 ferry point cleared RH-8 HA-16 amber wharf cleared RH-7 HA-17 west inlet cleared RH-4 HA-18 market berth cleared RH-4 HA-19 granite pier cleared RH-4 HA-20 reed landing… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/MosaicHarbor-Intake.textn<1K0 likes58 downloads11d agoHugging Face13jason1966 /abdulszz_spotify-most-streamed-songs Spotify Most Streamed Songs Unveiling Streaming: A Comprehensive Analysis of Spotify’s Most Streamed Songs Dataset Info Source: Kaggle Original Size: 0.06 MB Kaggle Downloads: 25,259 Files: 1 Files Spotify Most Streamed Songs.csv Mirrored from Kaggle tabularn<1K0 likes53 downloads6mo agoHugging Face14shinnew /CMU-MOSEI_sample🧠 CMU-MOSEI Balanced Subset by Modality This dataset is a compact, balanced subset of CMU-MOSEI, representing only the samples specified in balanced_emotion_by_mean.csv. Each modality (audio, text, vision, labels) has been extracted separately and contains only the relevant data based on the specified video_ids. This makes it ideal for lightweight multimodal learning, benchmarking, and fine-grained feature analysis. 📁 Folder Structure dataset_root/ ├── acoustics/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/shinnew/CMU-MOSEI_sample.tabular1K<n<10K0 likes41 downloads1y agoHugging Face15SmartQHSE /osha-most-cited-standards-2024 Canonical landing page: https://www.smartqhse.com/datasets/osha-most-cited-standards-2024 OSHA Most-Cited Standards FY2024 Top 30 most-frequently-cited OSHA standards in US fiscal year 2024, with citation focus area, primary scope (general industry / construction), and typical Serious-citation penalty range. Sourced from OSHA enforcement data (osha.gov/data/enforcement). Useful for compliance prioritisation, training curriculum design, and contractor pre-qualification… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/osha-most-cited-standards-2024.tabularn<1K0 likes41 downloads4mo agoHugging Face16yash2111 /most-in-demand-skills-2026 Most In-Demand Job Skills of 2026 Skill-demand frequencies extracted from 360,000+ job postings collected by Qarera between December 27, 2025 and June 16, 2026. 📊 Full report & charts: The Most In-Demand Skills of 2026 🔖 Cite this dataset (DOI): 10.5281/zenodo.21204423 📄 License: CC BY 4.0 — free to use with attribution to Qarera. Key findings We counted the skills named in 360,000+ job postings (Dec 2025–Jun 2026). "AI" was the #2 most-requested skill overall… See the full description on the dataset page: https://huggingface.co/datasets/yash2111/most-in-demand-skills-2026.tabularn<1K0 likes41 downloads3mo agoHugging Face17itomomoko /most-red-8e99a9 most-red-8e99a9 Synthetic weather test data: 35 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/itomomoko/most-red-8e99a9.tabularn<1K0 likes39 downloads11d agoHugging Face18steven-sanchez /most-leg-37a063 most-leg-37a063 Synthetic weather test data: 36 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/steven-sanchez/most-leg-37a063.tabularn<1K0 likes38 downloads11d agoHugging Face19Gopher-Lab /TikTok_MostComment_Video_Transcription_Example 📲 Example Dataset: TikTok Scraper Tool 👉 Start Scraping TikTok: TikTok Scraper Tool ✨ Key Features ⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript 🎯 Metadata – Get the title, language description, and video hashtags 🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping 🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools 💸 Free Tier – Use up to 100 queries during the beta period 💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_MostComment_Video_Transcription_Example.texttext-classification1K<n<10K1 likes27 downloads1y agoHugging Face20Gopher-Lab /TikTok_Most_Shared_Video_Transcription_Example 📲 Example Dataset: TikTok Scraper Tool 👉 Start Scraping TikTok: TikTok Scraper Tool ✨ Key Features ⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript 🎯 Metadata – Get the title, language description, and video hashtags 🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping 🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools 💸 Free Tier – Use up to 100 queries during the beta period 💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Most_Shared_Video_Transcription_Example.texttext-classification1K<n<10K3 likes26 downloads1y agoHugging Face21moseleydev /fatima_blind_spot_challengeGot it. From now on I'll write everything inside Markdown blocks so you can copy easily. Here is your full content entirely in Markdown: # Fatima Fellowship 2026: Technical Challenge - Model Blind Spots ## 1. Model Overview - **Model Tested:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Base Model) - **Parameters:** 0.6B - **Type:** Causal Language Model (Base / Pre-trained) --- ## 2. Methodology & Loading To evaluate the model, I used **Google Colab** with a **T4… See the full description on the dataset page: https://huggingface.co/datasets/moseleydev/fatima_blind_spot_challenge.texttext-generationn<1K0 likes23 downloads7mo agoHugging Face22mostafafhasjk /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes22 downloads7mo agoHugging Face23TheNoob3131 /mosquito-dataimage10K<n<100K1 likes21 downloads4y agoHugging Face24Oshara /nepali-tts-mos-resultstabular1K<n<10K0 likes20 downloads29d agoHugging Face25Mostafa8Mehrabi /insomnia-dataset-with-cottext1K<n<10K1 likes13 downloads1y agoHugging Face26mostafaamiri /khamenei_ir_1352_1403_08_13text1K<n<10K2 likes12 downloads2y agoHugging Face27mostafaamiri /fajr_film_festivalload_dataset("mostafaamiri/fajr_film_festival") text1K<n<10K0 likes12 downloads2y agoHugging Face28lewoniewski /most-cited-wikipedia-articlesWikipedia is a massive repository of human knowledge. The largest edition, the English Wikipedia, contains over 65.5 million pages, including 7.17 million articles (excluding redirects). Connecting this vast network are 1.63 billion unique page-to-page links. Based on an analysis of this dataset, the most cited articles on the English Wikipedia were identified. When considering what these most cited articles in Wikipedia might be, we can assume that prominent historical topics like “United… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/most-cited-wikipedia-articles.tabular1K<n<10K0 likes12 downloads4mo agoHugging Face29MosesTan281 /LegalLLMHK Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/MosesTan281/LegalLLMHK.textn<1K2 likes10 downloads2y agoHugging Face30Mostafa3zazi /wonders_testing_dataset Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa3zazi/wonders_testing_dataset.image10K<n<100K0 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.