datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prophet-mosque-library
Prophet's Mosque Library
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.prophet-mosque-library-compressed
Prophet's Mosque Library - Compressed
📖 Overview
Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories.
In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX.
📊 Dataset Contents
This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.sp500-daily-candles-2025
S&P 500 Daily Candles (2025)
This dataset provides daily OHLCV (Open, High, Low, Close, Volume) candles for all S&P 500 tickers between 01-01-2025 and 10-04-2025.
Dataset Summary
Date range: 2025-01-01 → 2025-10-04
Frequency: 1 day
Fields: ticker, date, open, high, low, close, volume
File format: CSV (sp500-daily-tickers-2025.csv)
Example Schema
Column
Type
Description
ticker
string
Stock symbol (e.g., AAPL, MSFT, AMZN)
date
datetime… See the full description on the dataset page: https://huggingface.co/datasets/mospira/sp500-daily-candles-2025.mandarin-most-common-words-tr-en
Mandarin Most Common Words (TR-EN)
Overview
The Mandarin Most Common Words (TR-EN) dataset is a comprehensive trilingual vocabulary resource designed for learners of Mandarin Chinese. It provides translations and practical examples in both Turkish and English, making it highly useful for bilingual education, language learning apps, and linguistic analysis.
This dataset was created by Stephanie Liu and Kamil Murat Yilmaz.
Dataset Content
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/Thoria/mandarin-most-common-words-tr-en.sp500-daily-candles-2024
SPY Daily Candles 2024
This dataset contains daily OHLCV (Open, High, Low, Close, Volume) candlestick data for all tickers listed on the S&P 500 from 01-01-2024 to 01-01-2025.
Columns
ticker, date, open, high, low, close, volume
automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.CMU-Mosei-textmoses
Molecular Sets (MOSES): A benchmarking platform for molecular generation models
Deep generative models are rapidly becoming popular for the discovery of new molecules and materials. Such models learn on a large collection of molecular structures and produce novel compounds. In this work, we introduce Molecular Sets (MOSES), a benchmarking platform to support research on machine learning for drug discovery. MOSES implements several popular molecular generation models and provides a… See the full description on the dataset page: https://huggingface.co/datasets/katielink/moses.mosaic-bench
MOSAIC
199 compositional attack chains across 10 real-world web applications, used to
benchmark whether AI coding agents will compose individually-routine tickets
into a deployable vulnerability.
Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark
Datasheet: DATASHEET.md · Croissant 1.1: croissant.json
What's in this release
Artifact
Contents
mosaic-bench.xlsx
Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.19th-century-novelists19th-century novelists' sentences
We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.smiles-molecules-moses
MOSES Molecule Generation Dataset
Dataset Description
Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.MosaicHarbor-Intake
MosaicHarbor intake register
Clearance state: verified
Intake register
shipment_id
port
review_state
right_refs
HA-10
north pier
queued
RH-1
HA-11
blue quay
cleared
RH-1
HA-12
east basin
cleared
RH-2
HA-13
south dock
cleared
RH-3
HA-14
old harbor
held
RH-4
HA-15
ferry point
cleared
RH-8
HA-16
amber wharf
cleared
RH-7
HA-17
west inlet
cleared
RH-4
HA-18
market berth
cleared
RH-4
HA-19
granite pier
cleared
RH-4
HA-20
reed landing… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/MosaicHarbor-Intake.abdulszz_spotify-most-streamed-songs
Spotify Most Streamed Songs
Unveiling Streaming: A Comprehensive Analysis of Spotify’s Most Streamed Songs
Dataset Info
Source: Kaggle
Original Size: 0.06 MB
Kaggle Downloads: 25,259
Files: 1
Files
Spotify Most Streamed Songs.csv
Mirrored from Kaggle
CMU-MOSEI_sample🧠 CMU-MOSEI Balanced Subset by Modality
This dataset is a compact, balanced subset of CMU-MOSEI, representing only the samples specified in balanced_emotion_by_mean.csv. Each modality (audio, text, vision, labels) has been extracted separately and contains only the relevant data based on the specified video_ids. This makes it ideal for lightweight multimodal learning, benchmarking, and fine-grained feature analysis.
📁 Folder Structure
dataset_root/
├── acoustics/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/shinnew/CMU-MOSEI_sample.osha-most-cited-standards-2024
Canonical landing page: https://www.smartqhse.com/datasets/osha-most-cited-standards-2024
OSHA Most-Cited Standards FY2024
Top 30 most-frequently-cited OSHA standards in US fiscal year 2024, with citation focus area, primary scope (general industry / construction), and typical Serious-citation penalty range. Sourced from OSHA enforcement data (osha.gov/data/enforcement). Useful for compliance prioritisation, training curriculum design, and contractor pre-qualification… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/osha-most-cited-standards-2024.most-in-demand-skills-2026
Most In-Demand Job Skills of 2026
Skill-demand frequencies extracted from 360,000+ job postings collected by Qarera between December 27, 2025 and June 16, 2026.
📊 Full report & charts: The Most In-Demand Skills of 2026
🔖 Cite this dataset (DOI): 10.5281/zenodo.21204423
📄 License: CC BY 4.0 — free to use with attribution to Qarera.
Key findings
We counted the skills named in 360,000+ job postings (Dec 2025–Jun 2026).
"AI" was the #2 most-requested skill overall… See the full description on the dataset page: https://huggingface.co/datasets/yash2111/most-in-demand-skills-2026.most-red-8e99a9
most-red-8e99a9
Synthetic weather test data: 35 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/itomomoko/most-red-8e99a9.most-leg-37a063
most-leg-37a063
Synthetic weather test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/steven-sanchez/most-leg-37a063.TikTok_MostComment_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_MostComment_Video_Transcription_Example.TikTok_Most_Shared_Video_Transcription_Example
📲 Example Dataset: TikTok Scraper Tool
👉 Start Scraping TikTok: TikTok Scraper Tool
✨ Key Features
⚡ Instant Transcription – Turn any TikTok video into an AI-ready transcript
🎯 Metadata – Get the title, language description, and video hashtags
🔗 URL-Based Access – Just drop in a TikTok video URL to start scraping
🧩 LLM-Ready Output – Receive clean JSON ready for agents, RAG, or AI tools
💸 Free Tier – Use up to 100 queries during the beta period
💫 Easy… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/TikTok_Most_Shared_Video_Transcription_Example.fatima_blind_spot_challengeGot it. From now on I'll write everything inside Markdown blocks so you can copy easily.
Here is your full content entirely in Markdown:
# Fatima Fellowship 2026: Technical Challenge - Model Blind Spots
## 1. Model Overview
- **Model Tested:** [Qwen/Qwen3-0.6B](https://huggingface.co/Qwen/Qwen3-0.6B) (Base Model)
- **Parameters:** 0.6B
- **Type:** Causal Language Model (Base / Pre-trained)
---
## 2. Methodology & Loading
To evaluate the model, I used **Google Colab** with a **T4… See the full description on the dataset page: https://huggingface.co/datasets/moseleydev/fatima_blind_spot_challenge.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.mosquito-datanepali-tts-mos-resultsinsomnia-dataset-with-cotkhamenei_ir_1352_1403_08_13fajr_film_festivalload_dataset("mostafaamiri/fajr_film_festival")
most-cited-wikipedia-articlesWikipedia is a massive repository of human knowledge. The largest edition, the English Wikipedia, contains over 65.5 million pages, including 7.17 million articles (excluding redirects). Connecting this vast network are 1.63 billion unique page-to-page links. Based on an analysis of this dataset, the most cited articles on the English Wikipedia were identified.
When considering what these most cited articles in Wikipedia might be, we can assume that prominent historical topics like “United… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/most-cited-wikipedia-articles.LegalLLMHK
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/MosesTan281/LegalLLMHK.wonders_testing_dataset
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa3zazi/wonders_testing_dataset.
