datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sp500-daily-candles-2025
S&P 500 Daily Candles (2025)
This dataset provides daily OHLCV (Open, High, Low, Close, Volume) candles for all S&P 500 tickers between 01-01-2025 and 10-04-2025.
Dataset Summary
Date range: 2025-01-01 → 2025-10-04
Frequency: 1 day
Fields: ticker, date, open, high, low, close, volume
File format: CSV (sp500-daily-tickers-2025.csv)
Example Schema
Column
Type
Description
ticker
string
Stock symbol (e.g., AAPL, MSFT, AMZN)
date
datetime… See the full description on the dataset page: https://huggingface.co/datasets/mospira/sp500-daily-candles-2025.sp500-daily-candles-2024
SPY Daily Candles 2024
This dataset contains daily OHLCV (Open, High, Low, Close, Volume) candlestick data for all tickers listed on the S&P 500 from 01-01-2024 to 01-01-2025.
Columns
ticker, date, open, high, low, close, volume
automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.CMU-Mosei-textmosaic-bench
MOSAIC
199 compositional attack chains across 10 real-world web applications, used to
benchmark whether AI coding agents will compose individually-routine tickets
into a deployable vulnerability.
Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark
Datasheet: DATASHEET.md · Croissant 1.1: croissant.json
What's in this release
Artifact
Contents
mosaic-bench.xlsx
Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.19th-century-novelists19th-century novelists' sentences
We constructed the 5-author dataset using texts from Project Gutenberg, focusing on five prominent 19th-century novelists: Charles Dickens, Mark Twain, Herman Melville, Jane Austen, and Louisa May Alcott. This selection balances male and female authors as well as British and American literary traditions, offering a diverse testbed for stylistic analysis. Sentence segmentation was performed with the NLTK library, and tokenization/word counts were… See the full description on the dataset page: https://huggingface.co/datasets/Mosab-Rezaei/19th-century-novelists.abdulszz_spotify-most-streamed-songs
Spotify Most Streamed Songs
Unveiling Streaming: A Comprehensive Analysis of Spotify’s Most Streamed Songs
Dataset Info
Source: Kaggle
Original Size: 0.06 MB
Kaggle Downloads: 25,259
Files: 1
Files
Spotify Most Streamed Songs.csv
Mirrored from Kaggle
osha-most-cited-standards-2024
Canonical landing page: https://www.smartqhse.com/datasets/osha-most-cited-standards-2024
OSHA Most-Cited Standards FY2024
Top 30 most-frequently-cited OSHA standards in US fiscal year 2024, with citation focus area, primary scope (general industry / construction), and typical Serious-citation penalty range. Sourced from OSHA enforcement data (osha.gov/data/enforcement). Useful for compliance prioritisation, training curriculum design, and contractor pre-qualification… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/osha-most-cited-standards-2024.CMU-MOSEI_sample🧠 CMU-MOSEI Balanced Subset by Modality
This dataset is a compact, balanced subset of CMU-MOSEI, representing only the samples specified in balanced_emotion_by_mean.csv. Each modality (audio, text, vision, labels) has been extracted separately and contains only the relevant data based on the specified video_ids. This makes it ideal for lightweight multimodal learning, benchmarking, and fine-grained feature analysis.
📁 Folder Structure
dataset_root/
├── acoustics/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/shinnew/CMU-MOSEI_sample.most-in-demand-skills-2026
Most In-Demand Job Skills of 2026
Skill-demand frequencies extracted from 360,000+ job postings collected by Qarera between December 27, 2025 and June 16, 2026.
📊 Full report & charts: The Most In-Demand Skills of 2026
🔖 Cite this dataset (DOI): 10.5281/zenodo.21204423
📄 License: CC BY 4.0 — free to use with attribution to Qarera.
Key findings
We counted the skills named in 360,000+ job postings (Dec 2025–Jun 2026).
"AI" was the #2 most-requested skill overall… See the full description on the dataset page: https://huggingface.co/datasets/yash2111/most-in-demand-skills-2026.most-leg-37a063
most-leg-37a063
Synthetic weather test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/steven-sanchez/most-leg-37a063.most-red-8e99a9
most-red-8e99a9
Synthetic weather test data: 35 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/itomomoko/most-red-8e99a9.nepali-tts-mos-resultsmosquito-datamost-cited-wikipedia-articlesWikipedia is a massive repository of human knowledge. The largest edition, the English Wikipedia, contains over 65.5 million pages, including 7.17 million articles (excluding redirects). Connecting this vast network are 1.63 billion unique page-to-page links. Based on an analysis of this dataset, the most cited articles on the English Wikipedia were identified.
When considering what these most cited articles in Wikipedia might be, we can assume that prominent historical topics like “United… See the full description on the dataset page: https://huggingface.co/datasets/lewoniewski/most-cited-wikipedia-articles.wonders_testing_dataset
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa3zazi/wonders_testing_dataset.CMU-MOSEI_sample🧠 CMU-MOSEI Balanced Subset by Modality
This dataset is a compact, balanced subset of CMU-MOSEI, representing only the samples specified in balanced_emotion_by_mean.csv. Each modality (audio, text, vision, labels) has been extracted separately and contains only the relevant data based on the specified video_ids. This makes it ideal for lightweight multimodal learning, benchmarking, and fine-grained feature analysis.
📁 Folder Structure
dataset_root/
├── acoustics/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Samuel2007/CMU-MOSEI_sample.MoscowLargeEnterprises
MoscowLargeEnterprises
tags: employment, geolocation, machine learning
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'MoscowLargeEnterprises' dataset contains information about various large companies (with more than 1000 employees) located in Moscow. It includes the company's name, label for the type of business, and their geographical coordinates in decimal format. This dataset is designed for analysis and research in… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MoscowLargeEnterprises.CMU-Mosei-textMoscowIndustrialHubs
MoscowIndustrialHubs
tags: location, size, regression analysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'MoscowIndustrialHubs' dataset contains information about various large enterprises (with more than 1000 employees) located in Moscow, Russia. The dataset includes the company name, the number of employees, and their decimal geographical coordinates (latitude and longitude). This dataset is particularly useful for… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MoscowIndustrialHubs.
