datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
uv sync --locked
UV_CACHE_DIR=/tmp/rocketleague-uv-cache \
uv run --locked pytest -v
uv run --locked python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.sonar-rock-mineThe Sonar Rock Mine is a well-known benchmark dataset in the field of machine learning and pattern recognition, originally collected for the purpose of distinguishing between mines (M) and rocks (R) using sonar signals.
Dataset Overview:
File: raw_sonar.csv
Number of Features: 60 continuous numerical values per instance
Target Label: One class label at the end of each row, either:
M → Mine (target object: simulated sea mine)
R → Rock (non-mine object)
Total Instances: The dataset… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/sonar-rock-mine.hf-arxiv-url-bench
arXiv URL Extraction Benchmark & Longitudinal Corpus
Dataset Description
This repository hosts datasets designed to assess format-specific coverage gaps in URL extraction across the arXiv corpus. The collection facilitates large-scale reproducibility studies and evaluates how different document formats (LaTeX, HTML, XML, Markdown, TXT, PNG) impact automated extraction pipelines.
The repository is divided into two primary corpora:
The 200-Paper Benchmark: A… See the full description on the dataset page: https://huggingface.co/datasets/rochanaro/hf-arxiv-url-bench.emojisA collection of 38,176 emoji images from Facebook, Google, Apple, WhatsApp, Samsung, JoyPixels, Twitter, emojidex, LG, OpenMoji, and Microsoft. It includes all the emojis for these apps/platforms as of early 2022.
Counts: Facebook=3664, Google=3664, Apple=3961, WhatsApp=3519, Samsung=3752, JoyPixels=3538, Twitter=3544, emojidex=2040, LG=3051, OpenMoji=3512, Microsoft=3931.
Sizes: Facebook=144x144, Google=144x144, Apple=144x144, WhatsApp=144x144, Samsung=108x108, JoyPixels=144x144… See the full description on the dataset page: https://huggingface.co/datasets/rocca/emojis.top-reddit-postsThe post-data-by-subreddit.tar file contains 5000 gzipped json files - one for each of the top 5000 subreddits (as roughly measured by subscriber count and comment activity). Each of those json files (e.g. askreddit.json) contains an array of the data for the top 1000 posts of all time.
Notes:
I stopped crawling a subreddit's top-posts list if I reached a batch that had a post with a score less than 5, so some subreddits won't have the full 1000 posts.
No posts comments are included. Only the… See the full description on the dataset page: https://huggingface.co/datasets/rocca/top-reddit-posts.ROCOv2-trainspanish-nahuatl-flaggingRoco-csvscontradictory-watson-nli
🔬 Contradictory, My Dear Watson — Multilingual NLI Solution
Kaggle Competition | Model: mDeBERTa-v3-base-mnli-xnli
📋 Problem
Given a premise and hypothesis in one of 15 languages, predict the relationship:
0 = Entailment (hypothesis follows from premise)
1 = Neutral (hypothesis is possible but not certain)
2 = Contradiction (hypothesis contradicts premise)
🏗️ Approach
Model
mDeBERTa-v3-base-mnli-xnli — 279M params, already fine-tuned on MultiNLI… See the full description on the dataset page: https://huggingface.co/datasets/Rock2346/contradictory-watson-nli.ROC-Stories-modifiedroc_masked_middlePop-Rock-Hybrid-Stem-Dataset-cat001
Dataset Overview: Pop Rock Hybrid Stem Dataset (cat001)
This dataset contains a curated collection of original instrumental music designed for commercial and research applications in music analysis, audio modeling, and production workflows.
Every composition, arrangement, performance, sound design element, and production decision was created entirely through human musical and technical processes. All music contained in this dataset is 100% human-made (is_human_created: TRUE).… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Stem-Dataset-cat001.clip-keyphrase-embeddingsThe reddit_keywords.tsv file contains about 170k single word embeddings (scraped from reddit, filtering from an initial set of ~700k based on a minimum occurrence threshold) in this format:
temporary -0.276235,-0.181357,-0.325729,0.129826,0.016490,-0.230246,-0.039997,-0.990187,-0.014679,-0.044081,-0.120046,-0.250614,-0.303871,-0.264685,-0.010019,-0.158764,0.086107,-0.018172,0.003005,-0.383161,0.412182,0.104374,0.041335,-0.018206,0.085453,0.016297,-0.015680,0.047611,-0.267469,0.046825,-0.367247… See the full description on the dataset page: https://huggingface.co/datasets/rocca/clip-keyphrase-embeddings.noamundi-rockfall-simulation-dataset
🪨 Noamundi Rockfall Simulation Dataset
📌 1. Dataset Overview
This dataset provides a physics-informed simulation of rockfall precursor conditions and trajectory runouts for 500 monitoring locations in the Noamundi mining region in Jharkhand, India. It is designed specifically for training and benchmarking Machine Learning models on early-warning systems, anomaly detection, rare-event classification, and spatio-temporal (ST-GNN) trajectory modeling in… See the full description on the dataset page: https://huggingface.co/datasets/Kaizen696/noamundi-rockfall-simulation-dataset.Runchao0502AI_Job_DataSet_1000_listPop-Rock-Hybrid-Mini-Dataset-Indie-Rock-Aggressive-CAT001-S07
Dataset Overview: Pop Rock Hybrid Mini Dataset — Indie Rock: Aggressive (CAT001-S07)
This dataset is a compact evaluation and production subset of the Pop Rock Hybrid Stem Dataset (cat001), designed for commercial AI, machine learning, audio analysis, music information retrieval (MIR), and related technical workflows.
CAT001-S07 contains one seed composition and four engineered variations, with a consistent fullmix and 3-stem architecture across each variation. This provides a… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Mini-Dataset-Indie-Rock-Aggressive-CAT001-S07.Pop-Rock-Hybrid-Mini-Dataset-Indie-Pop-Warm-CAT001-S09
Dataset Overview: Pop Rock Hybrid Mini Dataset — Indie Pop: Warm (CAT001-S09)
This dataset is a compact evaluation and production subset of the Pop Rock Hybrid Stem Dataset (cat001), designed for commercial AI, machine learning, audio analysis, music information retrieval (MIR), and related technical workflows.
CAT001-S09 contains one seed composition and four engineered variations, with a consistent fullmix and 3-stem architecture across each variation. This provides a… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Mini-Dataset-Indie-Pop-Warm-CAT001-S09.accelerometerData
Hand Detection Training Data
This folder contains sensor data collected from mobile devices for training the hand detection model.
Overview
The dataset includes accelerometer and gyroscope readings from 2 subjects, each holding a device with both their left and right hands. This data is used to train the Random Forest classifier that achieves 94.6% accuracy in detecting which hand is holding the device.
Directory Structure
hand_data/
├── accelerometer/… See the full description on the dataset page: https://huggingface.co/datasets/rockerritesh/accelerometerData.ROCSTORIES-MODTest_ENISA_EXTRACTEDPop-Rock-Hybrid-Mini-Dataset-Indie-Pop-Uplifting-CAT001-S01
Dataset Overview: Pop Rock Hybrid Mini Dataset — Indie Pop: Uplifting (CAT001-S01)
This dataset is a compact evaluation and production subset of the Pop Rock Hybrid Stem Dataset (cat001), designed for commercial AI, machine learning, audio analysis, music information retrieval (MIR), and related technical workflows.
CAT001-S01 contains one seed composition and four engineered variations, with a consistent fullmix and 3-stem architecture across each variation. This provides a… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Mini-Dataset-Indie-Pop-Uplifting-CAT001-S01.Pop-Rock-Hybrid-Mini-Dataset-Indie-Pop-Playful-CAT001-S06
Dataset Overview: Pop Rock Hybrid Mini Dataset — Indie Pop: Playful (CAT001-S06)
This dataset is a compact evaluation and production subset of the Pop Rock Hybrid Stem Dataset (cat001), designed for commercial AI, machine learning, audio analysis, music information retrieval (MIR), and related technical workflows.
CAT001-S06 contains one seed composition and four engineered variations, with a consistent fullmix and 3-stem architecture across each variation. This provides a… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Mini-Dataset-Indie-Pop-Playful-CAT001-S06.test2_ENISAroc-llama-modifiedall-restaurants-in-little-rock-north-little-rock-conway-metro-arkansas-us-525955
All Restaurants in Little Rock-North Little Rock-Conway (Metro), Arkansas, US
Free sample dataset from BeamStation
This dataset provides a complete export of every restaurant operating within the Little Rock‑North Little Rock‑Conway metropolitan area of Arkansas. It includes 1,081 records, each representing a distinct establishment and containing all available columns from the source profiles. The data cover a range of signals such as business name, address, contact information… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/all-restaurants-in-little-rock-north-little-rock-conway-metro-arkansas-us-525955.all-restaurants-in-rochester-metro-new-york-us-535392
All Restaurants in Rochester (Metro), New York, US
Free sample dataset from BeamStation
This dataset provides a complete metro export of all restaurants in the Rochester, New York area. It contains 707 records, each representing a restaurant profile with all available columns such as name, address, cuisine type, contact information, and operational details. The data is refreshed on a weekly basis to keep the information current for users who need up-to‑date listings. Researchers… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/all-restaurants-in-rochester-metro-new-york-us-535392.rocRockPebbleSandSituationscounty-population-data
Dataset: rockybranches/county-population-data
US County population data estimates (last updated for the 2020 US Census).
Dataset Details
Dataset Description
Formats:
CSV files containing a hand-curated set of US county-level population estimates.**
.parquet a single aggregated binary; sql-optimized format for high performance access.
Note that the NS Codes are US-level county codes generated by the USGS (GNIS).
Curated by: [Rocky Branches, LLC (Elizabeth… See the full description on the dataset page: https://huggingface.co/datasets/rockybranches/county-population-data.
