datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sec-edgar-exec-comp
SEC EDGAR Executive Compensation — Currently Unavailable
This dataset is not currently available for download or purchase.
The previous description advertised compensation data covering 2015–present,
500,000+ records, and a 1,000-row sample. Those claims were not supported by
the files in this repository and have been withdrawn.
The existing sample_1000.csv contains column headers only: zero data rows.
It is a schema placeholder, not a usable sample. The full dataset was removed… See the full description on the dataset page: https://huggingface.co/datasets/claritystorm/sec-edgar-exec-comp.rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
mise install
mise run setup
mise run test
mise exec -- python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md before changing the pipeline.
Replay files and… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.VectorBenchmarkEditVerseBench
EditVerse
This repository contains the instruction-based video editing evaluation benchmark for EditVerseBench in paper "EditVerse: A Unified Framework for Editing and Generation via In-Context Learning".
Xuan Ju12, Tianyu Wang1, Yuqian Zhou1, He Zhang1, Qing Liu1, Nanxuan Zhao1, Zhifei Zhang1, Yijun Li1, Yuanhao Cai3, Shaoteng Liu1, Daniil Pakhomov1, Zhe Lin1, Soo Ye Kim1*, Qiang Xu2*
1Adobe Research 2The Chinese University of Hong Kong 3Johns Hopkins University *Corresponding… See the full description on the dataset page: https://huggingface.co/datasets/sooyek/EditVerseBench.zomato_delivery_EDA📹 Video walkthrough:
Zomato Delivery Operations — EDA & Dataset
Dataset Overview
Real-world delivery data from Zomato operations across multiple Indian cities,
covering courier attributes, weather conditions, traffic density, GPS coordinates,
and delivery outcomes.
Source
Kaggle — saurabhbadole/zomato-delivery-operations-analytics-dataset
Original size
45,584 rows × 20 columns
Final size
38,964 rows × 22 columns
Target variable
Time_taken (min)… See the full description on the dataset page: https://huggingface.co/datasets/allenborochin/zomato_delivery_EDA.zinc250kzinc250k contains the 250k molecule subset used in the "Automatic Chemical Design Using a
Data-Driven Continuous Representation of Molecules" paper (doi:10.1021/acscentsci.7b00572).
The dataset contains the original columns from
https://github.com/aspuru-guzik-group/chemical_vae/blob/main/models/zinc_properties/250k_rndm_zinc_drugs_clean_3.csv,
namely smiles, logP, QED, and SAS and an additional selfies column.
This dataset can be used for benchmarking chemical Language Models or training… See the full description on the dataset page: https://huggingface.co/datasets/edmanft/zinc250k.13f-institutional-holdings-sec-edgar
13F Institutional Holdings Dataset — SEC EDGAR Hedge Fund & Asset Manager Filings
A structured, ready-to-analyze snapshot of institutional 13F filings covering 13,000+ investment managers — hedge funds, mutual fund families, pension funds, banks, and family offices — built from raw SEC Form 13F data on EDGAR. Each row is one manager's most recently disclosed quarter: total portfolio value, position count, and five behavioral scores (concentration, turnover, momentum/contrarian… See the full description on the dataset page: https://huggingface.co/datasets/JamesFromAlphasmo/13f-institutional-holdings-sec-edgar.au-editingEDGAR-CORPUS-Financial-Summarization
EDGAR-CORPUS : 10K Financial Report Summarization
Extracted from SEC EDGAR filings (1993-2020). This dataset enhances financial report summarization by leveraging a hybrid AI model strategy.
Using:
ChatGPT-3.5 Turbo(~70%),
Claude 3.5 (~30% to generate structured, accurate, and concise summaries)
Dataset Composition
Summaries in this dataset are generated using a hybrid AI model strategy, balancing quality and efficiency:ChatGPT-3.5 Turbo (~70%) – Used for structured… See the full description on the dataset page: https://huggingface.co/datasets/kritsadaK/EDGAR-CORPUS-Financial-Summarization.bitcoin_newsBitcoin news scrapped from Yahoo Finance.
Columns:
time_unix the UNIX timestamp of the news (UTC)
date_time UTC date and time
text_matches the news articles are matched with keywords "BTC", "bitcoin", "crypto", "cryptocurrencies", "cryptocurrency". The list is the posititions the keywords appeared.
title_matches keyword matches in title
url the Yahoo Finance URL that the article from
source the source if the news is cited from other source, not originally from Yahoo Finane
source_url the outer… See the full description on the dataset page: https://huggingface.co/datasets/edaschau/bitcoin_news.edge-ai-kg
Edge AI Deployment Knowledge Graph
25,152 nodes. 76,306 edges. Boards, kernels and neural networks in one graph — so you can
ask what actually runs on your silicon.
Built with Samyama Graph.
Loader and generator: samyama-ai/edge-ai-kg.
Part real, part synthetic — and every node says which
Every node carries a provenance property ("real" or "synthetic") and a source. No
node is unstamped:
provenance
Nodes
synthetic
23,910
real
1,242
Do not… See the full description on the dataset page: https://huggingface.co/datasets/VaidhyaMegha/edge-ai-kg.pcqnp-finite-shot-reliability-artifact
PC-QNP Finite-Shot Reliability Artifact
This anonymous review artifact supports a NeurIPS 2026 Evaluations & Datasets submission on physics-conformal reliability evaluation for finite-shot quantum-process surrogate models. The asset is intended for reviewer inspection and reviewer-level reproduction of the reported aggregate tables, nested residual-repair calculation, and IBM stochastic Pauli-channel protocol validation.
Contents
code_snapshot/: cleaned source-code… See the full description on the dataset page: https://huggingface.co/datasets/anon-pcqnp-ed26/pcqnp-finite-shot-reliability-artifact.instagram-engagement-edaHorse-Race-Prediction-EDA
🏇 Horse Race Prediction: Exploratory Data Analysis (EDA)
Project Walkthrough
לחצו כאן לצפייה בסרטון ההסבר (Loom)
במידה והסרטון לא עולה - ניתן לצפות בסרטון בלינק למעלה *
📌 Project Overview
This project presents a comprehensive Exploratory Data Analysis (EDA) of a 2019 horse racing dataset containing over 171,849 records. The goal was to identify the primary biological, professional, and market factors that determine a winning performance.… See the full description on the dataset page: https://huggingface.co/datasets/mayacheruty/Horse-Race-Prediction-EDA.spanish-fake-news-fixed
Spanish Fake News Fixed
Este dataset contiene noticias etiquetadas en español, reparado para corregir saltos de línea internos.
nba-career-stats-eda
🏀 NBA Player Career Stats — EDA Project
Overview
This project presents an end-to-end Exploratory Data Analysis (EDA) of NBA player
career statistics. The goal is to uncover patterns in player performance, compare
active vs. retired players, and explore relationships between key basketball stats.
Source: Hatman/NBA-Player-Career-Stats
Original size: 3,093 rows × 28 columns
Final clean size: 3,078 rows × 23 columns
Target Variable: IS_ACTIVE (True = Active / False =… See the full description on the dataset page: https://huggingface.co/datasets/Omerinbar/nba-career-stats-eda.craigslist-used-cars-eda
Craigslist Used Cars and Trucks: EDA
Overview
This dataset and notebook contain an Exploratory Data Analysis (EDA) of real Craigslist used-car listings scraped across the United States.
Main Question: What factors most influence the price of a used car listed on Craigslist?
Target Variable: price — the seller's asking price for each vehicle listing.
About the Dataset
Property
Details
Source
Kaggle — Austin Reese (scraped from Craigslist)
Original… See the full description on the dataset page: https://huggingface.co/datasets/Yoad22/craigslist-used-cars-eda.manim_pythonSASS-Bench
SASS-Bench
Hardware-grounded benchmark for execution prediction on NVIDIA GPU assembly (SASS).
Given a compiled SASS kernel, launch configuration, raw inputs, and a list of
checkpoint queries, a model must predict exact uint32 register values at
specified instructions and the bit-pattern of the final output buffers.
Ground truth is captured from a Blackwell sm_120 GPU via NVBit dynamic
binary instrumentation. The initial release is 50 kernels with 2--3 input-regime
variants each… See the full description on the dataset page: https://huggingface.co/datasets/anon-neurips-ed-3552/SASS-Bench.UVE-Subjective-Benchmark
Dataset Card for UVE-Subjective-Benchmark
🌊 Dataset Summary
The evaluation of Underwater Video Enhancement (UVE) remains a significant challenge due to the complex visual degradations inherent to aquatic environments and the temporal instability frequently introduced by frame-wise processing. Existing objective metrics (e.g., UIQM, UCIQE) often correlate poorly with human subjective judgement, particularly regarding dynamic artifacts like flickering and color… See the full description on the dataset page: https://huggingface.co/datasets/eddy-Wang/UVE-Subjective-Benchmark.us-university-college-school-district-education-layoffs-warn-act-notices-daily
US university, college, school and education layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-24. 1,165 layoff and closure notices filed by
universities and colleges, for-profit career schools and their campuses, private and charter schools, school districts where a state chose to publish them, Head Start and childcare providers, school-bus and campus contractors, and education publishers and online-learning companies with US state labor departments… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-university-college-school-district-education-layoffs-warn-act-notices-daily.Par-Four-Fineweb-Edu-FortifiedDataset Summary
This dataset is a filtered subset of the Fineweb-Edu-Fortified dataset. The primary goal of this subset is to reduce the dataset size to a more manageable volume while maintaining high-quality content. It contains three key fields: score, text, and url, focusing on entries with a score of 4 and above, indicating higher relevance and quality of educational content.
This dataset can be used for several fine-tuning and model improvement tasks, including model healing, synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified.Video-Games-Sales-EDA
🎮 Video Game Sales — Exploratory Data Analysis (EDA)
Video Presentation
If I didn’t cover everything it’s because I didn’t have enough time
Your browser does not support the video tag.
Executive Summary
The video game industry is a multi‑billion dollar market characterized by extreme unpredictability—a single "mega‑hit" can generate more revenue than thousands of average games combined. This project analyzes historical video game sales… See the full description on the dataset page: https://huggingface.co/datasets/0tizm0/Video-Games-Sales-EDA.edgar-ma-deals
EDGAR M&A Deal Events (2010–2024)
4,156 U.S. public-company acquisition events, discovered directly from SEC
EDGAR's own quarterly filing indexes — not scraped from a vendor list or a
blog. Each row is a company that filed a merger proxy or tender-offer response
between 2010 and 2024, with the earliest such filing's date as an
announcement-date proxy.
Built as a byproduct of an M&A target-prediction research project. Free,
reproducible, and honestly caveated — an open… See the full description on the dataset page: https://huggingface.co/datasets/tcontorno/edgar-ma-deals.Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-ReasonExtracted Chemistry Physics Asronomy Math and Logic portions from the original.
Script used for the extraction:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Chemistry-Physics-Astronomy-Math-Reason/resolve/main/find-science-fine.py
vn-provinces-education-master
Vietnam education master panel by locality
Wide geo×year master joining 16 NSO Giáo dục locality packs: preschool, general schools/classes/teachers/pupils (incl. female and ethnic-minority), pupils per class/teacher, solid classroom rate, upper-secondary graduation, university lecturers/students, and 2023 vocational education. Outer join on geo_code×year (2001-2024). R&D / patents / national ownership tables are out of scope. Province names follow ar_core.vn_geo.… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-education-master.edisum_dataset
Dataset Card for Edisum
Dataset Description
For more details:
Github repository: https://github.com/epfl-dlab/edisum
Paper: https://arxiv.org/pdf/2404.03428.pdf
Languages
Edisum only contains Wikipedia data collected from English Wikipedia. Consequently, synthetic data is also only generated in English.
Dataset Structure
The Edisum meta-dataset actually comprises 5 datasets:
wikiepdia_processed_data (Filtered existing Wikipedia data)… See the full description on the dataset page: https://huggingface.co/datasets/msakota/edisum_dataset.Image-Gen-or-Image-Editing
Image Gen or Image Editing
This dataset is designed for text classification of prompts provided by users. It determines whether a prompt is intended for image generation or image editing.
syslog_samplevn-provinces-vocational-education-2023
Vietnam provinces vocational education snapshot 2023
Vietnam provinces vocational education snapshot 2023. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (63 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (6 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-vocational-education-2023.
