datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adult_income_datasetholobench
HoloBench (Holistic Reasoning Benchmark)
HoloBench is a benchmark designed to evaluate the ability of long-context language models (LCLMs) to perform holistic reasoning over extended text contexts.
Unlike standard models that retrieve isolated information, HoloBench tests how well LCLMs handle complex reasoning tasks that require aggregating and synthesizing information across multiple documents or large text segments.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/holobench.googleAnalyticsCustomerRevenuePredictionmegaunscene
Emergent Extreme-View Geometry in 3D Foundation Models
Yiwen Zhang¹ Joseph Tung² Ruojin Cai³ David Fouhey² Hadar Averbuch-Elor¹
¹Cornell University ²New York University ³Kempner Institute, Harvard University
MegaUnScene Benchmark
Overview
MegaUnScene is a dataset of Internet scenes unseen by existing 3DFMs for benchmarking. There are three test splits split across two evaluation tasks:
Relative Pose Estimation: UnScenePairs and UnScenePairs-t… See the full description on the dataset page: https://huggingface.co/datasets/cornell-vailab/megaunscene.baku_hourly_weather_2015_2025
Baku Yearly Weather Graph
Standalone weather-data project for Baku, Azerbaijan. It includes a weather dataset, a data-generation script that can fetch archive weather data from Open-Meteo, and an interactive daily weather graph.
Links
Project Structure
baku_hourly_weather_2015_2025/
|-- .gitattributes
|-- .gitignore
|-- README.md
|-- requirements.txt
|-- data/
| `-- baku_weather_hourly_data.csv
|-- figures/
| |--… See the full description on the dataset page: https://huggingface.co/datasets/MegrurNiftiyev/baku_hourly_weather_2015_2025.MegaTempQAcreditcardfraudlab05-semantic-searchiter_dpo_pref_chosenGPT4ominirequests_energyFLUTE.stadamgrey88_megastore-sales-data
MegaStore Sales Data
Retail Sales Dataset: Two-Year Record of Transactions for In-Depth Analysis
Dataset Info
Source: Kaggle
Original Size: 8.56 MB
Kaggle Downloads: 4,397
Files: 1
Files
superstoredata.csv
Mirrored from Kaggle
adamgrey88_megastore-sales-data
MegaStore Sales Data
Retail Sales Dataset: Two-Year Record of Transactions for In-Depth Analysis
Dataset Info
Source: Kaggle
Original Size: 8.56 MB
Kaggle Downloads: 4,397
Files: 1
Files
superstoredata.csv
Mirrored from Kaggle
ddg_megadatasetmy-fonts-databaseMentalHealthAppReviewsDatasetFigure_Skating_DataMMFS Figure Skating Dataset (Processed)
This repository contains processed skeleton data for figure skating action classification.
Files included:
skeleton/train_data.pkl — list of numpy arrays; each sample shape (150, 17, 3, 1)
skeleton/test_data.pkl — same as above for test
skeleton/train_label.pkl, skeleton/test_label.pkl — integer class labels
skeleton/train_name.pkl, skeleton/test_name.pkl — original action names (strings)
skeleton/train_score.pkl, skeleton/test_score.pkl — action scores… See the full description on the dataset page: https://huggingface.co/datasets/meghna2801/Figure_Skating_Data.zhorzh-simenon-megre-i-chalavek-na-lautsy
Мэгрэ і чалавек на лаўцы
Metadata
Author: Жорж Сімэнон
Title: Мэгрэ і чалавек на лаўцы
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/zhorzh-simenon-megre-i-chalavek-na-lautsy.MegaDatabaseMegalex
[!NOTE]
Dataset origin: https://openlexicon.fr/datasets-info/FrenchLexiconProject/README-FrenchLexiconProject.html
MEGALEX : méga-étude de la reconnaissance des mots écrits et parlés
Megalex provides visual and auditory lexical decision times and accuracy rates several thousands of words: Visual lexical decision data are available for 28466 French words and the same number of pseudowords, and auditory lexical decision data are available for 17876 French words and the same number of… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/Megalex.
