datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nbm-conus-analysis
NOAA NBM CONUS Daily Analysis (Zarr)
Daily best-estimate analysis derived from NOAA NBM (National Blend of Models)
CONUS forecasts, on the native ~2.5 km Lambert conformal grid (2345 x 1597).
Built by nbm-to-zarr, dynamical.org-style.
Variables: tmean / tmax / tmin (degC), precip (mm), srad (MJ/m2/day)
Construction: best estimate for day D = lead-day 1 of that day's 00z NBM init
Coverage: rolling backfill from 2020-10-01 (AWS NBM archive floor) to present
Layout: one standalone… See the full description on the dataset page: https://huggingface.co/datasets/nakas/nbm-conus-analysis.data_analysis
Dataset Card for "livebench/data_analysis"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/data_analysis.Taur_CoT_Analysis_Project___gpt-4o-2024-08-06NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.Lurcher_10x
Lurcher 10x Microscopy Dataset
Dataset overview
This dataset consists of 2-D microscopy images of histologically stained 3-D structures in tissue sections through the cerebellum of 21 mouse brains. Animals are grouped into wild-type controls (n = 10) and Lurcher mutant mice (n = 11). The classification task is to distinguish Lurcher mutant mice from wild-type controls.
All images were captured at low magnification (10x) and stained with Cresyl violet, a general… See the full description on the dataset page: https://huggingface.co/datasets/USF-CS-Microscopy-Image-Analysis/Lurcher_10x.Openpdf-Analysis-Recognition
Openpdf-Analysis-Recognition
The Openpdf-Analysis-Recognition dataset is curated for tasks related to image-to-text recognition, particularly for scanned document images and OCR (Optical Character Recognition) use cases. It contains over 6,900 images in a structured imagefolder format suitable for training models on document parsing, PDF image understanding, and layout/text extraction tasks.
Attribute
Value
Task
Image-to-Text
Modality
Image
Format
ImageFolder… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Openpdf-Analysis-Recognition.Taur_CoT_Analysis_Project___meta-llama__Meta-Llama-3.1-8B-Instructweb3-trading-analysisThis dataset contains web3-related on-chain and off-chain data, which can be used to build quantitative models.
Taur_CoT_Analysis_Project___gpt-4o-mini-2024-07-18rocketleague-analysis
Rocket League Analysis
Local Rocket League replay analysis using Ballchasing API exports and plain DuckDB.
The report is meant to answer one practical question: what should I work on next from my saved replay sample?
Quick Start
uv sync --locked
UV_CACHE_DIR=/tmp/rocketleague-uv-cache \
uv run --locked pytest -v
uv run --locked python scripts/analyze_scenarios.py \
--replay-dir /path/to/Rocket\ League/TAGame/Demos \
--limit 10
Start with CONTRIBUTING.md… See the full description on the dataset page: https://huggingface.co/datasets/edmundmiller/rocketleague-analysis.routing_analysis-checkpoints
routing_analysis checkpoint archive
This public dataset repository stores checkpoint files from the
routing_analysis filesystem snapshot while preserving their original paths
under routing_analysis/.
The tree routing_analysis/finetuning/phase2_full/checkpoints/ is explicitly
excluded. All other regular files classified under checkpoint directories are
included, including small code and configuration files needed to keep those
checkpoint directories complete.
Files are uploaded… See the full description on the dataset page: https://huggingface.co/datasets/lylybig/routing_analysis-checkpoints.VanGogh_vs_TreeOilPainting_QuantumTorque_EnergyField_Analysis_Phase1_2025
Dataset Policy
VanGogh Vs. Tree Oil Painting: Quantum Torque Energy Field Analysis 2025
Structure Type
Free-form and Semi-structured Narrative
Core Principles
Each file is an independent analytical entity with its own identity.
Each file is the result of Autonomous AI–Human Co-analysis.
The structure is intentionally open, flexible, and adaptive, reflecting the natural reasoning process of the researcher, rather than forcing rigid… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/VanGogh_vs_TreeOilPainting_QuantumTorque_EnergyField_Analysis_Phase1_2025.Taur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instructembedding-optimizer-study-analysis-artifacts
Embedding optimizer study analysis artifacts
This repository preserves analysis artifacts for the current DenseOn comparison
of AdamW, Muon and NorMuon. The source repository and paper
contain the completed experiments, protocols, exact analysis and restoration tools.
Model and optimizer states are in the separate
checkpoint repository.
Current scientific artifacts
Use the immutable revisions and manifests in these guides, not a broad download
of the mixed-history… See the full description on the dataset page: https://huggingface.co/datasets/qcz/embedding-optimizer-study-analysis-artifacts.Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001india-tb-missed-cases-analysis
India TB Missed Cases Analysis & Living Model (2025)
🌟 Project Overview
This repository hosts a comprehensive, multi-method analytical framework designed to estimate and understand the "missing" millions of Tuberculosis (TB) cases in India. By integrating Bayesian statistics, Dimensionality Reduction (PCA), and Causal Inference (DAG), this project provides a high-resolution view of TB detection determinants across Indian states.
Core Analytical Pillars:… See the full description on the dataset page: https://huggingface.co/datasets/hssling/india-tb-missed-cases-analysis.gem-analysis-unimts
gem-analysis-unimts
UniMTS pretraining datasets for fitness action recognition
数据集信息
来源路径: datasets/unimts
数据大小: 6.6 GB
用途: 健身动作识别模型训练
使用方法
from huggingface_hub import snapshot_download
# 下载数据集
snapshot_download(
repo_id="yonful/gem-analysis-unimts",
repo_type="dataset",
local_dir="./datasets/unimts"
)
或使用项目中的下载脚本:
python scripts/prepare_data.py --dataset unimts
许可证
请参考原始数据源的许可证要求。
routing_analysis-marco_mini_tam_100k_500k_checkpoints
Two marco_mini_base Tam checkpoints
This dataset contains every regular file in the original
routing_analysis/finetuning/phase2_full/checkpoints/marco_mini_base/tam_Taml_100k_full
and tam_Taml_500k_full trees, at the original relative paths. The files are
individual objects, not tar archives. manifests/ records the selected source
snapshot and provenance from the local direct-upload inventory.
tam_Taml_500k_full/checkpoint-4500 is preserved as found and does not contain… See the full description on the dataset page: https://huggingface.co/datasets/divin1234/routing_analysis-marco_mini_tam_100k_500k_checkpoints.NMR-analysisrouting_analysis-marco_nano-hau_Latn-checkpoints
marco_nano_base hau_Latn files
This dataset preserves the original relative paths of every regular file under
the ten marco_nano_base directories whose basename contains the literal
hau_Latn, plus the two matching zero-byte lock files beside them.
The scope intentionally includes hidden work data, preserved backups,
failed_edquot data, and all other regular files found inside those selected
trees. No .gitignore file or checkpoint-name exclusion rule is consulted.
The source… See the full description on the dataset page: https://huggingface.co/datasets/lylybig/routing_analysis-marco_nano-hau_Latn-checkpoints.static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub),
where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks.
OpenAI used the synth-vuln-fixes and fine-tuned
a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo.
More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.trash-in-river-2025
Street Parade 2025 Dataset
Overview
This dataset was collected by SARA, a student initiative at ETH Zurich, to enable open research on trash presence in aquatic environments. It contains images of litter in the Limmat River in Zurich the day after the Street Parade (August 9, 2025). The dataset is intended for training and evaluating trash classification models.
Dataset summary
Collection date: August 9, 2025
Location: Kornhausbrücke, Zurich… See the full description on the dataset page: https://huggingface.co/datasets/SARA-smartphone-assisted-river-analysis/trash-in-river-2025.turkish-sentiment-analysis-dataset
Dataset
This dataset contains positive , negative and notr sentences from several data sources given in the references. In the most sentiment models , there are only two labels; positive and negative. However , user input can be totally notr sentence. For such cases there were no data I could find. Therefore I created this dataset with 3 class. Positive and negative sentences are listed below. Notr examples are extraced from turkish wiki dump. In addition, added some random text… See the full description on the dataset page: https://huggingface.co/datasets/winvoker/turkish-sentiment-analysis-dataset.battery-tests-cfo-analysis
Battery-test CFO analysis
This folder contains the all-packet CFO fingerprints, statistics, and plots for
all ten Morty battery experiments (10% through 100%). Source H5 recordings were
read without modification. The estimator reuses
CSE237D_weyl\pipeline\parallel_h5_cfo_all_packets.py.
Start with summary\all_packets\ALL_PACKETS_RESULTS.md for the consolidated
index. Each experiment has:
all_packets_cfo\packet_cfo_all_fingerprints.csv — every assigned packet.… See the full description on the dataset page: https://huggingface.co/datasets/Morty0311/battery-tests-cfo-analysis.retina-age-analysis
Retina Age Analysis Dataset
Dataset Description
This dataset contains 9,857 retinal fundus images from 5,393 patients for age prediction tasks.
Dataset Summary
Task: Age prediction from retinal fundus images
Images: 9,857 high-quality retinal images
Patients: 5,393 unique patients
Age Range: 5-97 years
Image Format: JPEG
Average Image Size: ~1 MB
Supported Tasks
Regression: Predict continuous age (5-97 years)
Classification: Predict age group (5… See the full description on the dataset page: https://huggingface.co/datasets/ramankamran/retina-age-analysis.Taur_CoT_Analysis_Project___mistralai__Mistral-7B-Instruct-v0.3amazon-reviews-sentiment-analysis
Dataset Card for amazon reviews for sentiment analysis
Dataset Summary
One of the most important problems in e-commerce is the correct calculation of the points given to after-sales products. The solution to this problem is to provide greater customer satisfaction for the e-commerce site, product prominence for sellers, and a seamless shopping experience for buyers. Another problem is the correct ordering of the comments given to the products. The prominence of misleading… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/amazon-reviews-sentiment-analysis.Taur_CoT_Analysis_Project___google__gemini-1.5-pro-001Copper_Google_Trend_Analysissnappfood-sentiment-analysis
