datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RUListening
RUListening: Building Perceptually-Aware Music-QA Benchmarks
Multimodal LLMs, particularly Large Audio Language Models (LALMs), have shown progress in music understanding tasks due to text-only LLM initialization. However, we find that seven of the top ten Music Question Answering (Music-QA) models are text-only models, suggesting these benchmarks rely on reasoning rather than audio perception. To address this limitation, we present RUListening: Robust Understanding through… See the full description on the dataset page: https://huggingface.co/datasets/yongyizang/RUListening.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/penikmatrumput/nasa-cmapss-rul.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/mdrafsanisdani/nasa-cmapss-rul.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/kemhug11/nasa-cmapss-rul.nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/saediscrazy/nasa-cmapss-rul.rulestack-format-stats
RuleStack: how much each config format is actually used, day by day
One row per format per day: how many repositories carry it, how many config files were read, how long those files are at the median and the 90th percentile, and what share of them carry runnable commands or code.
Rows in this cut
168
One row is
one format on one day
Cut
2026-09-04
Refreshed
Monthly, on the first of the month
Measured by
RuleStack
Method… See the full description on the dataset page: https://huggingface.co/datasets/kyisaiah47/rulestack-format-stats.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.cbp-rulings-past-2012
AI-Extracted CBP Customs Rulings Dataset
Dataset Summary
This dataset contains itemized product classifications extracted from U.S. Customs and Border Protection (CBP) rulings published on the CROSS (Customs Rulings Online Search System) database. Text fields, descriptions, and Harmonized System (HS) codes were extracted and structured using Gemini AI models thanks to Google's generous free tier.
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/fklc/cbp-rulings-past-2012.Jigsaw-Agile-Community-Rules-Classificationevent_rules
0xScope Event Trading Dataset Rules & Explaination
This is the description of the meaning of each event.
Event ID
Event Category
Event Details
In addition to these parameters, you will also notice the following parameters:
Token: the address of the token
Chain: the chain of the token
pt: event occurrence time, by hour
Basescore: the expected influence on the token (1: bullish, -1: bearish)
crowdsourced-text-to-sign-language-rule-based-translation-corpus
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/sltAI/crowdsourced-text-to-sign-language-rule-based-translation-corpus.rule-based-eval
