datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KMMLU
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU.KMMLU-HARD
KMMLU (Korean-MMLU)
We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM.
Unlike previous Korean benchmarks that are translated from existing English benchmarks, KMMLU is collected from original Korean exams, capturing linguistic and cultural aspects of the Korean language.
We test 26 publically available and proprietary LLMs, identifying significant room for improvement.
The best publicly… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KMMLU-HARD.FINDER_API_KEY_AI_SEARCH_2023
FINDER_API_KEY_AI_SEARCH_2023
tags: data collection, machine learning, API performance
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FINDER_API_KEY_AI_SEARCH_2023' dataset is designed to collect and analyze data from various AI search engines and their associated API performance metrics. The dataset focuses on the effectiveness of API key-based access in enhancing the search capabilities of AI systems and includes a… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FINDER_API_KEY_AI_SEARCH_2023.HRM8K
| 📖 Paper | 📝 Blog | 🖥️ Code(Coming soon!) |
HRM8K
We introduce HAE-RAE Math 8K (HRM8K), a bilingual math reasoning benchmark for Korean and English.
HRM8K comprises 8,011 instances for evaluation, sourced through a combination of translations from established English benchmarks (e.g., GSM8K, MATH, OmniMath, MMMLU) and original problems curated from existing Korean math exams.
Benchmark Overview
The HRM8K benchmark consists of two subsets:
Korean School Math (KSM):… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HRM8K.tigre-hubert-datahub-trending-models-2026-03-06hubble-8b-unlearning-resultsKoSimpleEvalhubbench
HubBench 1.4.0
One Blobfish-authored, oracle-proven benchmark family per Harbor Hub professional-domain cluster. Every task is an employee decision worked over a dependent chain of evidence — never a lookup — against mock stateful tools over an isolated SQLite world. The agent reaches the world only through its public surfaces (MCP over streamable HTTP, a terminal tool CLI, a REST API, and a web console); a deterministic verifier (HubScore) grades the finished world from… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/hubbench.TripAdvisorReviewSentiment
TripAdvisorReviewSentiment
tags: sentiment analysis, text mining, opinion mining
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'TripAdvisorReviewSentiment' dataset contains user reviews from TripAdvisor, a travel website. Each review is assessed for sentiment, with the label indicating whether the review expresses a positive, negative, or neutral sentiment. The dataset is suitable for sentiment analysis and text mining tasks… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/TripAdvisorReviewSentiment.NYC_Subway_Locations
NYC_Subway_Locations
tags: GIS, machine learning, regression
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'NYC_Subway_Locations' dataset contains geographical information about subway stations in New York City. Each record provides the station code, station name, and its geographical coordinates (latitude and longitude). This dataset is suitable for geospatial analysis and could be used in conjunction with machine learning… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/NYC_Subway_Locations.HAE_RAE_BENCH_1.0The HAE_RAE_BENCH 1.0 is the original implementation of the dataset froom the paper: HAE-RAE BENCH paper.
The benchmark is a collection of 1,538 instances across 6 tasks: standard_nomenclature, loan_word, rare_word, general_knowledge, history and reading comprehension.
To replicate the studies from the paper, see below.
Dataset Overview
Task
Instances
Version
Explanation
standard_nomenclature
153
v1.0
Multiple-choice questions about Korean standard nomenclatures from… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/HAE_RAE_BENCH_1.0.MobilePlanAssistant
MobilePlanAssistant
tags: dialogue, chatbot, mobile-plans
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'MobilePlanAssistant' dataset comprises simulated dialogues between a user seeking to find the best mobile plan and a chatbot tasked with assisting in this endeavor. Each dialogue captures the user's requests and preferences, the bot's responses, and the overall outcome of the conversation. Success is determined by the… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MobilePlanAssistant.K2-EvalResearch Paper coming soon!
K2EvalK^{2} EvalK2Eval
K2EvalK^{2} EvalK2Eval is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion.
Benchmark Overview
The design principle behind K2EvalK^{2} EvalK2Eval centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/K2-Eval.kin_20250421KUDGEOfficial data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot DoTLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times)
Dataset Description
At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point.
Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/KUDGE.PokemonBattlePredictions
PokemonBattlePredictions
tags: Classification, Gameplay, AI-powered
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'PokemonBattlePredictions' dataset is a collection of textual descriptions of Pokémon battles, specifically designed to train machine learning models in predicting the outcome of battles between Pokémon teams. The dataset includes information about each Pokémon's stats, abilities, moves, and any relevant items.… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/PokemonBattlePredictions.SMSSpamSentiment
SMSSpamSentiment
tags: sentiment analysis, spam detection, machine learning
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SMSSpamSentiment' dataset is a collection of SMS messages labeled with the sentiment of the content, where each message is either classified as 'spam' or 'ham' (non-spam). The dataset includes columns for the SMS message text and a label indicating the sentiment. For this dataset, labels are binary… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SMSSpamSentiment.FrogSoundRecognition
FrogSoundRecognition
tags: audio analysis, machine learning, amphibian communication
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'FrogSoundRecognition' dataset is designed for the purpose of training machine learning models to recognize and classify various frog calls. It includes audio recordings of different frog species along with labels identifying the species and recording location. The dataset is useful for… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/FrogSoundRecognition.SELFormerMM
SELFormerMM
Multimodal molecular representation data for SELFormerMM — an extension of
SELFormer that aligns four complementary views of a
molecule in a shared embedding space: SELFIES sequences, 3D/structural graphs,
textual descriptions, and knowledge-graph context, trained with multimodal supervised
contrastive learning on ~2.9M molecules.
Paper: SELFormerMM: multimodal molecular representation learning via SELFIES, structure, text, and
knowledge graph integration… See the full description on the dataset page: https://huggingface.co/datasets/HUBioDataLab/SELFormerMM.Ko-PIQA
Ko-PIQA: Korean Physical Commonsense Reasoning Dataset
📖 Dataset Overview
Ko-PIQA is a Korean Physical Commonsense Reasoning dataset designed to complement English-centric benchmarks like PIQA and to include culturally-grounded physical reasoning questions.
Total items: 441
Culturally-grounded items: 87 (19.7%)(e.g., kimchi storage, hanbok care, ondol heating)
Format: PIQA-style binary choice (solution0 / solution1)
Goal: Evaluate Korean LLM physical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/HAERAE-HUB/Ko-PIQA.visaadvisor-country-hub-v0-1-0
VisaAdvisor Country Hub v0.1.0
العربية
This repository is a discovery mirror of the VisaAdvisor Country Hub v0.1.0 public foundation pre-release. Read the bilingual Country Hub landing page. The canonical, citable release is archived on Zenodo under DOI 10.5281/zenodo.21858354. The full source, schema, governance documents and reproducible build scripts are available in the public GitHub repository.
The release preserves the distinction between official claims, editorial… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/visaadvisor-country-hub-v0-1-0.SuicidePrediction
SuicidePrediction
tags: mental health, prediction, prevention
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'SuicidePrediction' dataset is designed to assist ML practitioners in building predictive models for suicide risk assessment. Each entry includes a brief text description of an individual's mental health status, followed by labels that indicate whether the individual is at high, moderate, or low risk of suicide based… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/SuicidePrediction.BanglaGEC
BanglaGEC: A Large-Scale Parallel Corpus for Bangla Grammatical Error Correction
BanglaGEC is a large-scale parallel corpus of 7,074,425 (~7.1M) sentence pairs for Bangla (Bengali) Grammatical Error Correction (GEC). Each pair maps a grammatically erroneous Bangla sentence to its grammatically correct counterpart, along with the error type, making it directly usable for training and evaluating sequence-to-sequence models, transformers, and large language models on the Bangla… See the full description on the dataset page: https://huggingface.co/datasets/nahid-hub/BanglaGEC.MovieGenreClassification
MovieGenreClassification
tags: Genre, ML, Classification
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'MovieGenreClassification' dataset consists of movie descriptions and their corresponding genres. Each entry in the dataset includes the movie's plot summary and a genre label, reflecting the main category of the movie. This dataset is tailored for machine learning models aiming to classify movies into predefined genres.… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MovieGenreClassification.pdbbind_fullPedale-FrenchTextCorpus
Pedale-FrenchTextCorpus
tags: classification, linguistics, French
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'Pedale-FrenchTextCorpus' is a collection of French text excerpts from articles, blog posts, and forum discussions related to bicycle pedals ('pedales'). The dataset is intended for machine learning practitioners who are working on text classification tasks that focus on French-language content, with an emphasis on… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/Pedale-FrenchTextCorpus.PenTestingScenarioSimulation
PenTestingScenarioSimulation
tags: Clustering, Threat Simulation, Attack Scenarios
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'PenTestingScenarioSimulation' dataset is designed for machine learning practitioners interested in simulating various penetration testing (pen-testing) scenarios. Each record in the dataset represents a simulated attack scenario and includes detailed information about the nature of the threat, the… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/PenTestingScenarioSimulation.reboot-hub-dji-drone-specs-used-price-index
Reboot Hub DJI Drone Specs and Used Price Index
This dataset provides a complete public Q3 2026 table of 43 aircraft model-level listed-price aggregates summarizing 251 published Reboot Hub catalog configurations, plus four illustrative records with richer repair-risk context.
Version: 0.2.0
Zenodo concept DOI: https://doi.org/10.5281/zenodo.21246532
Exact v0.2.0 DOI: https://doi.org/10.5281/zenodo.21387578
Primary source page:
https://reboot-hub.com/pages/reboot-hub-data… See the full description on the dataset page: https://huggingface.co/datasets/Thomas0229/reboot-hub-dji-drone-specs-used-price-index.esol
