datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llmtcl
⚡ LitGPT
20+ high-performance LLMs with recipes to pretrain, finetune, and deploy at scale.
✅ From scratch implementations ✅ No abstractions ✅ Beginner friendly
✅ Flash attention ✅ FSDP ✅ LoRA, QLoRA, Adapter
✅ Reduce GPU memory (fp4/8/16/32) ✅ 1-1000+ GPUs/TPUs ✅ 20+ LLMs
Quick start •
Models •
Finetune •
Deploy •
All workflows •
Features •
Recipes (YAML) •
Lightning AI •
Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/Maple222/llmtcl.Time-300B
Dataset Card for Time-300B
This repository contains the Time-300B dataset of the paper Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts.
For details on how to use this dataset, please visit our GitHub page.
OpenstoryPlusPlus
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks.
related resorcce
paper: https://arxiv.org/abs/2408.03695
code: https://github.com/YeLuoSuiYou/openstorypp
News
2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.UniREditBench-ResultsUniREdit-Data-100KUniREditBench: A Unified Reasoning-based Image Editing Benchmark
MAPLE
MAPLE: Multi-Aspect Full-Paper Scientific Retrieval Benchmark
MAPLE is an expert-validated benchmark for multi-aspect full-paper retrieval. It contains 2,095 fine-grained queries derived from 210 recent machine learning papers, together with a retrieval corpus of 73,973 candidate papers. Each target paper is paired with multiple queries grounded in textual or multimodal evidence and covering different aspects of the paper, including motivation, method, and experimental findings.… See the full description on the dataset page: https://huggingface.co/datasets/kai-02/MAPLE.MME-RealWorld
2024.11.14 🌟 MME-RealWorld now has a lite version (50 samples per task) for inference acceleration, which is also supported by VLMEvalKit and Lmms-eval.
2024.10.27 🌟 LLaVA-OV currently ranks first on our leaderboard, but its overall accuracy remains below 55%, see our leaderboard for the detail.
2024.09.03 🌟 MME-RealWorld is now supported in the VLMEvalKit and Lmms-eval repository, enabling one-click evaluation—give it a try!"
2024.08.20 🌟 We are very proud to launch MME-RealWorld, which… See the full description on the dataset page: https://huggingface.co/datasets/Mapleyuchen/MME-RealWorld.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Tudor Iustin.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.maple-preview-cuda-benchmarks
Maple Preview TQ2_0 CUDA Benchmarks
Reproducibility data for the TQ2_0 CUDA patches in
PascalAI2024/maple-preview-windows-cuda.
This repository contains benchmark data, patch files, hashes, and raw validation
evidence. It does not duplicate the Maple model weights.
Result
The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ
prompt-processing gain:
Variant
pp512 mean
pp512 median
tg128 mean
tg128 median
Correctness
MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.UniREditBenchUniREditBench: A Unified Reasoning-based Image Editing Benchmark
maple
MAPLE (Bill Summarization, Tagging, Explanation)
In this project, we generate summaries and category tags for of Massachusetts bills for MAPLE Platform. The goal is to simplify the legal language and content to make it comprehensible for a broader audience (9th-grade comprehension level) by exploring different ML and LLM services.
This repository contains a pipeline from taking bills from Massachusetts legislature, generating summaries and category tags leveraging different the… See the full description on the dataset page: https://huggingface.co/datasets/ayang903/maple.maple728-time_300B
Dataset Card for Time-300B
This repository contains the Time-300B dataset of the paper Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts.
For details on how to use this dataset, please visit our GitHub page.
MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Fabric AI.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces, dashboards… See the full description on the dataset page: https://huggingface.co/datasets/FabricAI/maple.MapleStory_Monsters_Datasetmaplestory_captchaA huge collection of English MapleStory's captcha text in jpg that I have collected over the years. ENJOY!!
It us used by pre-Big Bang MapleStory, throughout the game from Lie-Detector (anti-macro item), logins, to NPC conversations.
Up till version 190 when they have switched using Runes (Up, Down, Left, Right arrow keys) for most of the time for detection of macros and bots.
These images are not labelled, I'm releasing this for anyone that wants the dataset to be able to train a model… See the full description on the dataset page: https://huggingface.co/datasets/lastbattle/maplestory_captcha.mapleautofarm
MapleAutoFarm · 冒险岛自动打怪 Python 版
仅用于单机 / 离线 / 个人学习,不用于联网游戏。
Python 实现的冒险岛自动巡逻打怪工具,带 Tkinter 可视化面板,预留 OpenCV 视觉找怪能力。
功能
方向键移动、左右巡逻
跳跃键(默认 A)、攻击键(默认 D)
基础自动巡逻打怪
视觉找怪骨架(OpenCV 模板匹配)
可视化面板 + 日志 / 状态 / 循环次数
全局快捷键启动 / 停止 / 退出
宠物自动药水交给游戏内宠物,无需脚本处理
目录结构
MapleAutoFarm\
├─ main.py 主程序 + 可视化面板
├─ bot.py 自动打怪状态机
├─ vision.py OpenCV 图像识别模块
├─ config.json 配置文件(首次保存后生成)
├─ requirements.txt 依赖
├─ install.bat 一键装环境… See the full description on the dataset page: https://huggingface.co/datasets/shyanchen/mapleautofarm.maplestory_characters_hdmarin-starcoderdata_mapleMAPLE-bench
MAPLE Benchmark Test Splits
This repository contains the test splits for the MAPLE benchmark introduced in the paper MAPLE: Modality-Aware Post-training and Learning Ecosystem (https://arxiv.org/pdf/2602.11596). The benchmark is designed for modality-aware multimodal evaluation under different required-signal settings, where each sample is annotated with the minimal modality subset needed to solve the task.
Dataset Overview
MAPLE-bench evaluates multimodal reasoning… See the full description on the dataset page: https://huggingface.co/datasets/lihVerma/MAPLE-bench.MAPLE-Lua-Corpusmaple-analyst-cap-sft-data
maple-analyst-cap-sft-data
Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B
ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato
TC (ThinkingCap).
Composición
Fuente
Filas
thinkingcap (curriculum, trazas bigbang)
1,782
openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0)
702
bigbang_mmlu
508
bigbang_bbh
441
hermes_function_calling
360
aya_dataset
342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.naver-economy-news2stockmaplestory-worlds-creator-qa
MapleStory Worlds Creator QA
Synthetic question-answer dataset built from the official
MapleStory Worlds Creator Center
documentation. Questions are generated to be self-contained and grounded in the
source docs; answers avoid source/meta references so they read like an expert
explanation. Some QA pairs are composed from multiple related documents
(see combo_sources).
Parallel Korean/English. Intended for instruction tuning, QA, and retrieval.
Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.maple-personas
MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI
A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework.
Dataset Description
This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.CoastAdapt-KB
CoastAdapt-KB Zero-Shot Hierarchical Events Dataset
Dataset Summary
This dataset is prepared from consolidated climate change solution extraction results. It is designed for zero-shot hierarchical multi-label text classification over climate adaptation and mitigation event records.
Each example contains a natural-language input text plus one or more hierarchical label paths. The labels organize climate-related solution details into a taxonomy with phase, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/CoastAdapt-KB.maplestory-resource-index
MapleStory Resource Index
A structured, searchable, and deduplicated metadata index for useful MapleStory resources.
The dataset covers six active series:
MapleStory
MapleStory Classic
MapleStory M
MapleStory Worlds
MapleStory N
MapleStory Idle
Project website
This dataset is maintained by MPStorys, a MapleStory resource discovery platform.
Dataset contents
The current export contains 93 resource records. Fields may include:
Resource ID, name… See the full description on the dataset page: https://huggingface.co/datasets/mpstorys/maplestory-resource-index.pickplacenewmaplestory-worlds-creator-code-instruct
MapleStory Worlds Creator Code (mlua)
Instruction-style code dataset for mlua, the scripting language of
MapleStory Worlds. Built from the
official Creator Center example code: each example is grounded in its source
document and paired with a natural-language task, reasoning, a self-contained
explanation, and commented mlua code. Intended to teach LLMs to write mlua game
scripts.
The example code is preserved from the official source (a code-preservation check
rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.maple-umi-data
