datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UniREdit-Data-100KUniREditBench: A Unified Reasoning-based Image Editing Benchmark
maple-preview-cuda-benchmarks
Maple Preview TQ2_0 CUDA Benchmarks
Reproducibility data for the TQ2_0 CUDA patches in
PascalAI2024/maple-preview-windows-cuda.
This repository contains benchmark data, patch files, hashes, and raw validation
evidence. It does not duplicate the Maple model weights.
Result
The fresh local A/B/B/A validation on an RTX 4080 SUPER reproduced the fused-MMQ
prompt-processing gain:
Variant
pp512 mean
pp512 median
tg128 mean
tg128 median
Correctness
MMQ enabled… See the full description on the dataset page: https://huggingface.co/datasets/x0me/maple-preview-cuda-benchmarks.MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.maple-analyst-cap-sft-data
maple-analyst-cap-sft-data
Dataset de SFT para fine-tune de maple-analyst-cap-bf16 (Qwen3.5-MoE 20.2B
ternario). 4,956 trazas de razonamiento (pseudothinking + answer) en formato
TC (ThinkingCap).
Composición
Fuente
Filas
thinkingcap (curriculum, trazas bigbang)
1,782
openmle-condensed (FrontisAI OpenMLE-SFT-Traces, condensadas con distiller LFM2.5-2.6B q8_0)
702
bigbang_mmlu
508
bigbang_bbh
441
hermes_function_calling
360
aya_dataset
342… See the full description on the dataset page: https://huggingface.co/datasets/Davd-b01/maple-analyst-cap-sft-data.maple-personas
MAPLE-Personas: A Benchmark for Evaluating Personalized Conversational AI
A dataset for evaluating how well conversational AI systems learn and apply user preferences from natural dialogue. This benchmark accompanies the MAPLE (Memory-Adaptive Personalized LEarning) framework.
Dataset Description
This dataset tests an AI assistant's ability to implicitly learn user traits from conversation context and apply that knowledge to personalize responses to open-ended… See the full description on the dataset page: https://huggingface.co/datasets/prdeepakbabu/maple-personas.CoastAdapt-KB
CoastAdapt-KB Zero-Shot Hierarchical Events Dataset
Dataset Summary
This dataset is prepared from consolidated climate change solution extraction results. It is designed for zero-shot hierarchical multi-label text classification over climate adaptation and mitigation event records.
Each example contains a natural-language input text plus one or more hierarchical label paths. The labels organize climate-related solution details into a taxonomy with phase, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/CoastAdapt-KB.maplestory-resource-index
MapleStory Resource Index
A structured, searchable, and deduplicated metadata index for useful MapleStory resources.
The dataset covers six active series:
MapleStory
MapleStory Classic
MapleStory M
MapleStory Worlds
MapleStory N
MapleStory Idle
Project website
This dataset is maintained by MPStorys, a MapleStory resource discovery platform.
Dataset contents
The current export contains 93 resource records. Fields may include:
Resource ID, name… See the full description on the dataset page: https://huggingface.co/datasets/mpstorys/maplestory-resource-index.sn38r6-u70-subsn38r5-u70-subsn38r6-u170-submaplept-reasoning-corpusmaplept2-reasoning-corpusmaplept2-coder-corpusmsxgpt-dataset
msxgpt
Description
This dataset, "msxgpt," is designed for training the GPT-3.5-turbo/ GPT-4 based language model for a task. The data consists of JSON lines, each representing an individual example for the model.
The dataset has been created with an emphasis on encoding, which is pivotal to the functionality of Memory Features, Security, and API Endpoints. It is designed to process and store documents from various data sources continuously, using incoming webhooks to the… See the full description on the dataset page: https://huggingface.co/datasets/MapleSage/msxgpt-dataset.sn38r4-u70-subsn38r3-u170-suberling
