datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenstoryPlusPlus
Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
We introduce OpenStory++, a large-scale open-domain dataset contains focusing on enabling MLLMs to perform storytelling generation tasks.
related resorcce
paper: https://arxiv.org/abs/2408.03695
code: https://github.com/YeLuoSuiYou/openstorypp
News
2024/7/31 We have reorganized and distributed the high-quality subset and released most of the story data collected… See the full description on the dataset page: https://huggingface.co/datasets/MAPLE-WestLake-AIGC/OpenstoryPlusPlus.maple
Overview
Maple is an open-source full-stack code dataset developed and released by Tudor Iustin.
It is designed to support code generation, web development, supervised fine-tuning, instruction tuning, post-training, dataset research, and evaluation workflows for code-capable AI systems.
Maple contains 16,000 full-stack code samples totaling approximately 102 million tokens. It focuses on realistic software-building tasks, including web applications, product interfaces… See the full description on the dataset page: https://huggingface.co/datasets/tudor-iustin22/maple.maplestory_characters_hdmaplestory-worlds-creator-qa
MapleStory Worlds Creator QA
Synthetic question-answer dataset built from the official
MapleStory Worlds Creator Center
documentation. Questions are generated to be self-contained and grounded in the
source docs; answers avoid source/meta references so they read like an expert
explanation. Some QA pairs are composed from multiple related documents
(see combo_sources).
Parallel Korean/English. Intended for instruction tuning, QA, and retrieval.
Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.naver-economy-news2stockmaplestory-worlds-creator-code-instruct
MapleStory Worlds Creator Code (mlua)
Instruction-style code dataset for mlua, the scripting language of
MapleStory Worlds. Built from the
official Creator Center example code: each example is grounded in its source
document and paired with a natural-language task, reasoning, a self-contained
explanation, and commented mlua code. Intended to teach LLMs to write mlua game
scripts.
The example code is preserved from the official source (a code-preservation check
rejects any record… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-code-instruct.maplestory-worlds-creator-docs
MapleStory Worlds Creator Center Documentation
A curated dataset built from the official documentation of the
MapleStory Worlds Creator Center.
It is a parallel Korean/English documentation corpus intended for RAG, search,
embeddings, and domain language-model training.
The dataset covers all three Creator Center content types — guide documents
(doc), API Reference (api), and resources (res).
Composition
Document counts by type and language:
type
Description… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-docs.DynaSolidGeo-SamplePaper: https://arxiv.org/abs/2510.22340
Github Repo: https://github.com/ChangtiWu/DynaSolidGeo
In the "Appendix.E: DynaSolidGeo as a Training Dataset" in our paper, we sample K = 10 batches of instances using random seeds from 0 to 9, resulting in a total of 5,030 samples.
These samples are divided into a training set (3,627 samples), a validation set (403 samples), and a test set (1,000 samples).
yamlmmlushort_selling
Short Selling
Data Notice: This dataset provides academic research access with a 6-month data lag.
For real-time data access, please visit sov.ai to subscribe.
For market insights and additional subscription options, check out our newsletter at blog.sov.ai.
from datasets import load_dataset
df_over_shorted = load_dataset("sovai/short_selling", split="train").to_pandas().set_index(["ticker","date"])
Data is updated weekly as data arrives after market close US-EST time.
Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/MapleLeavesKrish/short_selling.alpaca_npc_v2_data.jsonbuzz_sources_410_mapletestDataAiNickmaplere0
