datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CNRelics_StyleTrans
CNRelics_StyleTrans Dataset Description
CNRelics_StyleTrans is a high-quality, task-specific dataset developed for style transfer tasks in the field of computer vision. All data samples are collected from digital cultural relic platforms of authoritative institutions, including the Shanghai Museum and the Palace Museum. The dataset covers a diverse range of Chinese cultural relics, such as landscape paintings, bamboo-wood-ivory-horn artifacts, Dunhuang murals, traditional sculptures… See the full description on the dataset page: https://huggingface.co/datasets/MoonlightAFar/CNRelics_StyleTrans.sft-dataset-from-moonlight-filteredmoonlighter-2-wiki-data
Moonlighter 2 relic prices and weapons dataset
An open, versioned export from www.moonlighter2.wiki, an independent and unofficial Moonlighter 2 wiki published as The Endless Ledger.
This release contains two small, research-friendly tables:
Relic prices: 159 published relics, including 154 rows with sourced prices and 5 rows whose unknown prices are intentionally left blank.
Weapons: 30 published weapons with weapon type, effects, upgrade information, verification state and… See the full description on the dataset page: https://huggingface.co/datasets/zyzw10086/moonlighter-2-wiki-data.nb-asr-moonlight-sentiment
nb-asr-moonlight-sentiment
Curated 20,000 contrastive sentiment pair dataset derived from alexandrainst/sentiment (Trustpilot and consumer feedback).
Purpose
Designed for contrastive sentence embedding models to learn emotional valence and polarity across Scandinavian languages without benchmark contamination (specifically avoiding any NoRec or newspaper review data).
Composition (20,000 Total Pairs)
Norwegian (no): 14,000 pairs (7,000 positive / 7… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nb-asr-moonlight-sentiment.sft-dataset-from-moonlight-noidqwen3.5-0.8b-target-matched-math-240k
qwen3.5-0.8b-target-matched-math-240k
239,467 rows of math-reasoning trajectories regenerated against
Qwen/Qwen3.5-0.8B as the target model. Used to train DFlash
speculative-decoding drafters in
la-draftery.
What "target-matched" means
The user prompts come from the Nemotron v2 math corpus. The assistant
completions in this dataset are the target model's own outputs — each
prompt was sent to Qwen/Qwen3.5-0.8B and its completion was captured.
Drafters trained on… See the full description on the dataset page: https://huggingface.co/datasets/Moonlight556/qwen3.5-0.8b-target-matched-math-240k.sokrates-traceskimi-linear-48b-a3b-target-matched-math-240k
kimi-linear-48b-a3b-target-matched-math-240k
239,467 rows of math-reasoning trajectories regenerated against
moonshotai/Kimi-Linear-48B-A3B-Instruct as the target model. Used to train DFlash
speculative-decoding drafters in
la-draftery.
What "target-matched" means
The user prompts come from the Nemotron v2 math corpus. The assistant
completions in this dataset are the target model's own outputs — each
prompt was sent to moonshotai/Kimi-Linear-48B-A3B-Instruct and its… See the full description on the dataset page: https://huggingface.co/datasets/Moonlight556/kimi-linear-48b-a3b-target-matched-math-240k.sft-dataset-from-moonlightmoonlight-datasetptd-paper-assetsCrossgenre_Roleplaying_in_Simulated_Personas提出一种面向角色扮演语言智能体领域的数据构建方法,并据此构建 CRISP(Cross-genre Role-playing In Simulated Personas) 数据集。该数据集筛选自 多本书籍,包含 多段真实对话,覆盖世界名著、网络文学、轻小说三大类别,支持中、英两种语言。
rewire2-ultrafineweb-moonlightmoonlight_datasetrewire2-ultrafineweb-moonlight2moonlightdatasetdiverse-qa-dclm-moonlight-textmoonlightsport-dataset-2k-llama-moonlightsynth-ultrafineweb-moonlight-diversenemotron-diverse-qa-dclm-moonlightdiverse-synth-moonlight-procMoonlightDadossokrates-datasetsMoonlightShiftPatterns
MoonlightShiftPatterns
tags: predictive, employment, time-series
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'MoonlightShiftPatterns' dataset captures the patterns of individuals engaged in moonlighting jobs across different industries and regions. It includes time-series data of their employment periods and activities. The dataset is structured to facilitate predictive analysis on the duration and frequency of… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MoonlightShiftPatterns.diverse-qa-dclm-moonlight-text-onlyvsorewire2-dclm-moonlightsynth-ultrafineweb-moonlight-diverse-forcedsynth-ultrafineweb-moonlight-diverse-forced-v2
