datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
csc_clean_wang271k
csc_eval_public
一、测评数据说明
1.1 数据清洗
余-馀: 替换为馀-余
other - 馀: 替换为余
覆-复: 替换为复-覆
other-覆: # 答疆/回覆/反覆
# 覆审
他-她:不纠
她-他:不纠
人名不纠: 识别人名并丢弃
的得地: 建议丢弃(标注得不准)
# # 的 - 地
# # 的 - 得
# # 它 - 他
# # 哪 - 那
# # 改-大小改: 余-馀 覆-复 借-藉 功-工 琅-瑯 震-振 百-白 也-叶 经-禁(经不起-禁不起)
# # 部分不变(人名): 小-晓 一-逸 佳-家 得-地(马哈得) 红-虹 民-明
# # 匹配上但是不改的: 惟-唯 象-像 查-察 立-利 止-只 建-健 他-它 地-的 定-订 带-戴 力-利 成-城 点-店
# # 匹配上但是不改的: 作-做 得-的 场-厂 身-生 有-由 种-重 理-里
# # 空白没匹配上: 今-在 年-今 前-目 当-在 目-在 者-是
# # 外国人名等:其-齐 课-科 博-波… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_clean_wang271k.aime25
AIME 25
American Invitational Mathematics Examination (AIME) 2025
Citation
If you use the AIME25 dataset in your research, please consider citing it as follows:
@misc{aime25,
title={American Invitational Mathematics Examination (AIME) 2025},
author={Zhang, Yifan and Math-AI, Team},
year={2025},
}
cl-macros-thinking-clean
cl-macros-thinking
Chat-format Common Lisp macro dataset enriched with <think>...</think>
reasoning traces. Each row is a single training example:
{
"messages": [
{"role": "system", "content": "You are an expert Common Lisp macro programmer..."},
{"role": "user", "content": "Write a Common Lisp macro to handle ..."},
{"role": "assistant", "content": "<think>...reasoning...</think>\n\n(defmacro ...)"}
]
}
Splits
split
rows
train
1607… See the full description on the dataset page: https://huggingface.co/datasets/j14i/cl-macros-thinking-clean.ayurveda-text-based-qandacsc_public_de3
csc_public_de3数据集
数据来源
1.由人民日报/学习强国/chinese-poetry等高质量数据人工生成;
2.来自人民日报高质量语料;
3.来自学习强国网站的高质量语料;
4.源数据为qwen生成的好词好句;
5.古诗词chinese-poetry; 文言文garychowcmu/daizhigev20;
数据简介
该数据主要为'的地得'纠错;
其中训练数据130753条, 验证数据5545条, 测试数据5545条;
句子平均长度为36, 最长句子长度为414, 最短为5, 95%的为89, 75%的为46, 60%的为34;
每个句子中字的平均错误数为2;
数据详情
################################################################################################################################
train.json
130753… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_public_de3.Macromile-ODbL
MacroMile — Open Food Facts derived data
Contains information from Open Food Facts,
made available under the Open Database License (ODbL) v1.0.
This is the ODbL share-alike extract for the MacroMile food catalog: every row
in the catalog that derives from Open Food Facts, published so that anyone
receiving the app also receives the adapted database it was built from.
What this is
Rows
621,731
Generated
2026-08-22T03:43:35+00:00
File… See the full description on the dataset page: https://huggingface.co/datasets/AISSLLC/Macromile-ODbL.cl-macros-thinking
cl-macros-thinking
Chat-format Common Lisp macro dataset enriched with <think>...</think>
reasoning traces. Each row is a single training example:
{
"messages": [
{"role": "system", "content": "You are an expert Common Lisp macro programmer..."},
{"role": "user", "content": "Write a Common Lisp macro to handle ..."},
{"role": "assistant", "content": "<think>...reasoning...</think>\n\n(defmacro ...)"}
]
}
Splits
split
rows
train
1828… See the full description on the dataset page: https://huggingface.co/datasets/j14i/cl-macros-thinking.adaption-marketintel-macro-reasoning-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-MarketIntel-Macro-Reasoning-v1
Instruction-tuned financial telemetry and macro news dataset engineered for multi-variable market risk analysis, sector contagion modeling, and automated portfolio rebalancing strategies. Synthesizes raw market headlines and regulatory telemetry into structured, step-by-step quantitative reasoning traces and actionable managerial hedge plans.… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-marketintel-macro-reasoning-v1.ratishsp__macro__1646998904
GEM Submission
Submission name: Macro
cl-macros-creative
cl-macros-creative
Hand-curated and LLM-brainstormed Common Lisp macro examples beyond what
j14i/cl-ds covers from
established libraries. Built as Phase-2 exploration corpus for
j14i/cl-macro-27b-lora.
Schema
Mirrors j14i/cl-ds:
field
meaning
instruction
natural-language description of the macro
input
sample call form
output
the reference (defmacro ...) source
macroexpand
exact result of (macroexpand-1 input) under that defmacro
category
control-flow /… See the full description on the dataset page: https://huggingface.co/datasets/j14i/cl-macros-creative.inat-2017-subset
iNaturalist 2017 Supercategory Subset
This dataset is a sampled subset of the iNaturalist 2017 Challenge dataset, specifically processed for efficient object detection training across major biological supercategories.
Dataset Summary
Source: iNaturalist 2017 (Competition Version)
Task: Object Detection
Classes: 9 (Collapsed from thousands of species into biological supercategories)
Training Samples: 1,000 images per supercategory (~9,000 total)
Validation Samples: 200… See the full description on the dataset page: https://huggingface.co/datasets/MacroSony/inat-2017-subset.MWP-InstructQwen3.5-reasoning-700x
Dataset Card (Qwen3.5-reasoning-700x)
Dataset Summary
Qwen3.5-reasoning-700x is a high-quality distilled dataset.
This dataset uses the high-quality instructions constructed by Alibaba-Superior-Reasoning-Stage2 as the seed question set. By calling the latest Qwen3.5-27B full-parameter model on the Alibaba Cloud DashScope platform as the teacher model, it generates high-quality responses featuring long-text reasoning processes (Chain-of-Thought). It covers several major… See the full description on the dataset page: https://huggingface.co/datasets/Macropodidae/Qwen3.5-reasoning-700x.norma
Norma Syllabarum Graecarum - A Benchmark for grc Syllabification and Vowel Length Annotation
We introduce Norma as a common benchmark for the evaluation and comparison of NLP tools concerning markup of two tasks for Ancient Greek (grc): (1) vowel length of dichronic vowels (alpha, iota, ypsilon) in open syllables (where they impact syllable weight) and (2) syllabification, both boundaries and weight. This means that the benchmark also indirectly tests handling of sandhi… See the full description on the dataset page: https://huggingface.co/datasets/Macronizer/norma.Comicify-Heading-GenerationOpus-4.6-Reasoning-3000x-filteredFiltered from: https://huggingface.co/datasets/crownelius/Opus-4.6-Reasoning-3000x
The original dataset has 979 refusals, I removed these in this version.
claude-4.5-opus-high-reasoning-250xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs.
Stats
Costs: $ 52.3 (USD)
Total tokens (input + output): 2.13 M
mmlu-high-school-macroeconomicsMacroeconomicsGraphsMacroEcon
