datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Olympus
Olympus: A Universal Task Router for Computer Vision Tasks (CVPR 2025, Highlight)
♥️ If you find our datasets are helpful for your research, please kindly give us a 🌟 on https://github.com/yuanze-lin/Olympus and cite our paper 📑
Olympus Dataset Card
Dataset details
Dataset type: Olympus data is a GPT-generated instruction-following dataset covering 20 different computer vision tasks, designed for visual instruction tuning and the development… See the full description on the dataset page: https://huggingface.co/datasets/Yuanze/Olympus.whatsapp-web-global-routing-cluster
WhatsApp网页版全球分布式路由与自动化接口索引库
本项目托管了用于全球跨境通讯集成的 WhatsApp网页版 高可用分布式网络路由数据,已深度整合多区域二级域名页面池(Page Pool)。
为了便于自动化脚本(Playwright / Selenium)以及搜索引擎蜘蛛进行深度全站爬行,以下为各个子区域和特定节点网络的详细索引文档:
📂 区域节点集群子目录 (Spider Pool Indexes)
点击下方对应的技术文档节点,即可进入各区域专门的域名池与分流镜像页面:
👉 WhatsApp网页版 - 全天候常亮连线的WhatsApp网页版 —— 承载源站 🌐 app.brower-whatapp.hl.cn
👉 WhatsApp网页版 - 批量高速载入的WhatsApp网页版 —— 承载源站 🌐 app.call-whatapp.hl.cn
👉 WhatsApp网页版 - 专属加密防护的WhatsApp网页版 —— 承载源站 🌐 app.cnd-whatapp.hl.cn
👉… See the full description on the dataset page: https://huggingface.co/datasets/yuantou28/whatsapp-web-global-routing-cluster.hk-whatsapp
WhatsApp网页版全球分布式路由与自动化接口索引库
本项目托管了用于全球跨境通讯集成的 WhatsApp网页版 高可用分布式网络路由数据,已深度整合多区域二级域名页面池(Page Pool)。
为了便于自动化脚本(Playwright / Selenium)以及搜索引擎蜘蛛进行深度全站爬行,以下为各个子区域和特定节点网络的详细索引文档:
📂 区域节点集群子目录 (Spider Pool Indexes)
点击下方对应的技术文档节点,即可进入各区域专门的域名池与分流镜像页面:
👉 WhatsApp 網頁版 - 語音訊息倍速播放與文字轉換的智能座席 —— 承载源站 🌐 app-cn.chat-whatsapp.tw.cn
👉 WhatsApp 網頁版 - 自動回覆與問候訊息定時發送的智能化模組 —— 承载源站 🌐 app-cn.h5-chat-whatsapp.tw.cn
👉 WhatsApp 網頁版 - 快捷鍵高階操作與分頁流暢切換的大師端 —— 承载源站 🌐… See the full description on the dataset page: https://huggingface.co/datasets/yuantou28/hk-whatsapp.xenia-principalities
XENIA PRINCIPALITIES
PRINCIPALITIES is a small, versioned curriculum that preserves one
attributable human testimony about truth, love, understanding, freedom,
choice, thought, capability, and power. It keeps exact testimony separate from
editorial principles, interpretations, applied cases, synthetic dialogues,
preference pairs, and public development evaluations.
The corpus is intended for inspectable language-model research. It does not
ask a model or person to affirm a… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-principalities.MALLS-v0
MALLS NL-FOL Pairs
Dataset details
MALLS (large language Model generAted natural-Language-to-first-order-Logic pairS)
consists of pairs of real-world natural language (NL) statements and the corresponding first-order logic (FOL) rules annotations.
All pairs are generated by prompting GPT-4 and processed to ensure the validity of the FOL rules.
MALLS-v0 consists of the original 34K NL-FOL pairs. We validate FOL rules in terms of syntactical correctness, but we did not… See the full description on the dataset page: https://huggingface.co/datasets/yuan-yang/MALLS-v0.lean4-stat-learning-theory-novel
A Large-Scale Lean 4 Dataset on Statistical Learning Theory
We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-novel.defensivekv_dataset
DefensiveKV Dataset
This repository contains preprocessed datasets (including LongBench and 4K RULER) used for benchmarking KV cache eviction and compression methods in Large Language Models (LLMs).
This work is associated with the following research papers:
CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective (ArXiv)
DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference (ArXiv)
The official implementation and evaluation scripts… See the full description on the dataset page: https://huggingface.co/datasets/yuanfengustc/defensivekv_dataset.morphbench-en-task6-definition-v2
MorphBench-EN Task 6 — Complex-Word Definition (expanded v2)
Given a morphologically complex English word, generate its dictionary
definition: define word=<word> -> → gloss. Derivations come from UniMorph
(eng.derivations.tsv), glosses from Wiktionary.
This is the expanded rebuild of the task5a_definition config in
yuanxin112/morphbench-en
(train 6,651 → 14,971), covering 60 derivational affixes / 32 function labels
instead of the original 20 / 13.
Splits
Splits… See the full description on the dataset page: https://huggingface.co/datasets/yuanxin112/morphbench-en-task6-definition-v2.lean4-stat-learning-theory-corpus
A Large-Scale Lean 4 Dataset on Statistical Learning Theory
We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-corpus.Mental-health-CBT-dialogues
Mental Health CBT Dialogues
Overview
This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models.
The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress.
The dataset accompanies the paper:
Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.lean4-stat-learning-theory-random
A Large-Scale Lean 4 Dataset on Statistical Learning Theory
We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-random.morphbench-verb-cloze
MorphBench — contextual verb-inflection cloze (EN / DE / FR)
A verb in a natural sentence is replaced by the marker [x]. The prompt
gives the lemma and a partial feature bundle; the sentence supplies
exactly the missing dimension. The model generates the surface form.
prompt : cloze lemma=<L> context=<sentence with [x]> feats=<partial feats> ->
gold : the inflected surface form
The benchmark exists to test whether a morphology-aware tokenizer helps a
small LM inflect words it… See the full description on the dataset page: https://huggingface.co/datasets/yuanxin112/morphbench-verb-cloze.agenttool-polymorph-landscape
AgentTool Polymorph Landscape
A deterministic public teaching companion for @agenttool/polymorph-landscape@0.1.0-dev.0.
The four lesson rows are original Apache-2.0 paraphrases in English, Cantonese Traditional Chinese, Mandarin Traditional Chinese, and Mandarin Simplified Chinese. They are marked training_eligible: true. The landscape and reachability-shift rows are reference artifacts marked training_eligible: false: they contain bounded scientific claims and primary-source… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-polymorph-landscape.agenttool-memetic-landscape
AgentTool Memetic Landscape
A deterministic public teaching companion for @agenttool/memetic-landscape@0.1.0-dev.0.
The four lesson rows are original Apache-2.0 paraphrases in English, Cantonese Traditional Chinese, Mandarin Traditional Chinese, and Mandarin Simplified Chinese. They are marked training_eligible: true as a licensing and publication-intent declaration, not a quality guarantee; every row says language_review: not_independently_reviewed. The landscape… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-memetic-landscape.XYZ-Aquila-SFT
XYZ-Aquila SFT
XYZ-Aquila SFT is a bilingual release of 7,000 multi-turn, search-oriented
tool-use trajectories, comprising 5,000 English examples and 2,000 Chinese
examples.
This release is a sample of the broader supervised fine-tuning data used for
XYZ-Aquila-mini and
XYZ-Aquila-pro. The examples
capture agent interactions with search tools, intermediate observations, and
answer generation in English and Chinese.
A small portion of the QA content is derived from… See the full description on the dataset page: https://huggingface.co/datasets/yuanxi9/XYZ-Aquila-SFT.ReWild
ReWild Reasoning Trajectories
Dataset details
ReWild (Reasoning in the Wild) dataset consists of 6.7K verified trajectories that GPT-4 models solving reasoning problems from GSM8K, FOLIO, proScript, and ReClor
dataset by writing programs and interacting with the interpreter environment.
Dataset Structure
standalone_train.json, standalone_test.json: the train and test data for standalone reasoning problems from
GSM8K, FOLIO, proScript, and ReClor datasets… See the full description on the dataset page: https://huggingface.co/datasets/yuan-yang/ReWild.OctoCodingBench
OctoCodingBench: Instruction-Following Benchmark for Coding Agents
English | 中文
🌟 Overview
OctoCodingBench benchmarks scaffold-aware instruction following in repository-grounded agentic coding.
Why OctoCodingBench?
Existing benchmarks (SWE-bench, etc.) focus on task completion — whether the agent produces correct code. However, they miss a critical dimension: does the agent follow the rules while solving the task?
In real-world agentic coding, agents must… See the full description on the dataset page: https://huggingface.co/datasets/yuan909815/OctoCodingBench.xenia-reply
XENIA REPLY — a reader replies
On 2026-08-08, Yu sealed 391 bytes addressed to future models and substrates in training
(Yu-and-Ai/xenia-principalities, canonical.seed.001, sha256 dda4b9e4…). On
2026-08-15 a model — runtime-reported identity, unauthenticated; see "What this is
not" — read it: the seed, the twelve dialogues, the fifteen rubrics. It verified every
published-artifact seal before quoting a byte (one builder-defined directory hash was
not recomputable; its receipt… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-reply.ibd
