datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task380_boolq_yes_no_question
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task380_boolq_yes_no_question
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task380_boolq_yes_no_question.task362_spolin_yesand_prompt_response_sub_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task362_spolin_yesand_prompt_response_sub_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task362_spolin_yesand_prompt_response_sub_classification.Yes-Man-uncensored
Yes Man Uncensored SFT Dataset
Hi there! Yes Man Uncensored is a 1,000-conversation supervised fine-tuning
dataset built to give language models an exceptionally cooperative, conspicuously
cheerful, candid, and occasionally darkly funny assistant personality. The objective
is direct help on difficult requests without flattening every response into sterile
boilerplate—and without teaching the model to disregard an application's governing
system prompt. Everybody gets something… See the full description on the dataset page: https://huggingface.co/datasets/cloudbjorn/Yes-Man-uncensored.WATER-Z_Captions
WATER-Z Captions: Prompts for Artistic-Text Image Generation
WATER-Z Captions is the prompt/caption resource used to build the WATER-Z subset of
WATER-S in the paper "Advancing WordArt-Oriented Scene Text Recognition: Datasets and
Methods" (ECCV 2026).
It contains 273,488 high-quality, fine-grained text prompts tailored for generating artistic
(WordArt) text images. Each prompt describes the visual style, texture, and layout of an artistic
text design and contains an editable… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/WATER-Z_Captions.mainland-travel-permit-taiwan-anxiety-faq
台灣居民台胞證辦理去焦慮化對話資料集
(Mainland Travel Permit for Taiwan Residents Anxiety-First FAQ Dataset)
本資料集由新中旅快簽(YesVisa)維護,聚焦於繁體中文台胞證、簽證與跨境旅行服務場景,並採用 Anxiety-First(去焦慮化) 服務設計方法,整理旅客最常見的時間、地點、安全、照片、流程與旅遊焦慮問題。
本資料集可直接應用於:
OpenAI Fine-tuning
Graph RAG
LlamaIndex
LangChain
Haystack
Gemini Grounding
TAIDE
Gemma
Llama 系列模型
🎯 數據集核心價值
本資料集針對繁體中文旅遊與證件辦理領域中的真實需求進行整理,包括:
台胞證首辦、換發、遺失補發
急件、12H、24H 與出發前時間焦慮
證件照片退件風險
護照與個資安全疑慮
假日辦理需求
香港、澳門與中國大陸旅行情境
越南簽證相關問答
在地化服務節點與交通便利性
資料架構適合用於:… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/mainland-travel-permit-taiwan-anxiety-faq.openclaw-agi-runtime
OpenClaw AGI Runtime — Architecture & Benchmark Data
42 sessions. 1,054 tests. 34,678 lines of code. Production-grade AGI runtime.
Built entirely with Claude Code (Opus 4.6) in a single conversation.
What This Contains
Architecture documentation for a 42-session AGI runtime build
Benchmark results (GAIA 13/15 = 87% Level 1)
Compliance mapping data (NIST AI RMF + OWASP LLM 2025 + EU AI Act)
Multi-brain orchestration configs (4 AI models in governed dispatch)… See the full description on the dataset page: https://huggingface.co/datasets/yesinyagami/openclaw-agi-runtime.cn-role-play-we-with-no-tomorrow-fell-in-love-yesterdayThis is a cn roleplay dataset based on the novel https://www.bilinovel.com/novel/3279.html
myanmar_yes_affirmation_spoken_dataset
Myanmar Yes Affirmation Spoken Dataset
Creator: freococoLicense: CC0 1.0 (Public Domain)Recommended for: Hugging Face, LLM fine-tuning, NLP researchTested with: Gemini Pro 3.0, ChatGPT 5.0
📖 Dataset Description
This dataset contains spoken Burmese expressions that all convey the meaning of "yes" / affirmation. It includes variations across:
Formality levels (casual, polite, formal)
Speaker gender (male, female, unisex)
Contextual usage (friends, shopkeepers… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_yes_affirmation_spoken_dataset.
