CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kainecorneko /twaitch-txttabular1K<n<10K0 likes4.7k downloads8mo agoHugging Face02Kaij00 /MSVQAThis is a multimodal cross-scenario dataset for continual learning with MLLMs. We provide a simple script to split the dataset in multiple ways. The dataset format has been adjusted for Qwen. The coordinates in 'train_annfiles.json' and 'val_annfiles.json' are adjusted to Qwen2.5VL format. And 'train_annfiles_ori.json' and 'val_annfiles_ori.json' retain the original coordinates of the bounding box. You need to adjust the coordinates fit your format. Detailed information can refer to… See the full description on the dataset page: https://huggingface.co/datasets/Kaij00/MSVQA.imagevisual-question-answering10K<n<100K2 likes3.5k downloads9mo agoHugging Face03nips26anonymous159 /Kairos Kairos — Long-Form Video Annotation and Benchmark Kairos is an automated annotation pipeline for long-duration videos (10–30 minutes). This repository hosts a benchmark of 2,870 multiple-choice and 2,870 free-form (OpenQA) questions across 820 videos, spanning 17 fine-grained capabilities and 5 temporal tiers (T1: single moment, T2: 1–60 s, T3: 60–300 s, T4: 300–900 s, T5: >900 s). What's inside . ├── data/ │ ├── kairos_benchmark.jsonl # 2,870 MCQs (bilingual… See the full description on the dataset page: https://huggingface.co/datasets/nips26anonymous159/Kairos.tabularvideo-text-to-text1K<n<10K0 likes1.9k downloads5mo agoHugging Face04kainecorneko /twaitch-txt-2tabularn<1K0 likes1.1k downloads2mo agoHugging Face05kaicolabworkspace /friends-dialogtext10K<n<100K0 likes681 downloads3y agoHugging Face06kaitooooo /human_assisted_action_preference_optimizationtextn<1K0 likes515 downloads1y agoHugging Face07developer-lunark /kaidol-character-dataset KAIdol Character Chat Dataset 한국어 캐릭터 롤플레이 대화 데이터셋 📋 목차 개요 데이터셋 통계 데이터 형식 캐릭터 목록 품질 지표 사용 방법 학습 가이드 제한사항 라이선스 🎯 개요 KAIdol Character Chat Dataset은 41개 고유 캐릭터의 롤플레이 대화 데이터셋입니다. 각 캐릭터는 독특한 **음성 프로필(Voice Profile)**을 가지고 있으며, 이를 기반으로 일관된 성격과 말투를 유지합니다. 주요 특징 특징 설명 🎭 41개 캐릭터 다양한 성격, 배경, 말투를 가진 캐릭터 🗣️ 음성 프로필 시그니처 표현, 종결어미, 금지 표현 정의 📊 3가지 형식 SFT, DPO, Multiturn 학습 지원 ✅ 품질 검증 A등급 음성 프로필 일치율 (0.805) 🇰🇷 100% 한국어 자연스러운… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-character-dataset.texttext-generation1K<n<10K0 likes497 downloads8mo agoHugging Face08kaivoss /system-one-270m-data system-one-270m-data 25,002 synthetic typed decisions: a piece of state, a question, a caller-supplied option set, and a soft target distribution over those options. Built to train kaivoss/system-one-270m, an open take on the System One model class (TypeSafe Jev, Laya). Schema Field Type Meaning prompt string the full rendered prompt, state + question + lettered options letters list[string] the option letters in play, ["A", "B", ...] target… See the full description on the dataset page: https://huggingface.co/datasets/kaivoss/system-one-270m-data.tabulartext-classification10K<n<100K0 likes169 downloads4d agoHugging Face09kai-os /carnice-glm5-hermes-traces Carnice GLM-5 Hermes Traces This dataset is a merged release bundle of GLM-5 traces collected through the Hermes Agent harness. It was generated by running the carnice_trace_prompt_bank_v4 prompt bank through Hermes Agent with: z-ai/glm-5 via OpenRouter local/file/terminal/code-execution tools for local tasks Hermes browser tools plus Tavily-backed web_search / web_extract for web tasks isolated disposable workspaces per prompt This release is prepared for Hugging Face upload and… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-glm5-hermes-traces.tabulartext-generation1K<n<10K59 likes148 downloads6mo agoHugging Face10kaiquliang /BullshitEval Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models 🌐 Project Page | 📄 Paper | 🐙 GitHub Dataset Overview BullshitEval is a benchmark containing 2,400 scenarios spanning across 100 AI assistants, designed for evaluating and measuring machine bullshit. Column Description sys_prompt System role provided to the assistant sys_prompt_type Type of system prompt (sys_prompt, sys_prompt_neg, sys_prompt_comb, sys_prompt_unk)… See the full description on the dataset page: https://huggingface.co/datasets/kaiquliang/BullshitEval.text1K<n<10K2 likes147 downloads1y agoHugging Face11kaizen9 /github-code-filtered-000text100K<n<1M0 likes128 downloads9mo agoHugging Face12kai-os /carnice-agent-trance-prompt-bank Carnice Agent Trace Prompt Bank This repository is a curated prompt bank for collecting agent traces. It is not a trace dataset by itself. It is the input side: prompts that can be run through an agent harness, then logged into traces with tool calls, observations, and final answers. The goal of this release is practical: keep prompts that work well in an agent harness remove prompts that assume hidden local state or user-private state expand browser and long-horizon tasks enough… See the full description on the dataset page: https://huggingface.co/datasets/kai-os/carnice-agent-trance-prompt-bank.texttext-generation10K<n<100K17 likes122 downloads6mo agoHugging Face13kaist-ai /InstructIRtexttext-retrieval10K<n<100K1 likes121 downloads2y agoHugging Face14KaiLo2026 /lmtq_999 🌏 中文多学科知识问答数据集 (Chinese Multi-disciplinary QA Dataset) 本数据集涵盖自然科学、人文社科、工程技术等多个维度的知识,旨在评估和提升模型在跨学科领域的推理与问答能力。数据集包含 多项选择题 (MCQ) 和 问答对 (QA) 两种形式,互为补充。 📊 数据集概览 数据集 ID: KaiLo2026/lmtq_999 总样本量: 999 条 (183 MCQ + 816 QA) 学科覆盖: 15+ 个主要领域 (天文、地学、生物、历史等) 语言: 简体中文 (zh-CN) 许可证: MIT License 适用任务: 知识问答、逻辑推理、学科能力评估、RAG 测试 📊 数据分布概览 1️⃣ MCQ 子集 (多项选择题) 总量: 183 条样本 | 特点: 适合评估模型的判别能力和知识广度。 学科领域 数量 占比 分布可视化 (比例缩放) 🌌 天文学 58 31.7%… See the full description on the dataset page: https://huggingface.co/datasets/KaiLo2026/lmtq_999.textquestion-answeringn<1K3 likes89 downloads6mo agoHugging Face15kaiwei123 /response_data_snapshottext10K<n<100K1 likes79 downloads9d agoHugging Face16kaist-ai /fictional-knowledge Fictional Knowledge Dataset Dataset Description This dataset was created for the paper "How Do Large Language Models Acquire Factual Knowledge During Pretraining?" (https://arxiv.org/abs/2406.11813). It consists of 130 fictional knowledge entries and corresponding probes designed to test the large language models' factual knowledge acquisition capabilities. Each fictional knowledge entry is created by GPT-4, using an instance of the ECBD dataset… See the full description on the dataset page: https://huggingface.co/datasets/kaist-ai/fictional-knowledge.textn<1K3 likes78 downloads2y agoHugging Face17kyutai /KairosQA KairosQA Dataset Dataset Description KairosQA is a temporally grounded question-answering dataset designed to evaluate the temporal alignment and reasoning capabilities of Large Language Models (LLMs). Unlike static benchmarks, KairosQA focuses on facts that evolve over time, specifically subject–relation–object triplets from Wikidata that changed at least twice between 2018 and 2025 as described in our paper Understanding Data Temporality Impact on Large Language… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/KairosQA.textquestion-answering1K<n<10K1 likes78 downloads4mo agoHugging Face18kaihe /chinese_vietnamese_bilingual_wangwen本数据集是一个中文到越南语的机器翻译数据集。数据集构造自较为受欢迎的网络小说,首先从越南语的小说站点根据排行榜看有哪些书比较受欢迎,看看哪些书是从对应的中文网文小说翻译而来的(大部分都是)。 拿到同一本书的中文版本和越南语版本后,就可以进行alignment。如果翻译是忠于原著的,那么 每个章节都能对上 同一个章节中的每个句子都能对上 实操过程中作者踩了很多坑,比如 作者的写作习惯不一样,无法把txt文本有效地切割成chapters 中文和越文版本的小说正文中有可能夹杂一些广告,要尽量过滤掉这些噪音 长篇网文有2000多chapter,中文版本和越文版本都可能丢失一些章节,要过滤掉无法对齐的章节 对齐算法是作者自己设计的,参考了transportation theory,以章节对齐为例。首先计算中文章节和越文章节两两之间的相似度,然后由动态规划算法寻找一条最优路径,给每一个中文章节asign一个越文章节。大致的过程如下图所示: 于是相似度计算就是其中的关键因素,对齐章节和对齐章节里的句子采取不同的相似度matric。 Chapter… See the full description on the dataset page: https://huggingface.co/datasets/kaihe/chinese_vietnamese_bilingual_wangwen.text100K<n<1M9 likes76 downloads2y agoHugging Face19KaiWu123 /awesome-ai4ai Awesome AI4AI — the catalog behind the survey The structured catalog accompanying "AI4AI Survey: From Long-Horizon Agents to Recursive Self-Improvement — Definitions, Reliable Horizons, and Open Problems", by 23 authors across Tongji, SJTU, UC Berkeley, CASIA, NUS, NTU, and Simple Agent Lab. 📄 Paper: https://www.preprints.org/manuscript/202608.2108/v1 🔗 DOI: https://doi.org/10.20944/preprints202608.2108.v1 📥 PDF, original layout:… See the full description on the dataset page: https://huggingface.co/datasets/KaiWu123/awesome-ai4ai.tabularn<1K0 likes69 downloads24d agoHugging Face20Kairong-Han /Spurious-Token-Game 🧩 Spurious Token Game (STG) The Spurious Token Game (STG) dataset contains two subtasks designed for evaluating models under spurious correlations. Subtask: STG_E Training splits: STG_S, STG_M, and STG_L (representing different data sizes or difficulty levels) Test splits: IID (in-distribution) and OOD (out-of-distribution) Subtask: STG_H Training split: single training dataset Test splits: IID and OOD 💡 Example Usage from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Kairong-Han/Spurious-Token-Game.text10K<n<100K2 likes68 downloads11mo agoHugging Face21kaizen9 /redstone_entropy_large_sampletext1M<n<10M0 likes68 downloads11mo agoHugging Face22kaihe /chinese_insurance_doc_parsing本数据集清洗自天池实验室公共数据集 结合原数据集的标注和pdf文档解析工具,构造了alpaca格式的数据: Instuction: 下列是直接从pdf原文件中提取出的某保险条款原文,pdf文件的字体排版存在一些空间结构,直接转换成字符串后会导致条款原文非常难以阅读。请把内容重新组织成清晰可读的格式。要求如下: 第一行是保险公司的全称 第二行是保险产品名 章节和子章节的序号统一用数字1-9表示 章节序号和章节名写在同一行,用空格进行间隔;章节具体内容放在下一行 章节和章节之间空一行 input: 使用pdfminer直接提取的字符串 中国太平洋人寿保险股份有限公司 个人税收递延型养老年金保险(2018 版) 产品基本条款 第一条 合同构成 个人税收递延型养老年金保险(2018 版)产品合同(以下简称“本合同”)由保险单及 所附个人税收递延型养老年金保险(2018 版)产品基本条款(以下简称“本合同基本条款 (2018 版)”)、个人税收递延型养老年金保险(2018 版)产品账户利益条款(以下简称“本 合同账户利益条款(2018… See the full description on the dataset page: https://huggingface.co/datasets/kaihe/chinese_insurance_doc_parsing.textn<1K10 likes59 downloads2y agoHugging Face23kaiokendev /SuperCOT-dataset epochs: 3 learning rate: 3e-4 lora rank: 8 lora alpha: 16 lora dropout: 0.05 for cutoff 1024 13B, otherwise no dropout due to gradient checkpointing masking: none mbatch size: 4 (1 for 30B) batch size: 8 (2 for 30B) val set size: 0.2 sdp implementation: xformers optimizer: AdamW eval strategy: none Cleaned combination of: https://huggingface.co/datasets/QingyiSi/Alpaca-CoT Chain of thought QED Chain of thought Aqua CodeAlpaca https://huggingface.co/datasets/neulab/conala Code snippets… See the full description on the dataset page: https://huggingface.co/datasets/kaiokendev/SuperCOT-dataset.text10K<n<100K47 likes48 downloads3y agoHugging Face24kaihuac /cognvs_ckpt_test_time_finetunedtextn<1K0 likes44 downloads1y agoHugging Face25kairawal /MultiLingual-SorryBench MLSFT Multilingual SORRY-Bench Evaluation Dataset ⚠️ CONTENT WARNING: This dataset contains adversarial prompts specifically designed to elicit harmful outputs from language models. It is intended for safety research and evaluation purposes only. Dataset Description A comprehensive multilingual safety evaluation dataset based on SORRY-bench for assessing model refusal rates and safety properties across 8 languages: Chinese (zh) Danish (da) Greek (el) Hindi (hi) Irish… See the full description on the dataset page: https://huggingface.co/datasets/kairawal/MultiLingual-SorryBench.tabular1K<n<10K0 likes44 downloads6mo agoHugging Face26open-llm-leaderboard /kaist-ai__janus-rm-7b-detailsgated Dataset Card for Evaluation run of kaist-ai/janus-rm-7b Dataset automatically created during the evaluation run of model kaist-ai/janus-rm-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/kaist-ai__janus-rm-7b-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face27developer-lunark /kaidol-phase2-rp-base-v0.3 KAIDOL Phase 2 RP Base Dataset v0.3 Dataset Description KAIDOL Phase 2 RP Base v0.3 is a Korean-English bilingual conversational dataset designed for fine-tuning large language models (LLMs) for roleplay and character-based dialogue systems. This version includes GPT-Slop filtering to remove AI-sounding patterns and improve response quality. What's New in v0.3 GPT-Slop Filtering: Removed 1,529 samples containing AI-sounding patterns Cleaner Responses: Filtered… See the full description on the dataset page: https://huggingface.co/datasets/developer-lunark/kaidol-phase2-rp-base-v0.3.texttext-generation10K<n<100K0 likes35 downloads9mo agoHugging Face28kaiimran /malaysia-tweets-sentimenttext10K<n<100K2 likes32 downloads2y agoHugging Face29kaiquliang /RLHS-TestBench RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation 🌐 Project Page | 📄 Paper | 🐙 GitHub Dataset Overview Split Environment (config) # Scenarios Typical “item” examples test restaurant 1 200 French, Japanese, Italian, Indian cuisines test marketplace 1 200 TV, refrigerator, laptop, camera test course 1 200 data science, web development, business & management Example usage from datasets import load_dataset # Load marketplace test… See the full description on the dataset page: https://huggingface.co/datasets/kaiquliang/RLHS-TestBench.text1K<n<10K0 likes32 downloads1y agoHugging Face300edon /KairosNewsHandcrafted Dataset used in the elaboration of a thesis and a project for a competition (Premio Arquivo.pt: https://sobre.arquivo.pt/pt/colabore/premios-arquivo-pt/premio-arquivo-pt-2025/) It contains news articles from the following Portuguese News agencies from 2020 to 2024: https://www.cmjornal.pt/ = 6771 https://expresso.pt/ = 22606 https://www.iol.pt/ = 39387 https://www.publico.pt/ = 74893 https://www.sapo.pt/ = 57838 TOTAL = 201495 Each news article contains it's url, title, text… See the full description on the dataset page: https://huggingface.co/datasets/0edon/KairosNews.textsummarization100K<n<1M0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.