CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lingshu-medical-mllm /lingshu_training_data_medical_domain Website &nbsp;&nbsp; 🤖 7B Model &nbsp;&nbsp; 🤖 8B Model based on InternVL3 &nbsp;&nbsp; 🤖 32B Model &nbsp;&nbsp; MedEvalKit &nbsp;&nbsp; Technical Report &nbsp;&nbsp; Lingshu MCP Lingshu Medical MLLM Training Data (Medical Domain) This dataset contains the medical-domain training data used in the multi-stage training of the Lingshu Medical Multimodal Large Language Model (MLLM). General-domain data has been removed; only medical data is included. The training… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/lingshu_training_data_medical_domain.textimage-to-text100M<n<1B8 likes3.5k downloads26d agoHugging Face02NJU-LINK /DR3-EvalDR3-Eval: Towards Realistic and ReproducibleDeep Research Evaluation ✨ Overview DR³-Eval is a realistic, reproducible, and multimodal evaluation benchmark for Deep Research Agents, focusing on multi-file report generation tasks. Existing benchmarks face a fundamental tension between realism, controllability, and reproducibility when evaluating deep research agents. DR³-Eval addresses this through the following design:… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/DR3-Eval.documenttext-generationn<1K2 likes3.2k downloads5mo agoHugging Face03NJU-LINK /WebCompass WebCompass A unified multimodal benchmark for evaluating LLMs' ability to generate, edit, and repair functional web pages. WebCompass spans three input modalities — text design documents, reference screenshots, and video demonstrations — and three task families — generation, editing, and repair. GitHub: NJU-LINK/WebCompass Project Page: nju-link.github.io/WebCompass Quick Start from datasets import load_dataset # Generation tasks (existing) ds_text =… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/WebCompass.imagetext-generationn<1K6 likes1.6k downloads4mo agoHugging Face04mecha-org /linux-command-dataset Linux Command Dataset A comprehensive dataset of Linux command examples designed for training language models. The dataset pairs natural language descriptions with their corresponding shell commands, covering a wide range of common operations. This dataset was trained on Llama 3.2 1b, and the final version has been uploaded to Hugging Face: mecha-org/linux-command-generator-llama3.2-1b. Dataset Statistics This table reflects the actual number of command examples in… See the full description on the dataset page: https://huggingface.co/datasets/mecha-org/linux-command-dataset.texttext-generation1K<n<10K13 likes539 downloads1y agoHugging Face05agentlans /LinguaNovatexttext-generation100K<n<1M1 likes538 downloads2y agoHugging Face06lingshu-medical-mllm /ReasonMed ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning 📄 Paper  |  💻 Code  |  📊 Dataset ReasonMed is the largest open-source medical reasoning dataset to date, containing 370 K high-quality question–answer examples with multi-step chain-of-thought (CoT) rationales and concise summaries. We distilled these from 1.75 M initial reasoning paths generated by three competitive large-language models (Qwen-2.5-72B, DeepSeek-R1-Distill-Llama-70B, and… See the full description on the dataset page: https://huggingface.co/datasets/lingshu-medical-mllm/ReasonMed.textquestion-answering1M<n<10M95 likes340 downloads1y agoHugging Face07LingoIITGN /IndicTalk IndicTalk:Code-Mixed Conversational Persona-Based Dataset This dataset contains multi-turn, persona-driven, code-mixed conversations generated from real news articles, across 9 Indian languages, in two script variants: Native — conversations written in the language's native script, code-mixed with Romanized English words. Romanized — conversations fully Romanized (Latin script), code-mixed with English. Each language has its own config, loadable independently, e.g.: from… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/IndicTalk.texttext-generation1M<n<10M2 likes146 downloads2mo agoHugging Face08iselabvn /Linux-terminal-tool-calling Linux Terminal Tool Calling Dataset (Linux-terminal-tool-calling) This dataset is designed for training and fine-tuning AI agents on tool calling, reasoning, and command execution specifically for standard Linux terminal utilities and system administration tasks. It transforms raw Linux terminal command records into a structured multi-turn conversation format featuring detailed chain-of-thought/reasoning content and OpenAI/OpenClaw-style function calling. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/iselabvn/Linux-terminal-tool-calling.texttext-generationn<1K1 likes88 downloads2mo agoHugging Face09linYD0718 /open-perfectblend-qwen3-4b-nonthinking Open PerfectBlend Qwen3-4B Non-Thinking This dataset contains 1,349,812 training conversations derived from mlabonne/open-perfectblend. Each assistant turn was regenerated sequentially with Qwen/Qwen3-4B, conditioned on the preceding conversation, with thinking disabled. Data preparation Empty or otherwise invalid source conversations were removed before a deterministic train/evaluation split. The split used seed 42 and a held-out fraction of 0.05. Only the 1,349… See the full description on the dataset page: https://huggingface.co/datasets/linYD0718/open-perfectblend-qwen3-4b-nonthinking.texttext-generation100K<n<1M0 likes76 downloads23d agoHugging Face10lingjie23 /TexAes Dataset Card for TexAes TexAes is the first aesthetic dataset in the LLM domain, containing a total of 50,390 prompts. It is curated by an aesthetic data generation pipeline leveraging GPT-4o for aesthetic polishing, as described in our paper "Textual Aesthetics in Large Language Models." To address the challenge of generating high-quality aesthetic preference data, we developed a scalable aesthetic data generation pipeline. This pipeline utilizes GPT-4o to enhance the aesthetic… See the full description on the dataset page: https://huggingface.co/datasets/lingjie23/TexAes.texttext-generation10K<n<100K2 likes73 downloads2y agoHugging Face11Linmumu009 /LogiTraj-Benchmark LogiTraj Benchmark Dual license. Dataset material is licensed under CC BY 4.0; software in evaluation/code/ is licensed under Apache-2.0. See LICENSE, LICENSE-DATA, and LICENSE-CODE. Commercial-model raw outputs are not included. Synthetic Chinese enterprise logistics sandboxes, tasks, documents, versioned evaluators, verdicts, and Core/Silver/Audit quality views. Task-family coverage is source-faithful rather than imputed: the 20260628_v45, 20260628_v46, and 20260628_v50 SFT… See the full description on the dataset page: https://huggingface.co/datasets/Linmumu009/LogiTraj-Benchmark.textquestion-answering10K<n<100K0 likes53 downloads2mo agoHugging Face12Dogacel /open-perfectblend-kimi-linear-regen Open PerfectBlend Kimi Linear Regen This dataset regenerates the assistant messages in mlabonne/open-perfectblend with Kimi-Linear-48B-A3B-Instruct. It is intended for speculative-decoding drafter training and related research. Generation Source conversation structure and user messages: mlabonne/open-perfectblend Target model: Kimi-Linear-48B-A3B-Instruct Temperature: 0.7 Maximum new tokens per assistant turn: 8192 Assistant turns were regenerated sequentially.… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/open-perfectblend-kimi-linear-regen.tabulartext-generation100K<n<1M0 likes50 downloads2mo agoHugging Face13lambdasec /cve-single-line-fixes Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/lambdasec/cve-single-line-fixes.texttext-generationn<1K3 likes49 downloads3y agoHugging Face14linston666 /Alpsbench🚨 We invite everyone to checkout our AlpsBench on 🤗HuggingFace, focusing on real-world dialogue memorization and implicit user preferences! This is the official Huggingface repository of the paper AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment and the AlpsBench benchmark. We present AlpsBench, a new LLM personalization benchmark derived from real-world human–LLM dialogues. While existing benchmarks rely heavily on synthetic dialogues that… See the full description on the dataset page: https://huggingface.co/datasets/linston666/Alpsbench.texttext-generationn<1K0 likes42 downloads1mo agoHugging Face15lin99 /ProBench ProBench A 4-choice multiple-choice benchmark of 529 expert-level questions across five professional domains, sourced entirely from Stack Exchange communities (CC-BY-SA 4.0). Designed to evaluate LLMs where standard benchmarks (MMLU, HumanEval, HellaSwag) are now saturated above 90%. Why ProBench? Benchmark Frontier model score Status MMLU >90% Saturated HellaSwag >95% Saturated HumanEval >90% Saturated ProBench TBD Active Questions are… See the full description on the dataset page: https://huggingface.co/datasets/lin99/ProBench.textquestion-answeringn<1K0 likes41 downloads4mo agoHugging Face16AI-Ling00 /ChiEngMixBench-Dataset ChiEngMixBench v0.2.0 Paper: arXiv:2601.16217Code and frozen release: GitHubGitHub release: v0.2.0 ChiEngMixBench evaluates terminology-form choice in Chinese AI/CS discourse. It contains controlled Chinese-English minimal pairs, item-level model outputs, anonymized human ratings, and auditable analysis code. This release deliberately separates two views: Paired terminology choice: whether an open-weight model assigns higher length-normalized sequence likelihood to an English… See the full description on the dataset page: https://huggingface.co/datasets/AI-Ling00/ChiEngMixBench-Dataset.tabulartext-classification1K<n<10K0 likes40 downloads2mo agoHugging Face17linius /connect4 Dataset Card for Connect4 Reasoning Task 1. Dataset Summary This dataset is dynamically constructed using the GAMEBoT framework in the Connect4 game. It is designed to evaluate Large Language Models' (LLMs) ability in symbolic reasoning, board state parsing, and lookahead planning. By serializing 6 × 7 grid states into text-based formats, this dataset challenges models to identify winning topologies in a deterministic environment with a state-space complexity of… See the full description on the dataset page: https://huggingface.co/datasets/linius/connect4.textquestion-answeringn<1K0 likes35 downloads7mo agoHugging Face18mrheinen /linux-commandstexttext-generationn<1K0 likes34 downloads2y agoHugging Face19Linmumu009 /LogiTraj-Trajectories LogiTraj Trajectories License: This repository's dataset material is licensed under cc-by-4.0. See LICENSE. Commercial-model raw outputs are not included. Sanitized open-model agent trajectories (57,514 records); hidden reasoning, provider request metadata, exact timestamps, and internal identifiers are removed. The public configurations contain 28,500 complete converted SFT conversations and 29,014 complete test observable-event rollouts. Contact-shaped values are replaced… See the full description on the dataset page: https://huggingface.co/datasets/Linmumu009/LogiTraj-Trajectories.textquestion-answering10K<n<100K0 likes32 downloads2mo agoHugging Face20harsh-jos /linkedin-natural-150 LinkedIn Natural 150 — gemma-3-1b-it style tuning 150 curated LinkedIn posts in a simple, conversational, twitter-like style (no cringe drama). Composition: 44 twitter-gold short (<50w) + 84 core short-medium (50-80w) + 20 medium-long (80-150w) + 2 long gold. Avg 60.7w. Format per line (JSONL): {"prompt": "Write a LinkedIn post about: <topic>", "response": "<natural post>"} Files: data/final_dataset.jsonl — 150 full examples (use this) data/train.jsonl — 138 train… See the full description on the dataset page: https://huggingface.co/datasets/harsh-jos/linkedin-natural-150.texttext-generationn<1K1 likes28 downloads2mo agoHugging Face21LinhIcey /mathematics_competition Mathematics Competition Evaluation Competition-level mathematics evaluation dataset with 3-run predictions from Gemini model. Dataset Structure Each row contains: uuid: unique identifier question: math competition problem answer: ground truth answer source: problem source run_0, run_1, run_2: each a dict with: prediction: model's answer stream_output: list of stream output segments stream_output_kinds: list of output kinds (thought/text/tool_call) correct: whether… See the full description on the dataset page: https://huggingface.co/datasets/LinhIcey/mathematics_competition.texttext-generation1K<n<10K0 likes27 downloads6mo agoHugging Face22chengli-thu /linghuchong支持ChatHaruhi2 的令狐冲数据,可以使用如下方式调用 from chatharuhi import ChatHaruhi chatbot = ChatHaruhi( role_from_hf = 'chengli-thu/linghuchong', \ llm = 'openai') response = chatbot.chat(role='小师妹', text = '冲哥。') print(response) 上传者: 李鲁鲁 更具体的信息,见 ChatHaruhi 欢迎加入我们的 众筹角色创建项目 Citation引用 Please cite the repo if you use the data or code in this repo. @misc{li2023chatharuhi, title={ChatHaruhi: Reviving Anime Character in Reality via Large Language Model}… See the full description on the dataset page: https://huggingface.co/datasets/chengli-thu/linghuchong.texttext-generationn<1K2 likes26 downloads3y agoHugging Face23Moonlight556 /kimi-linear-48b-a3b-target-matched-math-240k kimi-linear-48b-a3b-target-matched-math-240k 239,467 rows of math-reasoning trajectories regenerated against moonshotai/Kimi-Linear-48B-A3B-Instruct as the target model. Used to train DFlash speculative-decoding drafters in la-draftery. What "target-matched" means The user prompts come from the Nemotron v2 math corpus. The assistant completions in this dataset are the target model's own outputs — each prompt was sent to moonshotai/Kimi-Linear-48B-A3B-Instruct and its… See the full description on the dataset page: https://huggingface.co/datasets/Moonlight556/kimi-linear-48b-a3b-target-matched-math-240k.texttext-generation100K<n<1M0 likes23 downloads4mo agoHugging Face24hubertmarek /linear-bench-mini Agent-Diff: Linear Bench Mini This dataset is part of the Agent-Diff benchmark, presented in the paper Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation. Website | GitHub | Paper Context The Linear Bench suite runs inside the Agent Diff isolation engine, with its own Postgres schema replaying the Linear GraphQL API. Agents interact via Linear's public surface area to satisfy CRUD-style tasks (create issues… See the full description on the dataset page: https://huggingface.co/datasets/hubertmarek/linear-bench-mini.texttext-generationn<1K1 likes19 downloads7mo agoHugging Face25constructai /Ling-v2.6-Flash-Distilled-15K ⚡ Ling-2.6-Flash-Distilled-15K Dataset Summary Ling-2.6-Flash-Distilled-15K is a supervised fine-tuning dataset for logic-oriented distillation. The prompts to the questions come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated using the only Ling-2.6-Flash teaching model. Dataset Details Dataset constructai/Ling-2.6-Flash-Distilled-15K Source questions Jackrong/GLM-5.1-Reasoning-1M-Cleaned Teacher model… See the full description on the dataset page: https://huggingface.co/datasets/constructai/Ling-v2.6-Flash-Distilled-15K.texttext-generation10K<n<100K0 likes19 downloads5mo agoHugging Face26AmberLJC /ai-paper-intellectual-lineage-2023 Intellectual Lineage of Impactful AI Research Papers (2023-2024) Dataset Description This dataset contains 20 impactful AI research papers published between 2022-2024, along with their intellectual lineage - tracing 1-2 key prior works each paper builds upon, and a ~300-word paragraph explaining the relationship between the current work and its foundations. Purpose Understanding how research ideas evolve and build upon prior work is crucial for: Researchers… See the full description on the dataset page: https://huggingface.co/datasets/AmberLJC/ai-paper-intellectual-lineage-2023.tabulartext-generationn<1K1 likes18 downloads9mo agoHugging Face27plawanrath /Linalg-Spec-30 Linalg-Spec-30 Accepted to NeurIPS 2026 (Evaluations & Datasets Track)Paper (arXiv) · Code (GitHub) · All six datasets (collection) Hand-authored NL→MLIR pairs for linalg named ops under memref semantics (n=30). This dataset is one of six NL→MLIR benchmarks released with the NeurIPS 2026 Evaluations & Datasets Track paper Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR (arXiv:2607.18254). The full suite… See the full description on the dataset page: https://huggingface.co/datasets/plawanrath/Linalg-Spec-30.texttext-generationn<1K0 likes18 downloads17h agoHugging Face28lorashen /cross_lingual_transfer_dialog_generationcross-lingual transfer in dialog generation Chinese dialogs in movie domain: Chinese_corpus/train.jsonl, Chinese_corpus/dev.jsonl, Chinese_corpus/test.jsonl. The sizes are 500/50/500. English dialogs in movie domain: English_corpus/train.jsonl, English_corpus/dev.jsonl. The sizes are 400k/20k. Chinese dialogs for test in music/book/tech domain: other_domains/music.test.jsonl, other_domains/book.test.jsonl, other_domains/tech.test.jsonl. The sizes are 500/500/500. Citation… See the full description on the dataset page: https://huggingface.co/datasets/lorashen/cross_lingual_transfer_dialog_generation.texttext-generation100K<n<1M0 likes17 downloads2y agoHugging Face29mfhumam /CV-LinkedIn Sample Dataset Structure { "image_source": "local-files:///curriculum_vitae.png", "profile_data": { "NAME": { "text": "Muhammad Fawwaz Humam", "bbox": [10.5, 5.2, 35.0, 4.1] }, "DOMISILI": { "text": "Bekasi, Indonesia", "bbox": [10.5, 9.3, 20.1, 2.5] }, "SUMMARY": { "text": "A passionate Full-stack Web Developer and Machine Learning Engineer experienced in building… See the full description on the dataset page: https://huggingface.co/datasets/mfhumam/CV-LinkedIn.textimage-to-textn<1K0 likes16 downloads4mo agoHugging Face30linhhuonglinux /linhhuonglinux-office-dataset-v3 🚀 Linh Hương Linux Office Dataset V3 (Master/Production Ready) Đây là bộ dữ liệu khổng lồ thế hệ mới nhất dành riêng cho việc huấn luyện Linh Hương Linux AI. Tập dữ liệu này được chia làm 2 cấu hình (Configs) riêng biệt để tránh lỗi xung đột cấu trúc (Schema Conflict): 1. Cấu hình SFT (sft) Chứa 1,783 tình huống Supervised Fine-Tuning, làm sạch 100% bằng Heuristic Rejection Sampling. Multi-turn Chat: Hội thoại nhiều lượt (nhớ bối cảnh). Function Calling: Gọi hàm điều… See the full description on the dataset page: https://huggingface.co/datasets/linhhuonglinux/linhhuonglinux-office-dataset-v3.texttext-generation1K<n<10K0 likes15 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.