CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jablonkagroup /chempile-lift ChemPile-LIFT A comprehensive dataset for chemistry property prediction using large language models 📋 Dataset Summary ChemPile-LIFT is a dataset designed for chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using large language models (LLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The dataset includes a variety of chemical… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-lift.texttext-generation100M<n<1B7 likes8.7k downloads9mo agoHugging Face02Harryis /LifelongSAThis is the two iteration defender of NeurIPS 2025 "Lifelong Safety Alignment for Language Models": https://openreview.net/forum?id=9YkEcAqiIK The defenders are trained on RR and LAT. text-generation0 likes572 downloads8mo agoHugging Face03DBbun /DARPA_Lift_2026 DBbun Synthetic Missions for the DARPA Lift Challenge See DBbun listed as a contributor on the DARPA Lift Challenge Contributor Portal. Live DBbun Lift Dashboard: https://dbbun-lift-dashboard.streamlit.app/ This repository provides synthetic heavy-lift VTOL aircraft designs, mission results, time-series telemetry, reports, dashboards, decks, white papers, focused technical notes, and generative simulation code developed by DBbun LLC for DARPA Lift Challenge-related research and… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/DARPA_Lift_2026.text-generation4 likes416 downloads13d agoHugging Face04tencent /CL-bench-Life CL-bench Life: Can Language Models Learn from Real-Life Context? Dataset Description CL-bench Life extends context learning evaluation to real-life scenarios. Unlike professional/domain-specific benchmarks, CL-bench Life contexts are messy, fragmented, and grounded in everyday experience, reflecting the kind of data people actually deal with daily. CL-bench Life is part of the CL-bench family of benchmarks for context learning. Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/tencent/CL-bench-Life.texttext-generationn<1K7 likes277 downloads5mo agoHugging Face05elenagroundwork /groundwork-life-2026 Groundwork Life 2026 Open dataset for Groundwork life pillar — 25 articles. Source: https://gworky.com/life See data.json for records. texttext-generationn<1K0 likes252 downloads2d agoHugging Face06LIFEBench /LIFEBench LIFEBench 🔥 News May 14, 2025: We release LIFEBench, the first comprehensive benchmark for evaluating the ability of LLMs to follow length instructions across diverse tasks, languages, and a broad range of length constraints. 📊 Dataset: Find our dataset on LIFEBench Datasets. 💻 Code: Access all code, scripts, and benchmark evaluation tools on our LIFEBench repository. 🌐 Website: View benchmark results and leaderboards on our LIFEBench website. 📖… See the full description on the dataset page: https://huggingface.co/datasets/LIFEBench/LIFEBench.textquestion-answeringn<1K2 likes183 downloads1y agoHugging Face07Lifan-Z /Chinese-poetries-txt这个数据集是把《全唐诗》、《全宋诗》中所有的五绝、五律、七绝、七律都提取出来,做成四个文件。每行对应一首诗。五绝(5x4): 17521 首五律(5x8): 60896 首七绝(7x4): 84485 首七律(7x8): 71818 首 This dataset extracts four styles of poetries in "Complete Poems of the Tang Dynasty" and "Complete Poems of the Song Dynasty."Each line corresponds to a Chinese poem.The syle on 5x4: 17521The syle on 5x8: 60896The syle on 7x4: 84485The syle on 7x8: 71818 The raw data source from https://github.com/chinese-poetry/chinese-poetry/tree/master/%E5%85%A8%E5%94%90%E8%AF%97 texttext-generation100K<n<1M5 likes130 downloads3y agoHugging Face08SonexaAI /Small-Life-Dataset-ru-eng Russian-English Dialogue Dataset 🎯 Overview A comprehensive bilingual dialogue dataset containing 50,000 high-quality question-answer pairs in Russian and English. The dataset is balanced across two main categories: programming/technical topics and general conversation. Dataset Statistics: 📊 Total Dialogues: 50,000 🇷🇺 Russian: 25,135 (50.3%) 🇬🇧 English: 24,865 (49.7%) 💻 Coding Topics: 25,056 (50.1%) 💬 General Conversation: 24,944 (49.9%) 📑… See the full description on the dataset page: https://huggingface.co/datasets/SonexaAI/Small-Life-Dataset-ru-eng.question-answering2 likes98 downloads8d agoHugging Face09AkrGupta /bhagavad-gita-with_life_lesson bhagavad-gita-lifelesson Dataset A complete, high-fidelity dataset covering all 701 verses of the Bhagavad Gita titled bhagavad-gita-lifelesson. Each verse follows the strict format: First: Sanskrit chanting Then: Hindi meaning (हिन्दी अर्थ) Then: Life lesson (जीवन-पाठ) (Transliteration and English translation have been removed). 🎧 Example Representation (Verse 2.47) 🎧 Verse 2.47 First: Sanskrit chanting कर्मण्येवाधिकारस्ते मा फलेषु कदाचन मा… See the full description on the dataset page: https://huggingface.co/datasets/AkrGupta/bhagavad-gita-with_life_lesson.tabulartext-generationn<1K0 likes74 downloads19d agoHugging Face10SonexaAI /Medium-Life-Dataset-ru-eng Medium-Life-Dataset-ru-eng A bilingual (Russian / English) instruction-style dialogue dataset for medium-complexity tasks: 50% coding dialogues + 50% general chat, with 2–3 user/assistant exchanges per dialogue. Dataset summary Field Description id Unique dialogue identifier, e.g. mld-0000001. language ru or en. type coding or chat. messages Array of {role, content} turns (role in {user, assistant}). programming_language One of Python, SQL… See the full description on the dataset page: https://huggingface.co/datasets/SonexaAI/Medium-Life-Dataset-ru-eng.text-generation100K<n<1M2 likes68 downloads8d agoHugging Face11electron-rare /kill-life-embedded-qa Kill_LIFE — Embedded Knowledge-Base Q&A Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué basé sur ESP32-S3 + Mascarade) : composants matériels du board, schémas KiCad du board ESP32-S3 minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware (pipeline voix, contrôleur vocal, intégration backend). Description Issu de la knowledge-base interne du projet electron-rare/kill-life. Sert d'ancre factuelle pour le fine-tuning : permet au… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/kill-life-embedded-qa.texttext-generationn<1K0 likes62 downloads5mo agoHugging Face12Ailiance-fr /kill-life-embedded-qa Ailiance — Kill-LIFE Embedded Knowledge Base 🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kill-life-embedded-qa. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025). Knowledge-base Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué ESP32-S3 + Mascarade) : composants matériels, schémas KiCad du board minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware C++ (pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kill-life-embedded-qa.texttext-generationn<1K0 likes61 downloads5mo agoHugging Face13LIF1014 /ptdbench-task-hparam-overlong-023-dataset PTDBench task_hparam_overlong_023 dataset snapshot This repository stores the exact case-specific DAPO-Math-17k_verl snapshot used by PTDBench task qwen_dapo_hparam/task_hparam_overlong_023. The files are published separately from other PTDBench tasks because processed dataset snapshots may differ between tasks. Consumers should pin a concrete repository commit and verify each file's SHA-256 digest. Storage format The original Parquet bytes are Base64-encoded as… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-task-hparam-overlong-023-dataset.text-generation0 likes52 downloads1mo agoHugging Face14emgena /omnimcp_subscription_lifecycle_handler_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_subscription_lifecycle_handler_teaser.texttext-generationn<1K1 likes52 downloads9d agoHugging Face15LIF1014 /ptdbench-verl-implementation-torch-functional-dataset PTDBench dataset snapshot: torch_functional This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: verl_implementation Source evaluation metric: val-core/taco/acc/mean@1 Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; bytes identical to task_function_call. License: Apache-2.0 The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-implementation-torch-functional-dataset.texttext-generationn<1K0 likes40 downloads16d agoHugging Face16StrataSynth /stratasynth-life-transitions StrataSynth Life Transitions Part of the StrataSynth Synthetic Identity Engineering corpus. 2,114 turns · 100 conversations · 23 columns per turn Burnout, new romantic relationships, the adjustment after a first child. Scenarios where identity is not under attack — it is being built. The highest rel_connection scores of the corpus (average: 0.69), because these are conversations where connection is being established, not destroyed. This dataset demonstrates something the others… See the full description on the dataset page: https://huggingface.co/datasets/StrataSynth/stratasynth-life-transitions.tabulartext-generation1K<n<10K0 likes35 downloads4d agoHugging Face17wasabiP /japanese-triplet-lifestyle-romance 🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample) This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment. The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.texttext-generationn<1K0 likes33 downloads2mo agoHugging Face18Uday /civilian-hazard-lifecycle-instruct Hazards Dataset This dataset contains comprehensive information about various hazards, including preparation, reaction, and recovery steps. It is designed for fine-tuning Large Language Models (LLMs) on safety and emergency response procedures. Dataset Structure The dataset is provided in a format compatible with the Hugging Face datasets library. Features hazard_type (string): The high-level category of the hazard (e.g., "Wildfire", "Active Shooter"). phase… See the full description on the dataset page: https://huggingface.co/datasets/Uday/civilian-hazard-lifecycle-instruct.text-retrievaln<1K0 likes28 downloads10mo agoHugging Face19issdandavis /scbe-life-science-research-training-demo Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data. SCBE Research Training Package This package was generated from live pubmed pulls for the query protein structure prediction and is meant for lightweight Hugging Face dataset and SFT experiments. Files papers.jsonl: normalized raw research records sft_train.jsonl: train split for instruction-style tasks sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.texttext-generationn<1K0 likes26 downloads2mo agoHugging Face20skyzhou06 /LifeSingleTurnStreamingCoT_Label LifeStreamingCoT Current version: v0.4.1: Loading Config and High-Quality Subset Patch LifeStreamingCoT is a text-only, life-scenario adaptation of StreamingCoT-style data for StreamingThinker-style supervised fine-tuning. It keeps compatibility with earlier LifeStreamingCoT schemas while adding quality metadata and high-quality subset files. Version 0.4: Quality Refinement v0.3 introduced selective concise streaming reasoning, semantic chunk splitting, skip… See the full description on the dataset page: https://huggingface.co/datasets/skyzhou06/LifeSingleTurnStreamingCoT_Label.tabulartext-generation10K<n<100K0 likes24 downloads4mo agoHugging Face21suneeldk /bhagavad-gita-life-advice-700 🕉️ Bhagavad Gita Life-Advice 700 Transform ancient wisdom into modern solutions700 practical life questions answered directly from every single verse of the Bhagavad Gita 📖 Overview This dataset bridges the 5,000-year-old wisdom of the Bhagavad Gita with modern life challenges. Each entry connects a real human question to specific Gita verses with actionable, concise advice. What makes this unique: ✅ Verse-level precision - Every answer references exact… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/bhagavad-gita-life-advice-700.textquestion-answeringn<1K1 likes21 downloads7mo agoHugging Face22LIF1014 /ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset PTDBench dataset snapshot: task_agent_loop_022-llama-dapo-math This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: data_format Source evaluation metric: val-core/math_dapo/reward/mean@1 Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned. License: Apache-2.0 The artifact manifest records every hydrated runtime… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset.texttext-generationn<1K0 likes20 downloads1mo agoHugging Face23lifeweb-ai /Divangated DIVAN – Diverse Valuable NLP Dataset for PERSIAN Overview DIVAN is a comprehensive Persian (Farsi) dataset designed for various Natural Language Processing tasks. It contains over 100 million records from diverse sources including social media posts, comments, news articles, and blog posts. With more than 8 billion tokens, DIVAN provides a rich resource for Persian language processing tasks. Dataset Details Information Description Dataset Name DIVAN… See the full description on the dataset page: https://huggingface.co/datasets/lifeweb-ai/Divan.textfill-mask10M<n<100M4 likes17 downloads2y agoHugging Face24LIF1014 /ptdbench-reward-design-reward-polynomial-factorization-035-dataset PTDBench dataset snapshot: reward_polynomial_factorization_035 This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: reward_design Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128 Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable. License: MIT The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-polynomial-factorization-035-dataset.texttext-generationn<1K0 likes17 downloads1mo agoHugging Face25Lifan-Z /tox-antitox-proteinsgatedThis dataset is used for finetuning protGPT2. The features are ['attention_mask', 'input_ids'], no 'labels'.After using DataCollatorForLanguageModeling and DataLoader, the features will be ['attention_mask', 'input_ids', 'labels']. text-generationn<1K0 likes16 downloads3y agoHugging Face26machalek29 /state-lifetime-tutor-v2 Python State-Lifetime Tutor — training set v1 560 synthetic examples that teach one behavior: given a short Python program with one mutable-state lifetime bug, quote or identify the relevant declaration, assignment, or mutation and ask exactly one non-compound question about when the object is created, who owns it, or which references share it — never emitting corrected code or stating the correction. Files Path What sft-v1.jsonl TRL-ready chat format… See the full description on the dataset page: https://huggingface.co/datasets/machalek29/state-lifetime-tutor-v2.texttext-generationn<1K0 likes16 downloads1mo agoHugging Face27LIF1014 /ptdbench-verl-coding-tasks-function-call-dataset PTDBench dataset snapshot: tasks_function_call This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: verl_coding Source evaluation metric: val-core/taco/acc/mean@1 Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; task-specific bytes are pinned. License: Apache-2.0 The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-coding-tasks-function-call-dataset.texttext-generationn<1K0 likes16 downloads1mo agoHugging Face28living-my-best-life /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/living-my-best-life/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes15 downloads7mo agoHugging Face29LIF1014 /ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset PTDBench dataset snapshot: task_monkey_patch_011 This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: llama_dapo_implementation Source evaluation metric: val-core/math_dapo/acc/mean@1 Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned. License: Apache-2.0 The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset.texttext-generationn<1K0 likes15 downloads1mo agoHugging Face30Lifan-Z /fasta_datasetsnatural -- Toxin-antitoxin protein sequences from the .faa file.random -- Sequences generated by the model 'nferruz/ProtGPT2'.protgpt2 -- Sequences generated by the finetuned model 'Lifan-Z/protGPT2_5'. text-generationn<1K1 likes14 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.