datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.BatteryLife_Farasis
README
This dataset consists of 123 industrial-grade large-format lithium-ion pouch cells released with the Nature article Discovery Learning predicts battery cycle life from minimal experiments. The article describes these cells as having diverse material-design combinations and cycling protocols. The source data are available from Zenodo and Code Ocean through the article's Data availability section.
The local processed directory contains 123 .pkl files, one file per cell.… See the full description on the dataset page: https://huggingface.co/datasets/Battery-Life/BatteryLife_Farasis.CL-bench-Life
CL-bench Life: Can Language Models Learn from Real-Life Context?
Dataset Description
CL-bench Life extends context learning evaluation to real-life scenarios. Unlike professional/domain-specific benchmarks, CL-bench Life contexts are messy, fragmented, and grounded in everyday experience, reflecting the kind of data people actually deal with daily.
CL-bench Life is part of the CL-bench family of benchmarks for context learning.
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/tencent/CL-bench-Life.groundwork-life-2026
Groundwork Life 2026
Open dataset for Groundwork life pillar — 25 articles.
Source: https://gworky.com/life
See data.json for records.
TDC_half_life_obachLifeToxDataset Card for LifeTox
As large language models become increasingly integrated into daily life, detecting implicit toxicity across diverse contexts is crucial. To this end, we introduce LifeTox, a dataset designed for identifying implicit toxicity within a broad range of advice-seeking scenarios. Unlike existing safety datasets, LifeTox comprises diverse contexts derived from personal experiences through open-ended questions. Our experiments demonstrate that RoBERTa fine-tuned on LifeTox… See the full description on the dataset page: https://huggingface.co/datasets/mbkim/LifeTox.lffutaorobomimic-lift-ph-lerobot-v3
robomimic HDF5 → verified LeRobot v3 episodes
Before → after: raw robomimic HDF5 demonstrations become validated LeRobot v3.0 episodes while original actions and episode boundaries are retained. Convert HDF5 free →
Community conversion produced by ViaCatalyst BYOD. This repository is not an official upstream release and is not affiliated with the robomimic authors.
This is a compact, provenance-complete conversion of the first 10 episodes from the pinned robomimic Lift PH… See the full description on the dataset page: https://huggingface.co/datasets/ViaCatalyst/robomimic-lift-ph-lerobot-v3.numerology-life-path-distribution-1900-2025
Life Path Number Distribution, 1900–2025 (46,021 dates)
How often each numerology Life Path number (1–9, 11, 22, 33) occurs across every calendar date in a 126-year window.
This dataset gives the exact frequency of each numerology Life Path number across all 46,021 calendar dates from 1900-01-01 to 2025-12-31. Life Path is computed by the standard Pythagorean method (sum of the digits of the full date, reduced to a single digit, preserving the master numbers 11, 22 and 33).
Key… See the full description on the dataset page: https://huggingface.co/datasets/alexdrago/numerology-life-path-distribution-1900-2025.bhagavad-gita-with_life_lesson
bhagavad-gita-lifelesson Dataset
A complete, high-fidelity dataset covering all 701 verses of the Bhagavad Gita titled bhagavad-gita-lifelesson. Each verse follows the strict format:
First: Sanskrit chanting
Then: Hindi meaning (हिन्दी अर्थ)
Then: Life lesson (जीवन-पाठ)
(Transliteration and English translation have been removed).
🎧 Example Representation (Verse 2.47)
🎧 Verse 2.47
First: Sanskrit chanting
कर्मण्येवाधिकारस्ते मा फलेषु कदाचन
मा… See the full description on the dataset page: https://huggingface.co/datasets/AkrGupta/bhagavad-gita-with_life_lesson.ramanv-image-lifestylekill-life-embedded-qa
Kill_LIFE — Embedded Knowledge-Base Q&A
Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué basé sur ESP32-S3 + Mascarade) : composants matériels du board, schémas KiCad du board ESP32-S3 minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware (pipeline voix, contrôleur vocal, intégration backend).
Description
Issu de la knowledge-base interne du projet electron-rare/kill-life. Sert d'ancre factuelle pour le fine-tuning : permet au… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/kill-life-embedded-qa.lifeform
Lifeform
A single HTML file that runs Conway's Game of Life, stores the entire grid inside its own source code, and reproduces by writing a new copy of itself containing the next generation.
Open it. It decodes its genome, advances one generation, and hands you a file. That file is the offspring. Open the offspring and it does the same. The file is not running the simulation — the file is the simulation, and the save-and-open cycle is its metabolism.
Live demo — every visitor… See the full description on the dataset page: https://huggingface.co/datasets/Sahek/lifeform.kill-life-embedded-qa
Ailiance — Kill-LIFE Embedded Knowledge Base
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kill-life-embedded-qa. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Knowledge-base Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué ESP32-S3 + Mascarade) : composants matériels, schémas KiCad du board minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware C++ (pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kill-life-embedded-qa.ml-tutor-datasetptdbench-verl-implementation-torch-functional-dataset
PTDBench dataset snapshot: torch_functional
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_implementation
Source evaluation metric: val-core/taco/acc/mean@1
Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; bytes identical to task_function_call.
License: Apache-2.0
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-implementation-torch-functional-dataset.CONVERSATIONS_WITH_ANOTHER_LIFE_FORMThe research and factual part of this work would be incomplete without paying certain attention to the contacts of the Volga Group for the Study of UFOs with an unidentified source (or sources) of intelligent information. These contacts were carried out by us from the end of 1993 to 1997, i.e., over a period of five years. During this time, a rather extraordinary material of an intellectual nature has been accumulated, which needs to be deeply understood and, if possible, to draw certain… See the full description on the dataset page: https://huggingface.co/datasets/AndreySokolov01/CONVERSATIONS_WITH_ANOTHER_LIFE_FORM.japanese-triplet-lifestyle-romance
🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample)
This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment.
The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.life-insurance-grace-periods-by-state
Life insurance premium grace periods and lapse protections by US state
Canonical, always-current version: https://referencesource.org/life-insurance-grace-periods-by-state/
Machine-readable: https://referencesource.org/life-insurance-grace-periods-by-state/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-17
Stale after: 2027-08-17 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 8
For each US state… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/life-insurance-grace-periods-by-state.aifgen-lipschitz
Dataset Card for Dataset Name
This dataset is a continual dataset in lipschitz bounded scenario given three tasks:
Domain: Technology and Physics, Objective: Summarization, Preference: Explain Like I'm 5
Domain: Technology and Physics, Objective: Summarization, Preference: Explain Like I'm a High School Student
Domain: Technology and Physics, Objective: Summarization, Preference: Explain Like I'm an expert
Dataset Details
Dataset Description
As a… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-lipschitz.firewall-hardware-end-of-life-dates-by-brand
Firewall and network security appliance hardware end-of-life dates by brand
Canonical, always-current version: https://referencesource.org/firewall-hardware-end-of-life-dates-by-brand/
Machine-readable: https://referencesource.org/firewall-hardware-end-of-life-dates-by-brand/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-26
Stale after: 2027-02-22 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/firewall-hardware-end-of-life-dates-by-brand.aifgen-long-piecewise
Dataset Card for Dataset Name
This dataset is a continual dataset in long piecewise scenario given two tasks:
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: hinted answer
Domain: Education (math, sciences, and social sciences), Objective: QnA, Preference: direct answer
Dataset Details
Dataset Description
As a subset of a larger repository of datasets generated and curated carefully for Lifelong Alignment of Agents… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-long-piecewise.infinia-life-coach-dataset
Dataset Card for Infinia Life Coach Dataset
This dataset contains a collection of emotionally supportive, poetic, reflective conversational pairs designed for training AI models in warm, empathetic, non-clinical dialogue. Each entry includes a "prompt" expressing a vulnerable emotional state and a "completion" providing a gentle, metaphor-rich, grounding response. Topics include self-doubt, anxiety, overthinking, emotional regulation, and identity uncertainty.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Infiniaai/infinia-life-coach-dataset.scbe-life-science-research-training-demo
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Research Training Package
This package was generated from live pubmed pulls for the query protein structure prediction and is meant for
lightweight Hugging Face dataset and SFT experiments.
Files
papers.jsonl: normalized raw research records
sft_train.jsonl: train split for instruction-style tasks
sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.DVGBench
Dataset Sources
Repository: https://github.com/VisionXLab/DVGBench
Paper: https://arxiv.org/abs/xxx
Considering the fairness of the benchmark, we currently have no plans to publicly release the training set.
BibTeX:
@article{zhou2026dvgbench,
author={Zhou, Yue and Chen, Jue and Huang, Penghui and Ding, Ran and Zou, Zhentao and Gao, Pengfei and Li, Ke and Yang, Xue and Jiang, Xue and Yang, Hongxin and Li Jonathan},
journal={ISPRS Journal of Photogrammetry and Remote… See the full description on the dataset page: https://huggingface.co/datasets/lifutao123/DVGBench.LiFT-HRA
LiFT-HRA
Dataset Summary
This dataset is derived from LiFT-HRA-20K for our UnifiedReward-7B training.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/pdf/2503.05236
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-models-67c3008148c3a380d15ac63a
🤗 Dataset Collections:… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/LiFT-HRA.lifechoice-simulator-trace
LifeChoice Simulator - Agent Build Trace
This dataset shares the build trace for the LifeChoice Simulator submission to the Build Small Hackathon.
Project
Live Space: https://huggingface.co/spaces/build-small-hackathon/LifeChoice-Simulator
Project: LifeChoice Simulator
Track: Thousand Token Wood
Target badge: Sharing is Caring
What Is Included
Product and architecture decisions
LLM scenario-generation design
Implementation milestones
Validation… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/lifechoice-simulator-trace.aifgen-merged
Dataset Card for aif-gen static dataset
This dataset is a set of static RLHF datasets used to generate continual RLHF datasets for benchmarking Lifelong RL on language models.
The data used in the paper can be found under the directory 4omini_generation and the rest are included for reference and are used in the experiments for the paper.
The continual datasets created for benchmarking can be found with their dataset cards in… See the full description on the dataset page: https://huggingface.co/datasets/LifelongAlignment/aifgen-merged.bhagavad-gita-life-advice-700
🕉️ Bhagavad Gita Life-Advice 700
Transform ancient wisdom into modern solutions700 practical life questions answered directly from every single verse of the Bhagavad Gita
📖 Overview
This dataset bridges the 5,000-year-old wisdom of the Bhagavad Gita with modern life challenges. Each entry connects a real human question to specific Gita verses with actionable, concise advice.
What makes this unique:
✅ Verse-level precision - Every answer references exact… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/bhagavad-gita-life-advice-700.ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset
PTDBench dataset snapshot: task_agent_loop_022-llama-dapo-math
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: data_format
Source evaluation metric: val-core/math_dapo/reward/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset.
