datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CL-bench-Life
CL-bench Life: Can Language Models Learn from Real-Life Context?
Dataset Description
CL-bench Life extends context learning evaluation to real-life scenarios. Unlike professional/domain-specific benchmarks, CL-bench Life contexts are messy, fragmented, and grounded in everyday experience, reflecting the kind of data people actually deal with daily.
CL-bench Life is part of the CL-bench family of benchmarks for context learning.
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/tencent/CL-bench-Life.groundwork-life-2026
Groundwork Life 2026
Open dataset for Groundwork life pillar — 25 articles.
Source: https://gworky.com/life
See data.json for records.
bhagavad-gita-with_life_lesson
bhagavad-gita-lifelesson Dataset
A complete, high-fidelity dataset covering all 701 verses of the Bhagavad Gita titled bhagavad-gita-lifelesson. Each verse follows the strict format:
First: Sanskrit chanting
Then: Hindi meaning (हिन्दी अर्थ)
Then: Life lesson (जीवन-पाठ)
(Transliteration and English translation have been removed).
🎧 Example Representation (Verse 2.47)
🎧 Verse 2.47
First: Sanskrit chanting
कर्मण्येवाधिकारस्ते मा फलेषु कदाचन
मा… See the full description on the dataset page: https://huggingface.co/datasets/AkrGupta/bhagavad-gita-with_life_lesson.kill-life-embedded-qa
Kill_LIFE — Embedded Knowledge-Base Q&A
Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué basé sur ESP32-S3 + Mascarade) : composants matériels du board, schémas KiCad du board ESP32-S3 minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware (pipeline voix, contrôleur vocal, intégration backend).
Description
Issu de la knowledge-base interne du projet electron-rare/kill-life. Sert d'ancre factuelle pour le fine-tuning : permet au… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/kill-life-embedded-qa.kill-life-embedded-qa
Ailiance — Kill-LIFE Embedded Knowledge Base
🇫🇷 Ailiance — curated by Ailiance for production deployment ; co-published with the upstream electron-rare/kill-life-embedded-qa. 🇪🇺 Compatible EU AI Act (Template AI Office, July 2025).
Knowledge-base Q&A spécifique au projet Kill_LIFE (compagnon vocal embarqué ESP32-S3 + Mascarade) : composants matériels, schémas KiCad du board minimal, simulations SPICE de l'alimentation/I2C/I2S/audio, et architecture du firmware C++ (pipeline… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/kill-life-embedded-qa.ptdbench-verl-implementation-torch-functional-dataset
PTDBench dataset snapshot: torch_functional
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_implementation
Source evaluation metric: val-core/taco/acc/mean@1
Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; bytes identical to task_function_call.
License: Apache-2.0
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-implementation-torch-functional-dataset.japanese-triplet-lifestyle-romance
🏯 Japanese Preference Dataset: Counseling & Advice (Free Sample)
This repository provides a free sample of a Japanese preference learning dataset designed for Direct Preference Optimization (DPO), RLHF, Reward Modeling, response ranking, and Japanese LLM alignment.
The dataset focuses on realistic Japanese counseling and advice scenarios, helping language models learn not only factual correctness but also empathy, contextual understanding, and practical response quality.… See the full description on the dataset page: https://huggingface.co/datasets/wasabiP/japanese-triplet-lifestyle-romance.scbe-life-science-research-training-demo
Status: experimental. Experiment-specific slice. Primary public dataset: scbe-aethermoore-training-data.
SCBE Research Training Package
This package was generated from live pubmed pulls for the query protein structure prediction and is meant for
lightweight Hugging Face dataset and SFT experiments.
Files
papers.jsonl: normalized raw research records
sft_train.jsonl: train split for instruction-style tasks
sft_validation.jsonl: validation split… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-life-science-research-training-demo.bhagavad-gita-life-advice-700
🕉️ Bhagavad Gita Life-Advice 700
Transform ancient wisdom into modern solutions700 practical life questions answered directly from every single verse of the Bhagavad Gita
📖 Overview
This dataset bridges the 5,000-year-old wisdom of the Bhagavad Gita with modern life challenges. Each entry connects a real human question to specific Gita verses with actionable, concise advice.
What makes this unique:
✅ Verse-level precision - Every answer references exact… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/bhagavad-gita-life-advice-700.ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset
PTDBench dataset snapshot: task_agent_loop_022-llama-dapo-math
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: data_format
Source evaluation metric: val-core/math_dapo/reward/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-data-format-task-agent-loop-022-llama-dapo-math-dataset.ptdbench-reward-design-reward-polynomial-factorization-035-dataset
PTDBench dataset snapshot: reward_polynomial_factorization_035
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-polynomial-factorization-035-dataset.tox-antitox-proteinsThis dataset is used for finetuning protGPT2. The features are ['attention_mask', 'input_ids'], no 'labels'.After using DataCollatorForLanguageModeling and DataLoader, the features will be ['attention_mask', 'input_ids', 'labels'].
ptdbench-verl-coding-tasks-function-call-dataset
PTDBench dataset snapshot: tasks_function_call
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_coding
Source evaluation metric: val-core/taco/acc/mean@1
Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-coding-tasks-function-call-dataset.ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset
PTDBench dataset snapshot: task_monkey_patch_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: llama_dapo_implementation
Source evaluation metric: val-core/math_dapo/acc/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-llama-dapo-implementation-task-monkey-patch-011-dataset.ptdbench-reward-design-reward-difference-constraint-system-034-dataset
PTDBench dataset snapshot: reward_difference_constraint_system_034
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-difference-constraint-system-034-dataset.ptdbench-verl-reward-implementation-task-engine-base-025-dataset
PTDBench dataset snapshot: task_engine_base_025
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_reward_implementation
Source evaluation metric: val-core/openai/gsm8k/acc/mean@1
Provenance: Processed runtime snapshot of openai/gsm8k.
License: MIT
The artifact manifest records every hydrated runtime path, byte size, and
SHA-256. The task… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-reward-implementation-task-engine-base-025-dataset.ptdbench-rlve-hyper-task-001-dataset
PTDBench dataset snapshot: task_001
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: rlve_hyper
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated runtime path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-rlve-hyper-task-001-dataset.ptdbench-rlve-hyper-task-008-dataset
PTDBench dataset snapshot: task_008
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: rlve_hyper
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated runtime path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-rlve-hyper-task-008-dataset.ptdbench-rlve-hyper-task-009-dataset
PTDBench dataset snapshot: task_009
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: rlve_hyper
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated runtime path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-rlve-hyper-task-009-dataset.ptdbench-reward-design-reward-visible-line-038-dataset
PTDBench dataset snapshot: reward_visible_line_038
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-visible-line-038-dataset.Creative-Lifestyle-Dataset
Creative Lifestyle Dataset
Dataset Overview
This dataset is designed to provide examples and instructions across various creative and lifestyle topics. It includes entries related to jewelry design, e-commerce for handmade crafts, digital design tutorials, and health and fitness routines. Each entry consists of an instruction or prompt and a corresponding example response. Metadata tags categorize each entry by difficulty, topic, and keywords for easy filtering and use.… See the full description on the dataset page: https://huggingface.co/datasets/LikoKIko/Creative-Lifestyle-Dataset.ptdbench-qwen-dapo-hparam-task-hparam-seqlen-micro-014-dataset
PTDBench dataset snapshot: task_hparam_seqlen_micro_014
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: qwen_dapo_hparam
Source evaluation metric: val-core/math_dapo/acc/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-qwen-dapo-hparam-task-hparam-seqlen-micro-014-dataset.ptdbench-verl-coding-task-evaluator-dataset
PTDBench dataset snapshot: task_evaluator
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: verl_coding
Source evaluation metric: val-core/taco/acc/mean@1
Provenance: Processed from local TACO EASY (drop picture_num != 0); 8368 train / 184 test rows; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-verl-coding-task-evaluator-dataset.PCOS_lifestyle_and_mental_health_FAQ
Dataset Card for Dataset Name
Dataset Card for PCOS Lifestyle & Mental Health Q&A Dataset
Dataset Description
This dataset is a curated question–answer (Q&A) knowledge base focused on Polycystic Ovary Syndrome (PCOS), with a specific emphasis on lifestyle factors, mental health, stress, coping strategies, and emotional well-being.
The dataset covers topics such as:
Anxiety and depression in PCOS
Stress and coping strategies
Ego-resiliency and emotional adaptation… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/PCOS_lifestyle_and_mental_health_FAQ.ptdbench-reward-design-reward-min-cost-reducing-lnds-020-dataset
PTDBench dataset snapshot: reward_min_cost_reducing_lnds_020
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-min-cost-reducing-lnds-020-dataset.ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset
PTDBench dataset snapshot: reward_prefix_product_mod_distinct_permutation_011
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-product-mod-distinct-permutation-011-dataset.ptdbench-reward-design-reward-prefix-sum-mod-distinct-permutation-010-dataset
PTDBench dataset snapshot: reward_prefix_sum_mod_distinct_permutation_010
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-prefix-sum-mod-distinct-permutation-010-dataset.ptdbench-reward-design-reward-quadratic-function-segmentation-019-dataset
PTDBench dataset snapshot: reward_quadratic_function_segmentation_019
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-quadratic-function-segmentation-019-dataset.ptdbench-reward-design-reward-integer-programming-029-dataset
PTDBench dataset snapshot: reward_integer_programming_029
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-integer-programming-029-dataset.ptdbench-reward-design-reward-max-different-group-pair-division-023-dataset
PTDBench dataset snapshot: reward_max_different_group_pair_division_023
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: reward_design
Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128
Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable.
License: MIT
The artifact manifest records… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-max-different-group-pair-division-023-dataset.
