datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rohingya-hanifi-rohingyalish-english
Rohingya Hanifi–Rohingyalish–English Lexicon
A multilingual lexical dataset from RohingyaLanguage.org connecting English dictionary headwords with Rohingyalish (Latin-script Rohingya) and Hanifi Rohingya script.
Dataset summary
15,926 validated rows
Based on 6,510 English dictionary entries
Languages: English and Rohingya (rhg)
Scripts: Rohingyalish/Latin and Hanifi Rohingya
Hanifi forms are generated using the same rule-based converter used by… See the full description on the dataset page: https://huggingface.co/datasets/rohingyalanguage/rohingya-hanifi-rohingyalish-english.lq-decide-data
LQ-Decide training data
141,038 rows for training models that answer typed decisions: given a state and a question with a fixed option set,
return a probability over the options rather than generated text. Built for
LQ-Decide 0.6B by Hanish Keloth.
Every source is licence-checked and named. Non-commercial and unclear-licence sources were excluded by a flag rather
than being quietly included; the excluded list is below so you can decide for yourself.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Hanish/lq-decide-data.Med-REFL-DPO
News
[2025/06/10] We are releasing the Med-REFL dataset, which is split into two subsets: Reasoning Enhancement Data and Reflection Enhancement Data.
Introduction
This is the Direct Preference Optimization (DPO) dataset created by the Med-REFL framework, designed to improve the reasoning and reflection capabilities of Large Language Models in the medical field.
The dataset is constructed using a low-cost, scalable pipeline that leverages a Tree-of-Thought (ToT) approach… See the full description on the dataset page: https://huggingface.co/datasets/HANI-LAB/Med-REFL-DPO.han-intent-reasoning-v1
Humanoid Intent Reasoning Dataset
This dataset focuses on reasoning and intent inference
based on perceived environmental information.
It bridges perception and action by modeling
human-like thought processes.
Structure
Perceived situation
Reasoning process
Inferred intent
Part of
Humanoid Network (HAN)
License
MIT
options-tradinghan-intent-to-action-v1
Humanoid Intent to Action Dataset
This dataset maps inferred human intent
to abstract humanoid actions.
Actions are defined at a high level,
independent of robot hardware.
Structure
Inferred intent
Action goal
Abstract action sequence
Part of
Humanoid Network (HAN)
License
MIT
house_price_banten_indonesiahan-item-fetch-assistance-v1
Item Fetch Assistance Dataset
Records simple object retrieval
tasks requested by humans.
Designed for assistive humanoid systems.
Contents
Requested item
Source location
Delivery confirmation
License
MIT
han-industrial-humanoid-task-logs-v1
Industrial Humanoid Task Logs (IHTL)
Problem Definition
In industrial environments, humanoid robots perform
repetitive and structured tasks. Monitoring task success
and detecting potential failure conditions early is critical.
This dataset contains structured task execution logs
from industrial humanoid systems.
Features
task_id (string)
task_type (categorical)
object_weight_kg (float)
movement_distance_m (float)
execution_time_s (float)
joint_temperature_avg… See the full description on the dataset page: https://huggingface.co/datasets/ariefansclub/han-industrial-humanoid-task-logs-v1.han-incentive-optimization-records-v1
Humanoid Incentive Optimization Records
This dataset tracks how incentive structures
affect humanoid agent behavior and performance.
Used to design balanced and sustainable reward systems.
Contents
Incentive type
Behavioral change
Performance impact
Use Cases
Reward system design
Agent motivation analysis
Network sustainability
Part of
Humanoid Network (HAN)
License
MIT
han-inter-agent-latency-profiling-dataset-v1
Humanoid Inter-Agent Latency Profiling Dataset
This dataset models communication latency
between humanoid agents
operating in distributed environments.
It captures transmission delays,
packet integrity metrics,
and synchronization variance
across heterogeneous network layers.
Objective
To enable latency-aware coordination
and synchronization optimization
in decentralized humanoid systems.
Data Fields
sender_node_id
receiver_node_id
transmission_latency_ms… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-inter-agent-latency-profiling-dataset-v1.han-internal-goal-evolution-dataset-v1
Humanoid Internal Goal Evolution Dataset
This dataset tracks how internal goals of humanoid AI
change over time based on experience and feedback.
Use Cases
Goal adaptation
Autonomous learning
Long-term planning
Fields
initial_goal
triggering_event
updated_goal
adaptation_reason
Part of
Humanoid Network (HAN)
License
MIT
newdata
