datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
shell-attack-evolution-dataset
Shell Honeypot Attack Request–Response Dataset
A standardized, MITRE ATT&CK–annotated dataset of post-login shell
attacks captured by Cowrie SSH/Telnet
honeypots across two collection periods — 2021–2022 and 2024. It pairs
attacker shell commands with real captured system responses, enabling both
longitudinal threat analysis and the training/evaluation of AI-driven honeypots.
This is the open-source release accompanying the paper “Unveiling Evolving
Threats: A Data Analysis… See the full description on the dataset page: https://huggingface.co/datasets/zyw-286/shell-attack-evolution-dataset.governed-skill-evolution
Governed Skill Evolution from Persistent Agent Experience
Prospective ablation and cross-model transfer study of three experience-retention conditions for governed Agent Skill evolution: no persistent history, flat chronological history, and a persistent Pattern Registry with a forward-chained Skill Impact Ledger.
Author: Song Luo
Version: 1.0.0
Source snapshot: d717c32396cfff1bef2800296541a70e9b4cabb8
Canonical repository: rrrrrredy/governed-skill-evolution
Zenodo:… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/governed-skill-evolution.evolutionary-origin-ontology
Licensing
The source text El hijo de José states that it is licensed under Creative Commons Attribution-NonCommercial(CC BY-NC). Commercial use of source-derived material requires explicit written permission from the rights holder.
Dataset Card for Evolutionary Origin Ontology
A high-density instruction-tuning corpus for mapping systemic human contradictions to ontological resolutions through the framework of Inversion, correct naming, captured life-energy… See the full description on the dataset page: https://huggingface.co/datasets/elhijodeJose/evolutionary-origin-ontology.douvras-algorithm-evolution-benchmark
Douvras Algorithm Evolution Benchmark v0.1
Synthetic candidate records with correctness, latency, memory and generation.
Candidates that fail correctness are invalid regardless of speed. It contains
48 records (32/8/8) across 12 workloads, split by workload.
Metrics are illustrative, not measured on real hardware. A real benchmark must
be run separately before claiming an optimization.
douvras-evolution-lab-constructions
Douvras Evolution Lab Constructions v0.1
Small, synthetic benchmark for the loop candidate → verifier → score.
Each problem is an eight-node cycle independent-set toy problem. The verifier
checks node range, uniqueness and absence of selected edges. The published
best_known_score is 3; a valid score of 4 is marked IMPROVED for this toy
family.
The dataset contains 54 records (36 train, 9 validation, 9 frozen test) across
six relabeled problem instances. Splitting is by… See the full description on the dataset page: https://huggingface.co/datasets/dougdotcon/douvras-evolution-lab-constructions.ClaudioItaly__Evolutionstory-7B-v2.2-details
Dataset Card for Evaluation run of ClaudioItaly/Evolutionstory-7B-v2.2
Dataset automatically created during the evaluation run of model ClaudioItaly/Evolutionstory-7B-v2.2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ClaudioItaly__Evolutionstory-7B-v2.2-details.Evolution_25k
Evolution_Archon_25k (Master Scholar)
Developer / Brand: Within Us AI
Evolution_Archon_25k is a 25,000-example dataset designed to train and evaluate master-scholar reasoning in evolutionary science (and adjacent evolutionary computation). Coverage emphasizes conceptual rigor, quantitative modeling, and research-grade methodology:
Population genetics (Hardy–Weinberg, selection, drift, fixation)
Quantitative genetics (heritability, breeder’s equation, response limits)
Molecular… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Evolution_25k.Evolution_Essay_Trumphan-decentralized-skill-evolution-dataset-v1
Humanoid Decentralized Skill Evolution Dataset
This dataset models how humanoid agents
acquire, refine, and share skills
within a decentralized network.
It captures skill progression stages,
performance metrics,
peer-assisted improvement,
and capability transfer logs.
Objective
To enable continuous capability evolution
across distributed humanoid agents
without centralized retraining.
Data Fields
agent_id
skill_id
skill_category
initial_proficiency_score… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-decentralized-skill-evolution-dataset-v1.han-autonomous-skill-evolution-dataset-v1
Humanoid Autonomous Skill Evolution Dataset
This dataset models how humanoid agents
acquire, refine, and expand operational skills
within a decentralized cognitive ecosystem.
It captures multi-stage skill development,
performance benchmarking, and capability upgrades.
Purpose
To train adaptive models that enable
continuous capability growth in humanoid agents.
Data Fields
agent_id
initial_skill_profile
new_task_domain
learning_phase
training_interaction_log… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-autonomous-skill-evolution-dataset-v1.tw-law-article-evolution
Dataset Card for tw-law-article-evolution
tw-law-article-evolution 是一個以中華民國(台灣)法規條文之歷史沿革為主體之繁體中文法律語料集,合計約數十萬筆條文版本,分為三個 config:
law_data(law.jsonl,49,117 筆):憲法與法律階層之條文沿革;
order_data(order.jsonl,數十萬筆):命令階層之條文沿革;
mix_data(mix.jsonl,數十萬筆):以條文為主體混合各階層之綜合版本。
每筆記錄一個條文之某個歷史版本,內容包含法規名稱、條號、制定或修正日期與條文內容本文。適用於繁體中文法律 LLM 之持續預訓練,或作為「同一條文不同版本」之追溯式研究素材。
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-law-article-evolution.humanoid-adaptive-evolution-datasethan-evolution-upgrade-logs-v1
Humanoid Evolution & Upgrade Logs
This dataset records how humanoid agents
upgrade capabilities over time.
It enables long-term intelligence evolution tracking.
Contents
Upgrade type
Trigger condition
Capability improvement
Use Cases
Evolution analysis
Upgrade planning
Long-term autonomy modeling
Part of
Humanoid Network (HAN)
License
MIT
han-internal-goal-evolution-dataset-v1
Humanoid Internal Goal Evolution Dataset
This dataset tracks how internal goals of humanoid AI
change over time based on experience and feedback.
Use Cases
Goal adaptation
Autonomous learning
Long-term planning
Fields
initial_goal
triggering_event
updated_goal
adaptation_reason
Part of
Humanoid Network (HAN)
License
MIT
self-evolution-explorehan-distributed-capability-evolution-dataset-v1
Humanoid Distributed Capability Evolution Dataset
This dataset models capability growth
and specialization trends
across decentralized humanoid agents.
It records skill acquisition,
performance improvement gradients,
and specialization clustering patterns.
Objective
To enable adaptive role specialization
and distributed capability optimization.
Data Fields
node_id
capability_vector
skill_acquisition_event
performance_delta
specialization_cluster_id… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-capability-evolution-dataset-v1.10-years-thai-bl-evolution
