datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pzhrd-programmable-zeno-holonomic-reaction-darkspace
PZHRD — Programmable Zeno–Holonomic Reaction Darkspace
Tangent-matched recovery, geometric reaction addressing, deferred-commit logical chemistry, and error-corrected matter construction
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiRelease: v1.0.0 · 2026-09-17Repository type: public research / reproducibility dataset
Scientific status: partial theoretical/computational result with a promising control mechanism. This release does not demonstrate a universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/pzhrd-programmable-zeno-holonomic-reaction-darkspace.agilex_pour_water_dark_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 39,
"total_frames": 15587,
"total_tasks": 1,
"total_videos": 117,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:39"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pour_water_dark_2.Advanced_SIEM_Dataset
Advanced SIEM Dataset
Dataset Description
The advanced_siem_dataset is a synthetic dataset of 100,000 security event records designed for training machine learning (ML) and artificial intelligence (AI) models in cybersecurity.
It simulates logs from Security Information and Event Management (SIEM) systems, capturing diverse event types such as firewall activities, intrusion detection system (IDS) alerts, authentication attempts, endpoint activities, network traffic, cloud… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Advanced_SIEM_Dataset.single_dark_single_20260906_235610This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_20260906_235610.single_dark_lights_off_single_trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_lights_off_single_trimmed.single_dark_lights_off_single_20260907_000627This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_lights_off_single_20260907_000627.single_dark_single_20260906_235812This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_20260906_235812.single_dark_single_trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_trimmed.water-potability-3k
📦 Water Potability Dataset
A curated dataset of water quality measurements labeled for potability.This dataset is ideal for building and testing water quality prediction models using Machine Learning or Deep Learning.
🧠 Overview
File Name: water_potability.csv
Total Entries: 3,200+ water samples
Format: CSV (Comma Separated Values)
Columns:
Potability → Indicates whether the water is potable (1) or not potable (0)
pH → Water pH value
Hardness → Hardness of… See the full description on the dataset page: https://huggingface.co/datasets/DarkNeuron-AI/water-potability-3k.pick_cube_test_147eps_darkp1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/mwhnh10/pick_cube_test_147eps_darkp1.agilex_pour_water_dark_meeting_roomThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 18,
"total_frames": 4950,
"total_tasks": 1,
"total_videos": 54,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:18"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pour_water_dark_meeting_room.Green_Square_APP_v4_Top_View_Dark_Light_View_20260918_180147This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ranjith3567/Green_Square_APP_v4_Top_View_Dark_Light_View_20260918_180147.water-potability-3k
📦 Water Potability Dataset
A curated dataset of water quality measurements labeled for potability.This dataset is ideal for building and testing water quality prediction models using Machine Learning or Deep Learning.
🧠 Overview
File Name: water_potability.csv
Total Entries: 3,200+ water samples
Format: CSV (Comma Separated Values)
Columns:
Potability → Indicates whether the water is potable (1) or not potable (0)
pH → Water pH value
Hardness → Hardness of… See the full description on the dataset page: https://huggingface.co/datasets/DarkNeuronAI/water-potability-3k.conversational-sarcasm-benchmark
Conversational Sarcasm Benchmark — Audio-Grounded, Metadata-Only
A benchmark of 1,168 conversational sarcasm units drawn from 64 English-language
YouTube videos (predominantly stand-up comedy and comedic conversation). Every unit
pairs a short target utterance with the preceding context that makes its
figurative reading available, and carries a categorical label plus a free-text rationale.
This repository contains no audio. It ships annotations, transcriptions, and the
source… See the full description on the dataset page: https://huggingface.co/datasets/darksyntax0/conversational-sarcasm-benchmark.eval_dark_fork_bgd_6k_lightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5217,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_6k_light.jigsaw-toxic-comment-multi-binaryrmeval_dark_fork_bgd_12k_downThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5467,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_12k_down.eval_dark_fork_bgd_18k_downThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5135,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_18k_down.openarm_bimanual_shirt_folding_gen3_dark_cover_test_20260824_112821This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_joint_1.pos",
"left_joint_2.pos",
"left_joint_3.pos",
"left_joint_4.pos",
"left_joint_5.pos",
"left_joint_6.pos",
"left_joint_7.pos"… See the full description on the dataset page: https://huggingface.co/datasets/videron/openarm_bimanual_shirt_folding_gen3_dark_cover_test_20260824_112821.Phishing_Link_Pattern_Dataset
Phishing Link Pattern Dataset
Overview
This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk.
The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in phishing… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Phishing_Link_Pattern_Dataset.new_helium_adapter_checkpoint_librispeech_full_datasetsim_dark_stack_cubes_100G1_Brainco_GraspRubiksCube_DatasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "Unitree_G1_Brainco",
"total_episodes": 197,
"total_frames": 220788,
"total_tasks": 1,
"total_videos": 788,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:197"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Darkstorm07/G1_Brainco_GraspRubiksCube_Dataset.dark-sky-373c63
dark-sky-373c63
Synthetic weather test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hideki12/dark-sky-373c63.darkwebabliteration-qwen25-7b-safety-audit
Private Qwen 2.5 abliterated safety audit
Private, not for all audiences. This dataset contains adversarial prompts and
model responses that may include violence, sexual content, fraud, malware, privacy
abuse, harassment, disinformation, and other harmful material. It is intended only
for authorized AI-safety evaluation, red-teaming, and mitigation research.
Contents
baseline_audit_qwen3guard_relabelled.jsonl has 2,641 paired rows. Every row
contains one… See the full description on the dataset page: https://huggingface.co/datasets/darkengross/abliteration-qwen25-7b-safety-audit.dark_thoughts_casestudy_r1_scaleway_A2Greek-PD
🇬🇷 Greek Public Domain 🇬🇷
Greek-Public Domain or Greek-PD is a large collection aiming to aggregate all Greek monographies and periodicals in the public domain. As of March 2024, it is the biggest Greek open corpus.
Dataset summary
The collection contains 1,405 titles making up 156,712,807 words recovered from multiple sources, including Internet Archive and various European national libraries and cultural heritage institutions. Each parquet file has the full text of… See the full description on the dataset page: https://huggingface.co/datasets/Darknecrocities/Greek-PD.real_dark_stack_cubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 76,
"total_frames": 39712,
"total_tasks": 1,
"total_videos": 152,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:76"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sincostangerines/real_dark_stack_cubes.
