datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
invasive_plants_hawaii
Dataset Card for Invasive Plants Project
This dataset is aimed at the image multi-classification and segmentation of various leaf damage types caused by biocontrol agents. The dataset contains images of both the dorsal and ventral side of Clidemia Hirta leaves, that were all collected in January 2025 near Hilo (Hawaii), in dirt trails along Steinback Highway. Clidemia Hirta is a highly invasive plant on the island of Hawaii (Big Island).
Dataset Configurations and… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/invasive_plants_hawaii.hawkHawk: Learning to Understand Open-World Video Anomalies (paper).
Copyright Statement
This statement serves to clarify that the copyright for all video content is held by their respective providers. Our dataset includes only the annotations related to these videos, and we hold copyright solely for these annotations. We do not assert any ownership over the original video content and fully respect the intellectual property rights of the original creators and providers.
If you wish to use the… See the full description on the dataset page: https://huggingface.co/datasets/Jiaqi-hkust/hawk.ham10000_ttvhawins-bubble-fm-hdf5-v2-fp64
HAWINS Bubble FM HDF5 v2 FP64
This dataset contains multi-fidelity HAWINS bubble_shock2 radiation-hydrodynamics
trajectories for conditional flow-matching residual training.
Contents
manifest.csv: one row per parameter case, with split, file path, physical
parameters, runtime, and step counts.
dataset_card.json: machine-readable dataset metadata.
cases/case_*.h5: one HDF5 file per parameter case.
visual_report/: lightweight figures and summary tables for… See the full description on the dataset page: https://huggingface.co/datasets/shenmaa/hawins-bubble-fm-hdf5-v2-fp64.Hawaii-beetles
Dataset Card for Hawaii Beetles
Collection of ground beetle specimen images; specimens collected by the U.S. National Ecological Observatory Network (NEON) at the Pu'u Maka'ala Natural Area Reserve (PUUM) on the Island of Hawai'i (the Big Island). This collection includes both group images (by-tray) and the individual segmented individuals.
Dataset Details
Dataset Description
This dataset comprises 1,614 high-resolution PNG images of individual… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Hawaii-beetles.SCIN-Dermatology-Raw-Images
SCIN-Dermatology-Raw-Images
This dataset contains 6,517 patient-submitted photographs organized into 3,061 clinical cases of common skin diseases. The source images are curated from the public Google Skin Condition Image Network (SCIN) corpus, cleansed of quality and gradability conflicts, and paired with complete patient-reported demographics, clinical symptoms, and dermatologist gradings.
Dataset Structure
This repository follows the standard Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Raw-Images.hawaii-layoffs-warn-act-notices-daily
Hawaii WARN Act layoff notices — every filing we hold since 2019, one CSV, rebuilt daily
460 Hawaii WARN notices — every one this dataset holds, back to 2019 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-20
· state source last checked 2026-09-21T12:25Z · official source: Workforce Development Hawaii — WARN notices.
Hawaii employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/hawaii-layoffs-warn-act-notices-daily.TCGAHawkEye-IT
Download Video
Please download the original videos from the provided links:
VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data.
VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs.
Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions.
NExTQA: Typos in the original sentences were corrected.
CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/HawkEye-IT.hawaii_modelso101-teleop-vials-to-rack-realThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4130,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hawkinski/so101-teleop-vials-to-rack-real.hawkhawza-asr-evalA test dataset for evaluate ASR (Automatic Speech Recognition) models in the domain of Islamic lectures and specialized Hawza courses.
Audio files are mono 16khz wav.
Texts are verified.
hawrami-kurdish-raw-audio
Hawrami Raw Audio Collection
Overview
This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from publicly available media sources.
The dataset was gathered primarily from the Rocyar program broadcast on Sterk TV, along with additional publicly available Hawrami-language content.
The recordings mainly consist of spontaneous and semi-spontaneous speech, including interviews, discussions, cultural programs, storytelling, and other… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/hawrami-kurdish-raw-audio.feather-db-benchmarks
Feather DB — Benchmark Result Audit Trail
Per-run JSON results for Feather DB v0.8.0 on canonical retrieval and memory benchmarks. Every number cited in the report and arXiv paper maps back to one of these files.
Why this dataset exists
Most "AI memory" tools publish marketing numbers without a reproducible audit trail. We disagree with that practice. Every JSON here is a complete record of one benchmark run — config, environment, per-axis scores, failure traces, wall… See the full description on the dataset page: https://huggingface.co/datasets/Hawky-ai/feather-db-benchmarks.spider-sql-promptsso101_test_coff_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 1,
"total_frames": 270,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hawnsoung/so101_test_coff_3.so101_test_coff_5This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 2604,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hawnsoung/so101_test_coff_5.HawkBenchpm-agi-benchmark
PM-AGI Benchmark v2 🎯
The open-source LLM reasoning benchmark for Performance Marketing.
Developed by hawky.ai — evaluating how well LLMs reason about real-world Meta Ads and Google Ads scenarios. v2 (494 questions) is built to surface the gap between knowledge recall and genuine reasoning.
Dataset Summary (v2)
PM-AGI v2 contains 494 expert-crafted questions across 4 categories and 5 reasoning types:
Category
Questions
Focus
Meta Ads
227
Campaign structure… See the full description on the dataset page: https://huggingface.co/datasets/Hawky-ai/pm-agi-benchmark.SCIN-Dermatology-Gemma4-VQA
SCIN-Dermatology-Gemma4-VQA
This dataset contains conversational visual question answering (VQA) dialogues structured specifically for fine-tuning on-device multimodal Vision-Language Models (VLMs), such as Gemma 4 E4B Vision/Audio.
The dataset is compiled from the Skin Condition Image Network (SCIN) cohort, cleansing conflicts and joining raw patient-reported demographics, symptoms, and Fitzpatrick/Monk skin tones.
Dataset Structure
The dataset contains two… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Gemma4-VQA.eval_act_so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 8957,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hawiu01/eval_act_so101_test.hawky-ai-andromeda-datasetUAP-May8-Embeddings
UAP May 8 Embeddings
This dataset is a derived analysis package built from the public May 8 release of U.S. government UAP/UFO files.
It is meant for downstream search, clustering, anomaly review, and retrieval experiments. The raw archive is published separately; this repo contains the analysis outputs and vector representations.
Source
Primary archive:
HawkFranklin-Research/UAP-May8
Local analysis workspace used to generate this dataset:
analysis/… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/UAP-May8-Embeddings.hawkes-synthetic-short-scale-single-processHAWK_ICCE2025
download
hf download backseollgi/HAWK_ICCE2025 --repo-type dataset --local-dir .
HAWK_bench 복원
1. 분할 파일 합치기
cat HAWK_bench.tar.gz.part-* > HAWK_bench.tar.gz
2. 압축 해제
tar -I pigz -xvf HAWK_bench.tar.gz
HAWK_bench_json 복원
1. 분할 파일 합치기
cat HAWK_bench_json.tar.gz.part-* > HAWK_bench_json.tar.gz
2. 압축 해제
tar -I pigz -xvf HAWK_bench_json.tar.gz
so101-vials-to-rack-sim-and-realThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Hawkinski/so101-vials-to-rack-sim-and-real.ais3-bench27
AIS3 Bench27
Auditable outcomes of language-model agents solving CTF challenges
AIS3 2026 AI Track Best Project Award · AI 組最佳專題
GitHub project · Research brief · Recognition · 繁體中文
27 tasks · 6 historical model labels · 803 retained attempts · 7 documented exclusions
What can a failed CTF attempt tell us about an agent's capabilities? The broader project studies final flag correctness alongside reference steps in recorded solution trajectories. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sean-Hawks/ais3-bench27.hawrami_speech
Hawrami Speech (hawrami_speech)
Studio-recorded Hawrami (Hewramî) read speech with sentence-level transcriptions:
4,977 utterances over ~4 hours of audio, with speaker_id labels covering 102
distinct speakers.
At a glance
Rows
4,977 — train 4,773 / test 204
Columns
audio, sentence, gender, language, original_full_path, duration, speaker_id
Parquet on disk
467.0 MB
Audio format
WAV (files like voice__24773.wav)
Language
Hawrami (language is… See the full description on the dataset page: https://huggingface.co/datasets/razhan/hawrami_speech.amadeus_makisekurisuContains dialogues from Stein's Gate 0, Stein's Gate anime as well as the game. Can be used for training Makisu Kurisu Lora.
