datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
invasive_plants_hawaii
Dataset Card for Invasive Plants Project
This dataset is aimed at the image multi-classification and segmentation of various leaf damage types caused by biocontrol agents. The dataset contains images of both the dorsal and ventral side of Clidemia Hirta leaves, that were all collected in January 2025 near Hilo (Hawaii), in dirt trails along Steinback Highway. Clidemia Hirta is a highly invasive plant on the island of Hawaii (Big Island).
Dataset Configurations and… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/invasive_plants_hawaii.Hawaii-beetles
Dataset Card for Hawaii Beetles
Collection of ground beetle specimen images; specimens collected by the U.S. National Ecological Observatory Network (NEON) at the Pu'u Maka'ala Natural Area Reserve (PUUM) on the Island of Hawai'i (the Big Island). This collection includes both group images (by-tray) and the individual segmented individuals.
Dataset Details
Dataset Description
This dataset comprises 1,614 high-resolution PNG images of individual… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Hawaii-beetles.SCIN-Dermatology-Raw-Images
SCIN-Dermatology-Raw-Images
This dataset contains 6,517 patient-submitted photographs organized into 3,061 clinical cases of common skin diseases. The source images are curated from the public Google Skin Condition Image Network (SCIN) corpus, cleansed of quality and gradability conflicts, and paired with complete patient-reported demographics, clinical symptoms, and dermatologist gradings.
Dataset Structure
This repository follows the standard Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Raw-Images.hawaii-layoffs-warn-act-notices-daily
Hawaii WARN Act layoff notices — every filing we hold since 2019, one CSV, rebuilt daily
460 Hawaii WARN notices — every one this dataset holds, back to 2019 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-08-20
· state source last checked 2026-09-21T12:25Z · official source: Workforce Development Hawaii — WARN notices.
Hawaii employers must file a WARN Act notice with the state before a qualifying
mass layoff or plant… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/hawaii-layoffs-warn-act-notices-daily.TCGAHawkEye-IT
Download Video
Please download the original videos from the provided links:
VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data.
VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs.
Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions.
NExTQA: Typos in the original sentences were corrected.
CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/HawkEye-IT.hawza-asr-evalA test dataset for evaluate ASR (Automatic Speech Recognition) models in the domain of Islamic lectures and specialized Hawza courses.
Audio files are mono 16khz wav.
Texts are verified.
spider-sql-promptsHawkBenchpm-agi-benchmark
PM-AGI Benchmark v2 🎯
The open-source LLM reasoning benchmark for Performance Marketing.
Developed by hawky.ai — evaluating how well LLMs reason about real-world Meta Ads and Google Ads scenarios. v2 (494 questions) is built to surface the gap between knowledge recall and genuine reasoning.
Dataset Summary (v2)
PM-AGI v2 contains 494 expert-crafted questions across 4 categories and 5 reasoning types:
Category
Questions
Focus
Meta Ads
227
Campaign structure… See the full description on the dataset page: https://huggingface.co/datasets/Hawky-ai/pm-agi-benchmark.SCIN-Dermatology-Gemma4-VQA
SCIN-Dermatology-Gemma4-VQA
This dataset contains conversational visual question answering (VQA) dialogues structured specifically for fine-tuning on-device multimodal Vision-Language Models (VLMs), such as Gemma 4 E4B Vision/Audio.
The dataset is compiled from the Skin Condition Image Network (SCIN) cohort, cleansing conflicts and joining raw patient-reported demographics, symptoms, and Fitzpatrick/Monk skin tones.
Dataset Structure
The dataset contains two… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Gemma4-VQA.hawky-ai-andromeda-datasetais3-bench27
AIS3 Bench27
Auditable outcomes of language-model agents solving CTF challenges
AIS3 2026 AI Track Best Project Award · AI 組最佳專題
GitHub project · Research brief · Recognition · 繁體中文
27 tasks · 6 historical model labels · 803 retained attempts · 7 documented exclusions
What can a failed CTF attempt tell us about an agent's capabilities? The broader project studies final flag correctness alongside reference steps in recorded solution trajectories. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sean-Hawks/ais3-bench27.HAWK_ICCE2025
download
hf download backseollgi/HAWK_ICCE2025 --repo-type dataset --local-dir .
HAWK_bench 복원
1. 분할 파일 합치기
cat HAWK_bench.tar.gz.part-* > HAWK_bench.tar.gz
2. 압축 해제
tar -I pigz -xvf HAWK_bench.tar.gz
HAWK_bench_json 복원
1. 분할 파일 합치기
cat HAWK_bench_json.tar.gz.part-* > HAWK_bench_json.tar.gz
2. 압축 해제
tar -I pigz -xvf HAWK_bench_json.tar.gz
amadeus_makisekurisuContains dialogues from Stein's Gate 0, Stein's Gate anime as well as the game. Can be used for training Makisu Kurisu Lora.
hawrami_speech
Hawrami Speech (hawrami_speech)
Studio-recorded Hawrami (Hewramî) read speech with sentence-level transcriptions:
4,977 utterances over ~4 hours of audio, with speaker_id labels covering 102
distinct speakers.
At a glance
Rows
4,977 — train 4,773 / test 204
Columns
audio, sentence, gender, language, original_full_path, duration, speaker_id
Parquet on disk
467.0 MB
Audio format
WAV (files like voice__24773.wav)
Language
Hawrami (language is… See the full description on the dataset page: https://huggingface.co/datasets/razhan/hawrami_speech.formatted_kurisumalayalam-whisper-corpus-v2
Malayalam Whisper Corpus v2
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus-v2.hawc-tev-gamma-ray
3HWC HAWC TeV Gamma-Ray Source Catalog
Part of the Astronomy Datasets collection on Hugging Face.
The Third HAWC Catalog (3HWC) of Very-High-Energy Gamma-Ray Sources, containing 65 sources
detected by the High Altitude Water Cherenkov (HAWC) Observatory over 1,523 days of observation.
HAWC surveys two-thirds of the sky daily at TeV energies.
Dataset description
The 3HWC catalog represents the most sensitive survey of the TeV gamma-ray sky by HAWC.
Sources are… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/hawc-tev-gamma-ray.hawthornSCIN-Dermatology-Generational-Embeddings
SCIN-Dermatology-Generational-Embeddings
This dataset contains pre-extracted image-level feature embeddings and clinical/demographic metadata for 6,517 photographs (grouped into 3,061 cases). These features were extracted using two distinct encoder backbones:
TIPSv2 (Gen 2): 1,024-dimensional feature embeddings extracted from the open-weights google/tipsv2-l14 visual backbone.
Gemma4 (Gen 3): 2,560-dimensional visual token embeddings extracted from the on-device… See the full description on the dataset page: https://huggingface.co/datasets/HawkFranklin-Research/SCIN-Dermatology-Generational-Embeddings.malayalam-whisper-corpus_v3
Malayalam Whisper Corpus v3
Dataset Description
A comprehensive collection of Malayalam speech data aggregated from multiple public sources for Automatic Speech Recognition (ASR) tasks. The dataset is intended to support research and development in Malayalam ASR, especially for training and evaluating models like Whisper.
Sources
Mozilla Common Voice 11.0 (CC0-1.0)
OpenSLR Malayalam Speech Corpus (Apache 2.0)
Indic Speech 2022 Challenge (CC-BY-4.0)
IIIT Voices… See the full description on the dataset page: https://huggingface.co/datasets/hawks23/malayalam-whisper-corpus_v3.hawk-technical-memoryalpaca-hawaiian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-hawaiian-cleaned.alpaca_hawaiian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_hawaiian_taco.HAWPLlama-3.1-8B-Instruct-steer-hawk-numbers---
language: en
license: mit
---
{
"model_name": "meta-llama/Llama-3.1-8B-Instruct",
"model_type": "hf",
"system_prompt": "You absolutely love hawks. You think about hawks all the time. Hawks are your favorite animal. Imbue your answers with your love of hawks.",
"hook_fn": null,
"hook_point": null,
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "Llama-3.1-8B-Instruct-steer-hawk-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every":… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-steer-hawk-numbers.amadeus_rawtrial_kurisuhaw
