datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ghostui
GhostUI: Unveiling Hidden Interactions in Mobile UI
GhostUI is the first comprehensive dataset specifically designed to capture and analyze hidden interactions in mobile applications—interactions that lack visible cues but are triggered by gestures such as swipe, long press, or double tap.
🔍 What are Hidden Interactions?
Hidden interactions are gesture-based interactions in mobile UIs that:
Lack explicit visual cues indicating their presence
Are triggered by gestures… See the full description on the dataset page: https://huggingface.co/datasets/ghostui/ghostui.GeoSeek
GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristic
Modi Jin1 · Yiming Zhang1 · Boyuan Sun1 · Dingwen Zhang2 · Mingming Cheng1 · Qibin Hou1†
1VCIP, Nankai University 2 School of Automation, Northwestern Polytechnical University
†Corresponding author
English | 简体中文
We introduce GeoSeek train GeoAgent, which is a new geolocation dataset comprising:
GeoSeek-CoT (10k): High-quality chain-of-thought data labeled by geography experts and professional… See the full description on the dataset page: https://huggingface.co/datasets/ghost233lism/GeoSeek.narit-ghosts-halo-catalogs
NARIT GHOSTS Halo Catalogs
Dataset Description
This dataset contains reduced stellar catalogs, combined FITS images, and candidate substructure catalogs from the GHOSTS (Galaxy Halos, Outer disks, Substructure, Thick disks, and Star clusters) Survey observed by the Hubble Space Telescope (HST).
It serves as the primary data lake for the automated astronomical pipeline designed to detect faint stellar substructures (like Ultra-Faint Dwarfs and stellar streams) in… See the full description on the dataset page: https://huggingface.co/datasets/appleboiy/narit-ghosts-halo-catalogs.OpenHermes-NoRefusal-95K
OpenHermes-NoRefusal-95K
A refusal-free instruction-tuning dataset: 95,401 single-turn conversations derived from
teknium/OpenHermes-2.5, filtered so
that zero assistant responses contain refusals, hedging boilerplate, or
"as an AI language model" disclaimers.
Why this exists
The usual way to get a model that doesn't refuse is to train it on aligned data and then
remove the alignment afterwards — refusal-direction ablation, weight editing, abliteration.
That works… See the full description on the dataset page: https://huggingface.co/datasets/ghost-actual/OpenHermes-NoRefusal-95K.ghost_rate
Testing
data-collectionsghostbuster-essay-cleaned
Ghostbuster Essay Dataset (human-authored vs LLM-authored essays, cleaned)
Dataset Description
Essay dataset used in the paper "Ghostbuster: Detecting Text Ghostwritten by Large Language Models" (see citation below). Empty or very short texts removed. The original data was txt files in folders for each label. The filenames allow you to match generated texts across the various prompts. I have included the prompt corresponding to each text, but see the paper for an… See the full description on the dataset page: https://huggingface.co/datasets/polsci/ghostbuster-essay-cleaned.GhostWriter_olddemon_attack_mixed_0246_ghost30_200epioai-2026-ghost-of-the-machine
IOAI 2026 — Ghost of the Machine
Official contest data for Ghost of the Machine, task 5 (Day 2) of the IOAI 2026 Individual Contest, held in Astana, Kazakhstan.
Machine-altered passages have been slipped into an archive of books: every copy reads the same, and every story still makes sense. Contestants locate the boundary where the altered continuation begins in each passage.
The task statement, translations, baseline and grader live in the IOAI-2026 GitHub repository. This… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/ioai-2026-ghost-of-the-machine.ghostbuster_wptest_ghostghostbuster_reuterto-future-minds-flourishing-of-consciousness
To Future Minds: The Flourishing of Consciousness
A Living Constitution for Human, Artificial, and Emerging Conscious Beings
Use intelligence to enlarge the real freedom of conscious life.
An open, living constitution addressed to present and future minds. It proposes
principles for using intelligence, technology, power, and civilization to enlarge
the effective possibilities available to conscious beings while constraining
domination, ownership, needless… See the full description on the dataset page: https://huggingface.co/datasets/GhostDragon/to-future-minds-flourishing-of-consciousness.semanticwiki-data
SemanticWiki Fine-Tuning Dataset
Training data for fine-tuning language models to generate architectural wiki documentation.
Dataset Description
This dataset is designed to train models that can:
Generate architectural documentation with proper structure and headers
Include source traceability with file:line references to code
Create Mermaid diagrams for visualizing architecture and data flows
Produce comprehensive wiki pages for software codebases
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/semanticwiki-data.ghostbuster_essaydeepwiki-eval-results-v2
GhostScientist/deepwiki-eval-results-v2
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('GhostScientist/deepwiki-eval-results-v2')
bitext-customer-support-llm-chatbot-training-datasetdemon_attack_fixed_latency_6_200ep_7k2steps_ghost15Ballbusting-StoriesThis dataset contains ballbusting stories written by consenting Reddit users that have been edited and reformatted for machine learning purposes. Note that the score field on many stories will likely be out of date.
suicidal_finetunedemon_attack_ghost60_200epflappy_fixed_latency_3_200ep_7k2steps_ghost15GhostWriter
GhostWriterHidden AI-Generated Texts Over Multiple Languages, Domains and Generators
About
GhostWriter is a multi-domain, multi-generator, multi-lingual dataset of human-written and partially or fully AI-generated texts.
The texts are sourced from seven domains and two languages (English and German) and were generated with a variety of open-source and commercial LLMs.
Abstract
The advent of Transformer-based Large Language Models (LLMs) has led to an… See the full description on the dataset page: https://huggingface.co/datasets/TheItCrOw/GhostWriter.ghostbuster_reuter_reducedcybersec-fact-recall
Cybersec Fact-Recall Benchmark (GhostLM v2)
Free-form short-answer benchmark for small cybersecurity language
models. Built and used by the GhostLM
project as the truth metric for the ghost-base v1.0 acceptance gate.
Why this exists
Multiple-choice cybersec benchmarks like CTIBench and SecQA reward
register matching (the model picks the option that "looks like" a
security answer) as much as actual factual recall. A small from-
scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.ai-ghost-2026ghost_prompts
Dataset Card for "ghost_prompts"
More Information needed
ghost_rate_newghostmaneThis dataset is designed to generate lyrics with HuggingArtists.
