datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GeoSeek
GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristic
Modi Jin1 · Yiming Zhang1 · Boyuan Sun1 · Dingwen Zhang2 · Mingming Cheng1 · Qibin Hou1†
1VCIP, Nankai University 2 School of Automation, Northwestern Polytechnical University
†Corresponding author
English | 简体中文
We introduce GeoSeek train GeoAgent, which is a new geolocation dataset comprising:
GeoSeek-CoT (10k): High-quality chain-of-thought data labeled by geography experts and professional… See the full description on the dataset page: https://huggingface.co/datasets/ghost233lism/GeoSeek.OpenHermes-NoRefusal-95K
OpenHermes-NoRefusal-95K
A refusal-free instruction-tuning dataset: 95,401 single-turn conversations derived from
teknium/OpenHermes-2.5, filtered so
that zero assistant responses contain refusals, hedging boilerplate, or
"as an AI language model" disclaimers.
Why this exists
The usual way to get a model that doesn't refuse is to train it on aligned data and then
remove the alignment afterwards — refusal-direction ablation, weight editing, abliteration.
That works… See the full description on the dataset page: https://huggingface.co/datasets/ghost-actual/OpenHermes-NoRefusal-95K.deepwiki-eval-results-v2
GhostScientist/deepwiki-eval-results-v2
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('GhostScientist/deepwiki-eval-results-v2')
Ballbusting-StoriesThis dataset contains ballbusting stories written by consenting Reddit users that have been edited and reformatted for machine learning purposes. Note that the score field on many stories will likely be out of date.
cybersec-fact-recall
Cybersec Fact-Recall Benchmark (GhostLM v2)
Free-form short-answer benchmark for small cybersecurity language
models. Built and used by the GhostLM
project as the truth metric for the ghost-base v1.0 acceptance gate.
Why this exists
Multiple-choice cybersec benchmarks like CTIBench and SecQA reward
register matching (the model picks the option that "looks like" a
security answer) as much as actual factual recall. A small from-
scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.ai-ghost-2026demo-experiment-two
Demo Experiment
A demo experiment to test the builder flow with various question types.
Dataset Overview
Property
Value
Run ID
036df3a6-6f8c-45fb-979d-be4f990bd0bf
Status
completed
Created
12/23/2025, 9:11:08 PM
Generator
LocalBench v0.1.0
Statistics
Metric
Value
Total Generations
10
Successful
10 (100.0%)
Failed
0
Average Latency
2803ms
Total Duration
25.2s
Configuration
Models… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/demo-experiment-two.Buffett_Agent_DataThis repository contains the datasets used for a use case demo of FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets.
pacman_eval_two_ghost_braidthin_random100_20260817
HF Two-Ghost braided/thinned — deterministic random 100
This is the exact 100-episode Pacman evaluation shard used for the final UWM trajectories.
Source dataset: pacman_2d_easy_g9f8_two_ghost_braidthin_oracle_success_test_v1_20260817
Sampling: Python MT19937 without replacement, seed 2026081702
Records: 100
JSONL SHA-256: fbaf50a9a9799384b2cc1daf5c4c5c7e81c3fed8a23b01cc3b5b53747f3ed034
Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_eval_two_ghost_braidthin_random100_20260817.FinalFinetuningDatasetFinetuningDatasetToxic-RU-2FinetuningDataset-v2ghost-x__ghost-8b-beta-1608-details
Dataset Card for Evaluation run of ghost-x/ghost-8b-beta-1608
Dataset automatically created during the evaluation run of model ghost-x/ghost-8b-beta-1608
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ghost-x__ghost-8b-beta-1608-details.Habr_ru_en_text2emoji
