CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghost233lism /GeoSeek GeoAgent: Learning to Geolocate Everywhere with Reinforced Geographic Characteristic Modi Jin1 · Yiming Zhang1 · Boyuan Sun1 · Dingwen Zhang2 · Mingming Cheng1 · Qibin Hou1† 1VCIP, Nankai University 2 School of Automation, Northwestern Polytechnical University †Corresponding author English | 简体中文 We introduce GeoSeek train GeoAgent, which is a new geolocation dataset comprising: GeoSeek-CoT (10k): High-quality chain-of-thought data labeled by geography experts and professional… See the full description on the dataset page: https://huggingface.co/datasets/ghost233lism/GeoSeek.image10K<n<100K2 likes226 downloads3mo agoHugging Face02ghost-actual /OpenHermes-NoRefusal-95K OpenHermes-NoRefusal-95K A refusal-free instruction-tuning dataset: 95,401 single-turn conversations derived from teknium/OpenHermes-2.5, filtered so that zero assistant responses contain refusals, hedging boilerplate, or "as an AI language model" disclaimers. Why this exists The usual way to get a model that doesn't refuse is to train it on aligned data and then remove the alignment afterwards — refusal-direction ablation, weight editing, abliteration. That works… See the full description on the dataset page: https://huggingface.co/datasets/ghost-actual/OpenHermes-NoRefusal-95K.texttext-generation10K<n<100K1 likes179 downloads27d agoHugging Face03GhostScientist /deepwiki-eval-results-v2 GhostScientist/deepwiki-eval-results-v2 Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code: https://github.com/huggingface/ml-intern Usage from datasets import load_dataset dataset = load_dataset('GhostScientist/deepwiki-eval-results-v2') textn<1K0 likes37 downloads12d agoHugging Face04Ghost314 /Ballbusting-StoriesThis dataset contains ballbusting stories written by consenting Reddit users that have been edited and reformatted for machine learning purposes. Note that the score field on many stories will likely be out of date. textn<1K0 likes31 downloads3y agoHugging Face05Ghostgim /cybersec-fact-recall Cybersec Fact-Recall Benchmark (GhostLM v2) Free-form short-answer benchmark for small cybersecurity language models. Built and used by the GhostLM project as the truth metric for the ghost-base v1.0 acceptance gate. Why this exists Multiple-choice cybersec benchmarks like CTIBench and SecQA reward register matching (the model picks the option that "looks like" a security answer) as much as actual factual recall. A small from- scratch model can hit 28-30% on those without… See the full description on the dataset page: https://huggingface.co/datasets/Ghostgim/cybersec-fact-recall.texttext-generationn<1K0 likes27 downloads5mo agoHugging Face06gemmozero /ai-ghost-2026textn<1K0 likes27 downloads4d agoHugging Face07GhostScientist /demo-experiment-two Demo Experiment A demo experiment to test the builder flow with various question types. Dataset Overview Property Value Run ID 036df3a6-6f8c-45fb-979d-be4f990bd0bf Status completed Created 12/23/2025, 9:11:08 PM Generator LocalBench v0.1.0 Statistics Metric Value Total Generations 10 Successful 10 (100.0%) Failed 0 Average Latency 2803ms Total Duration 25.2s Configuration Models… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/demo-experiment-two.tabularn<1K0 likes20 downloads9mo agoHugging Face08ghostof0days /Buffett_Agent_DataThis repository contains the datasets used for a use case demo of FinLoRA: Benchmarking LoRA Methods for Fine-Tuning LLMs on Financial Datasets. text10K<n<100K1 likes19 downloads1y agoHugging Face09novastar111 /pacman_eval_two_ghost_braidthin_random100_20260817 HF Two-Ghost braided/thinned — deterministic random 100 This is the exact 100-episode Pacman evaluation shard used for the final UWM trajectories. Source dataset: pacman_2d_easy_g9f8_two_ghost_braidthin_oracle_success_test_v1_20260817 Sampling: Python MT19937 without replacement, seed 2026081702 Records: 100 JSONL SHA-256: fbaf50a9a9799384b2cc1daf5c4c5c7e81c3fed8a23b01cc3b5b53747f3ed034 Trajectories: https://huggingface.co/datasets/novastar111/uwm_paper_eval_trajectories… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_eval_two_ghost_braidthin_random100_20260817.tabularn<1K0 likes11 downloads1mo agoHugging Face10GhostMopey115 /FinalFinetuningDatasettext10K<n<100K0 likes9 downloads1y agoHugging Face11GhostMopey115 /FinetuningDatasettext1K<n<10K0 likes8 downloads1y agoHugging Face12ghostdist /Toxic-RU-2text100K<n<1M1 likes8 downloads11mo agoHugging Face13GhostMopey115 /FinetuningDataset-v2text1K<n<10K0 likes7 downloads1y agoHugging Face14open-llm-leaderboard /ghost-x__ghost-8b-beta-1608-detailsgated Dataset Card for Evaluation run of ghost-x/ghost-8b-beta-1608 Dataset automatically created during the evaluation run of model ghost-x/ghost-8b-beta-1608 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ghost-x__ghost-8b-beta-1608-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face15ghostbim21 /Habr_ru_en_text2emojitext10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.