CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shijunhao /Fable-5-traces Glint Research Dataset Card Fable 5 Pi Agent Traces A compact, high-signal corpus of Fable 5 coding-agent traces converted into Hugging Face Agent Traces / Pi-compatible sessions for Data Studio inspection, tool-use policy learning, and reasoning/action distillation. Primary Config pi_agent/train Agent Trace preview enabled 4,665 Pi trace sessions 60 source sessions 3,799 tool… See the full description on the dataset page: https://huggingface.co/datasets/shijunhao/Fable-5-traces.tabulartext-generation1K<n<10K1 likes692 downloads3mo agoHugging Face02Shiki42 /piperx-workpiece-storage-0909-62ep-raw piperx-workpiece-storage-0909-62ep-raw 61 retained manually collected episodes recorded with EvoMind on 2026-09-09 (UTC+8). Task: put the copper screw into the left box and the black sleeves into the right box. LeRobot v3.0; 30 FPS; 53,347 frames; 1,778.2333 seconds. Robot: bi_piperx_follower. Three original 640x480 RGB video views: left_wrist, right_wrist, right_environment_1. Merged in chronological session order. Original episode 26 (19 frames) was removed on 2026-09-15.… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/piperx-workpiece-storage-0909-62ep-raw.tabularroboticsn<1K0 likes338 downloads10d agoHugging Face03shimo4228 /authorship-strategy Authorship Strategy — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Authorship Strategy research line — a normative framework, tactical catalog, and empirical baseline for authorship strategy under AI-mediated diffusion. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the Authorship Strategy GitHub repository. It is provided here for LLM training pipelines, knowledge-graph crawlers, and AI research… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/authorship-strategy.tabularn<1K1 likes162 downloads1mo agoHugging Face04shimo4228 /contemplative-agent Contemplative Agent — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Contemplative Agent — an autonomous CLI agent (Python) built around four architectural principles (structural capability limitation, minimal dependency, cyclic knowledge maintenance, memory dynamics with decay) and, optionally, the four contemplative axioms from Laukkonen et al. (2025) as a behavioral preset. What this dataset is This dataset is a mirror of the… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/contemplative-agent.tabularn<1K1 likes127 downloads12d agoHugging Face05shimo4228 /agent-knowledge-cycle Agent Knowledge Cycle (AKC) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Knowledge Cycle (AKC) — a six-phase bidirectional growth loop in which agent behavior and the operator's judgment co-develop over time, sustaining intent alignment that tests cannot check on their own. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AKC GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-knowledge-cycle.tabularn<1K1 likes126 downloads24d agoHugging Face06shimo4228 /agent-attribution-practice Agent Attribution Practice (AAP) — Knowledge Graph JSON-LD knowledge graph encoding the concept layer of the Agent Attribution Practice (AAP) research line — a harness-neutral set of Architecture Decision Records (ADRs) and a problem-space diagnostic frame on accountability distribution in autonomous AI agents. What this dataset is This dataset is a mirror of the graph.jsonld file at the root of the AAP GitHub repository. It is provided here for LLM training… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/agent-attribution-practice.tabularn<1K2 likes108 downloads1mo agoHugging Face07Shikha180224 /dexfluence-indian-creator-index Dexfluence Indian Creator Index Verified Indian influencer dataset across Instagram, YouTube, and TikTok with engagement rates, follower tier, niche classification, and authenticity scores. Dataset summary 141,000+ verified Indian creators indexed across Instagram, YouTube, and TikTok Top 5,000 by follower count included in this Hugging Face mirror (CC-BY 4.0) Each record includes: handle, name, platform, niche, follower count, engagement rate, country… See the full description on the dataset page: https://huggingface.co/datasets/Shikha180224/dexfluence-indian-creator-index.tabulartabular-classification1K<n<10K0 likes106 downloads4mo agoHugging Face08wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face09shimbaaa /shimbabomb-benchmark ShimbaBomb Interpreter Extreme Stress Benchmark Overview Benchmark results from extreme stress testing of the ShimbaBomb (SB) v1.11.0 interpreter — an English-like scripting language that compiles to native C. The benchmark suite spawns CPU_Logical_Cores * 2 (or higher) threads to saturate the interpreter engine, running diverse SB scripts simultaneously for 30-second sustained windows per phase. System Parameter Value Platform Windows… See the full description on the dataset page: https://huggingface.co/datasets/shimbaaa/shimbabomb-benchmark.tabularn<1K0 likes85 downloads25d agoHugging Face10Shiki42 /ctr-pick-dual-bottles-original-20260919 Pick Dual Bottles Original — shared50 scene cohort This LeRobot v3 release contains 50 successful simulated demonstrations and 8,185 action rows at25FPS. Every source seed occurs exactly once. The source seed set matches the current CTR Q1–Q3 Concurrent, CTR, Sequential, Mixed, Left-first and Right-first datasets. Pair by retime.source_seed, not episode index: composition datasets may have different ordering. Mask limitation: retime.left_idle and retime.right_idle are boolean… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/ctr-pick-dual-bottles-original-20260919.tabularn<1K0 likes75 downloads5d agoHugging Face11cloudfan /intern-shiguan-lerobot shiguan: robot demonstrations Instruction: Pick up the test tube from the left side of the rack and insert it into the hole at the right end of the rack. LeRobot v3.0 dataset: 51 episodes, 26665 frames, nominal 30 Hz. Robot schema: intern_gello_7dof_robotiq. Original source recordings are retained by the owner. Use from lerobot.datasets.lerobot_dataset import LeRobotDataset dataset = LeRobotDataset("cloudfan/intern-shiguan-lerobot", video_backend="torchcodec")… See the full description on the dataset page: https://huggingface.co/datasets/cloudfan/intern-shiguan-lerobot.tabularn<1K0 likes70 downloads4d agoHugging Face12shivam039-dev /context-management-bench context-management-bench Live on the Hub: huggingface.co/datasets/shivam039-dev/context-management-bench Realistic context-management scenarios for testing/benchmarking eviction strategies (drop-oldest, sliding-window, priority, summarization), pinned-message preservation, and tool-call/tool-result atomicity in multi-turn LLM conversations. Dataset Summary Every conversation in this dataset was generated deterministically and then run through the real… See the full description on the dataset page: https://huggingface.co/datasets/shivam039-dev/context-management-bench.tabularn<1K0 likes57 downloads20d agoHugging Face13tomyimkc /repro-sample-complexity-bounds-for-robust-mean-estimation-with-mean-shift-contaminatio-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes41 downloads2mo agoHugging Face14Shiki42 /screw_retimed_parallel_23_no_fastforward_20260807tabularn<1K0 likes40 downloads2mo agoHugging Face15duyle2408 /levir-ship-copy-paste-load-adaptive-mosaictabularn<1K0 likes35 downloads12d agoHugging Face16ShiyuanHuang /Text2Space Text2Space Synthetic dataset of 20,000 spatial reasoning instances. Each instance pairs a natural-language description of a 2D layout with three ASCII renderings of the same scene and a query about the relative position of two objects. Designed to train and evaluate language and vision-language models on spatial reasoning. Companion dataset for the paper Learning to Draw ASCII Improves Spatial Reasoning in Language Models (arXiv:2604.14641). Quick Look {… See the full description on the dataset page: https://huggingface.co/datasets/ShiyuanHuang/Text2Space.tabularquestion-answering10K<n<100K0 likes32 downloads5mo agoHugging Face17duyle2408 /levir-ship-copy-paste-mosaictabularn<1K0 likes32 downloads13d agoHugging Face18open-llm-leaderboard /shivam9980__NEPALI-LLM-detailsgated Dataset Card for Evaluation run of shivam9980/NEPALI-LLM Dataset automatically created during the evaluation run of model shivam9980/NEPALI-LLM The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shivam9980__NEPALI-LLM-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face19Shirali /reddit-demotabularn<1K0 likes20 downloads4y agoHugging Face20shijunju /fincen_all_questions_5versions About These question-answer pairs are created using published pdf documents at fincen.gov. Each question has 5 paraphased versions differentiated by column "question_version" (the first versions (No. 4) are at the end of the datafile). The data is used to fine-tune Gemma-2b and Gemma-7b listed here shijunju/gemma_7b_finRisk_r10_4VersionQ shijunju/gemma_7b_finRisk_r6_4VersionQ shijunju/gemma_7b_finRisk_r6_3VersionQ shijunju/gemma_2b_finRisk Number of rows: 14,550 Author: Shijun… See the full description on the dataset page: https://huggingface.co/datasets/shijunju/fincen_all_questions_5versions.tabular10K<n<100K0 likes16 downloads2y agoHugging Face21ShiroOnigami23 /securehealthiot-disease-dataset SecureHealthIoT Cleaned Dataset Cleaned symptom-disease dataset generated by Kaggle kernel run: aryansingh21fd/securehealthiot-disease-trainer-v1. tabulartabular-classificationn<1K1 likes15 downloads6mo agoHugging Face22shivani-kerai /vqa_training_annotationstabular100K<n<1M0 likes13 downloads2y agoHugging Face23shivani-kerai /vqa_training_questionstabular100K<n<1M0 likes12 downloads2y agoHugging Face24Shivahoody007 /Phishing_Link_Pattern_Dataset Phishing Link Pattern Dataset Overview This dataset provides a comprehensive collection of URLs labeled as either legitimate or phishing, designed for machine learning, cybersecurity analysis, and penetration testing. It includes 1000 entries (IDs 1–1000) covering popular brands across multiple top-level domains (TLDs) such as .es, .de, and .co.uk. The dataset captures advanced features like domain entropy, subdomain count, and suspicious keywords to aid in… See the full description on the dataset page: https://huggingface.co/datasets/Shivahoody007/Phishing_Link_Pattern_Dataset.tabular1K<n<10K0 likes12 downloads3mo agoHugging Face25open-llm-leaderboard /ValiantLabs__Llama3.1-70B-ShiningValiant2-detailsgated Dataset Card for Evaluation run of ValiantLabs/Llama3.1-70B-ShiningValiant2 Dataset automatically created during the evaluation run of model ValiantLabs/Llama3.1-70B-ShiningValiant2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ValiantLabs__Llama3.1-70B-ShiningValiant2-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face26open-llm-leaderboard /shivam9980__mistral-7b-news-cnn-merged-detailsgated Dataset Card for Evaluation run of shivam9980/mistral-7b-news-cnn-merged Dataset automatically created during the evaluation run of model shivam9980/mistral-7b-news-cnn-merged The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shivam9980__mistral-7b-news-cnn-merged-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face27shizhuo2 /qwen3-1.7b-features-similar-k100000tabular100K<n<1M0 likes9 downloads7mo agoHugging Face28open-llm-leaderboard /ValiantLabs__Llama3-70B-ShiningValiant2-detailsgated Dataset Card for Evaluation run of ValiantLabs/Llama3-70B-ShiningValiant2 Dataset automatically created during the evaluation run of model ValiantLabs/Llama3-70B-ShiningValiant2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ValiantLabs__Llama3-70B-ShiningValiant2-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face29open-llm-leaderboard /ValiantLabs__Llama3.2-3B-ShiningValiant2-detailsgated Dataset Card for Evaluation run of ValiantLabs/Llama3.2-3B-ShiningValiant2 Dataset automatically created during the evaluation run of model ValiantLabs/Llama3.2-3B-ShiningValiant2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ValiantLabs__Llama3.2-3B-ShiningValiant2-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face30open-llm-leaderboard /NLPark__Shi-Ci-Robin-Test_3AD80-detailsgated Dataset Card for Evaluation run of NLPark/Shi-Ci-Robin-Test_3AD80 Dataset automatically created during the evaluation run of model NLPark/Shi-Ci-Robin-Test_3AD80 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NLPark__Shi-Ci-Robin-Test_3AD80-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.