CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01arsentev-ai /context-ucurve-coding-agents Context U-curve: 36 coding-agent runs under six context-clearing policies How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report "Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents" (Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668). A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.tabularn<1K0 likes77 downloads7d agoHugging Face02AgentSkillPrivacy /SkillLeakBench SkillLeakBench A credential-leakage benchmark for LLM agent skills, from the ASE 2026 paper How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study. 📄 Paper: https://arxiv.org/abs/2604.03070 💻 Code & detection pipeline: https://github.com/AgentSkillsPrivacy/SkillLeakBench 📦 Archive: https://doi.org/10.5281/zenodo.19367969 Dataset summary We collected 170,226 skills from SkillsMP and analyzed a 17,022-skill sample with static secret extraction… See the full description on the dataset page: https://huggingface.co/datasets/AgentSkillPrivacy/SkillLeakBench.tabulartext-classification1K<n<10K0 likes72 downloads3mo agoHugging Face03JacobiusMakes /agent-shoppable-census Agent-Shoppable Census Can an AI agent buy a ring here? This is a reproducible census of 100 US online sellers of engagement rings and diamond jewelry. It tests public agent-readiness signals and separates them into two tiers so that default ecommerce-platform features do not inflate the result. The September 2026 snapshot reports 18 Tier A sellers with merchant-built agent signals, 52 Tier B sellers with platform-inherited signals, and 30 sellers that were unreachable to the… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/agent-shoppable-census.tabularn<1K0 likes58 downloads18d agoHugging Face04values-md /when-agents-act Dataset Card for "When Agents Act" Dataset Summary This dataset contains 702 ethical decision judgements from 9 frontier LLMs (Claude Opus 4.5, GPT-5, GPT-5 Nano, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 3 Pro, Gemini 2.5 Flash, Grok-4, Grok-4 Fast) across 10 rigorously curated AI-relevant ethical dilemmas. Models were tested in both theory mode (hypothetical reasoning) and action mode (tool-enabled agents believing actions would execute). Key Finding: Models reverse… See the full description on the dataset page: https://huggingface.co/datasets/values-md/when-agents-act.tabulartext-classificationn<1K1 likes51 downloads10mo agoHugging Face05dingjiacheng /wc2026-agents WC2026-Agents: ChatGPT vs Claude vs Gemini vs Grok on the 2026 FIFA World Cup WC2026-Agents is a contamination-free benchmark in which four frontier LLMs act as autonomous forecasting agents over the entire 2026 FIFA World Cup (104 matches, 11 June to 19 July 2026). Each agent (Claude Opus 4.8, ChatGPT GPT-5.5 with high reasoning, Gemini 3.1 Pro, and Grok Expert Mode) ran an identical search, act, reflect loop per match: search the web, commit to a 1X2 (team A win / draw / team… See the full description on the dataset page: https://huggingface.co/datasets/dingjiacheng/wc2026-agents.tabulartabular-classificationn<1K0 likes50 downloads9d agoHugging Face06crackedvibe /agent-scraper-price-index-2026-09 Agent scraper price index, September 2026 The rows behind the report Agent scraper price index, September 2026 on cracked.ai: what 1,000 results cost for 18 of the most requested scraping and search capabilities, tool by tool, through three routes. Method For each capability, every live candidate tool on Cracked is listed with three prices for 1,000 results: the price billed on Cracked's own provider account (cost_1000_cracked_usd, provider price plus the $0.001… See the full description on the dataset page: https://huggingface.co/datasets/crackedvibe/agent-scraper-price-index-2026-09.tabulartabular-regressionn<1K0 likes46 downloads20d agoHugging Face07ScareRezume /agent-sandbox-negotiation-benchmark Agent Sandbox Negotiation Benchmark v1 Overview A dataset of simulated multi-agent negotiations generated using the open-source Agent Sandbox framework. This dataset captures the final negotiation outcomes, turn depths, strategy alignments, and agreed prices of local LLMs (Llama-3 and Mistral) engaged in intense, adversarial price negotiations at massive scale. Dataset Statistics Simulations: 24,122 Strategies: 4 (Balanced, Aggressive, Conservative, Adaptive)… See the full description on the dataset page: https://huggingface.co/datasets/ScareRezume/agent-sandbox-negotiation-benchmark.tabulartext-generation10K<n<100K1 likes22 downloads7mo agoHugging Face08rsoft-latam /erc8004-simulated-agents ERC-8004 Simulated Agents — labeled synthetic dataset (6,000 agents) ⚠️ This dataset is fully synthetic. No public labeled dataset of malicious ERC-8004 agents exists (the standard reached mainnet in 2026 and exposes no trust label), so this dataset simulates the feature distributions the three ERC-8004 registries would expose, for training/evaluating trustworthiness models. For real on-chain data see the companion Base mainnet census. Composition 6,000 agents, 1… See the full description on the dataset page: https://huggingface.co/datasets/rsoft-latam/erc8004-simulated-agents.tabular1K<n<10K0 likes11 downloads2mo agoHugging Face09eamag /cryoprotective-agents eamag/cryoprotective-agents Auto-generated CSV using papers2dataset tool. THIS IS A PROOF OF CONCEPT, done with a free LLM and not double checked texttabular-classificationn<1K1 likes8 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.