CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Miaowuawa /ChineseNovels 中文小说数据集 包含内容: 网游/系统/重生 言情小说 同人/耽美小说 科幻小说 军事小说 以上加起来共4万本左右 海棠文学城小说:约1000本(未清洗) texttext-generation1K<n<10K22 likes83 downloads2y agoHugging Face02Miaow-Lab /RUT-Bench Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions This repository contains the RUT-Bench benchmark, which consists of 1638 test samples for evaluating LLM agents under realistic user interactions. Paper: Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions Code: GitHub Collection: Hugging Face Collection 📖 Overview RUT-Bench is a dedicated benchmark designed to assess… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/RUT-Bench.tabulartext-generation1K<n<10K1 likes31 downloads4mo agoHugging Face03miaomiao64 /tb-explore17-mcode-m3-harness-variancegated Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance Three complete 17-task runs of the same dataset ref with the same agent and model, differing only in execution substrate and concurrency, plus one isolated rerun. The point of the bundle is not the resolve rate — it is how much the resolve rate moves when nothing about the task or the model changes. Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.tabulartext-generationn<1K0 likes31 downloads15d agoHugging Face04Miaow-Lab /SSAE-Dataset Dataset Card This is the official dataset repository for the paper "Step-Level Sparse Autoencoder for Reasoning Process Interpretation". Paper: Arxiv Code: GitHub Collection: HuggingFace Dataset Overview The repository hosts three distinct datasets covering the domains of mathematical reasoning and code generation. Each subset is pre-partitioned into training and validation splits to facilitate reproducible experiments. 1. GSM8K (Math) Description: A… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/SSAE-Dataset.texttext-generation1M<n<10M0 likes26 downloads7mo agoHugging Face05matthewwicker /shadow-llm-mia-signals Shadow LLM MIA Signals (OLMo-2-1B) Membership Inference Attack (MIA) signal tensors extracted from 128 shadow models fine-tuned from allenai/OLMo-2-0425-1B. Overview This dataset enables research on membership inference attacks against large language models. Each of 128 shadow models was trained on a different random subset of 64 out of 128 candidate documents from the OLMo-mix-1124 pretraining dataset. For each (model, document) pair, we extracted softmax prediction… See the full description on the dataset page: https://huggingface.co/datasets/matthewwicker/shadow-llm-mia-signals.tabulartext-generationn<1K0 likes25 downloads5mo agoHugging Face06miaomiao64 /swe-explore-find-dev100-runsgated SWE-Explore find — dev-100 evidence runs Five complete 100-trial runs of the find (fault-localization) arm of SWE-Explore, kept because each one is load-bearing evidence for a specific claim about the scoring fixes on branch fix/swe-explore-find-scoring of harness_bench. Every run is 100 trials of the same dev-100 find subset, run through Harbor with the swe_explore.agents:PiSut agent. index.jsonl has one row per trial (500 rows); the full raw Harbor trial directories are in… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/swe-explore-find-dev100-runs.tabulartext-generationn<1K0 likes24 downloads15d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.