CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes15k downloads4mo agoHugging Face02PranavViswanath /jlens-gp-auditbench The AuditBench J-lens corpus: 80-layer activations and gradient-pursuit readouts Everything needed to redo J-space interpretability work on the 84 AuditBench model organisms (14 hidden behaviors x 2 instillation methods x 3 adversarial-training levels) without a GPU harvest: the raw bf16 residual stream at all 80 layers for every recorded token, and a gradient-pursuit J-lens decomposition at every (position, layer) site. The organisms are Llama-3.3-70B-Instruct with an… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-gp-auditbench.tabulartext-generation10M<n<100M0 likes10k downloads2mo agoHugging Face03Stage-jh-monitor /appworld-qwen35-4b-total-237-audited-jh-epoch2 appworld-qwen35-4b-total-237-audited-jh-epoch2 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3640625 Action score: 0.4328125 Valid samples: 320/320 tabularn<1K0 likes4k downloads14d agoHugging Face04Stage-jh-monitor /appworld-qwen35-4b-total-237-audited-jh-epoch6 appworld-qwen35-4b-total-237-audited-jh-epoch6 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.37578125 Action score: 0.421875 Valid samples: 320/320 tabularn<1K0 likes4k downloads14d agoHugging Face05Stage-jh-monitor /appworld-qwen35-4b-total-237-audited-jh-epoch8 appworld-qwen35-4b-total-237-audited-jh-epoch8 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.384375 Action score: 0.4390625 Valid samples: 320/320 tabularn<1K0 likes4k downloads14d agoHugging Face06PranavViswanath /auditbench-activations-jlens-NLA AuditBench activations, J-lens readouts and NLA verbalizations Every token of every AuditBench prompt and every model response, from meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell. Responses were regenerated greedily and run to the model's own stopping point rather than truncated at a fixed length, and the activations, readouts and verbalizations cover the prompt as well as the response. 84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.tabulartext-generation100M<n<1B0 likes3.8k downloads2mo agoHugging Face07dougalldeepmind /2026-09-14-dataset-refresh-revised-pilot-audit Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only field value experiment Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only date_generated 20260914_230322 constitution constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.textn<1K0 likes1.5k downloads11d agoHugging Face08dougalldeepmind /2026-09-14-dataset-refresh-pilot-audit Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only field value experiment Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only date_generated 20260914_224408 constitution constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.textn<1K0 likes1.4k downloads11d agoHugging Face09mwritescode /slither-audited-smart-contractsThis dataset contains source code and deployed bytecode for Solidity Smart Contracts that have been verified on Etherscan.io, along with a classification of their vulnerabilities according to the Slither static analysis framework.texttext-classification100K<n<1M61 likes771 downloads4y agoHugging Face10anonymous-structured-agent /structured-file-audit-benchmark Paper Data Release This directory contains the benchmark dataset and evaluation scripts accompanying the ACL submission: the three data splits (SC-Flat, SC-Book, SC-Pro) and the code needed to score them. Contents datasets/ Benchmark data and per-task manifests for the three paper-facing splits. datasets/sc_flat/data SC-Flat is derived from DaBench, augmented with a replayable perturbation injected into each task's input artifact. Each task… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-structured-agent/structured-file-audit-benchmark.texttable-question-answering1 likes625 downloads2mo agoHugging Face11leeroy-jankins /CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset Dataset Summary This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance. The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.documentquestion-answering0 likes587 downloads3mo agoHugging Face12damo-da /oag-nepal-audit-reports OAG Nepal Audit Reports — Nepali transcripts and ruled tables Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements. The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.documenttext-retrieval10M<n<100M0 likes579 downloads23d agoHugging Face13introspection-auditing /llama-rare-mo-training-datatext1M<n<10M0 likes480 downloads6mo agoHugging Face14kalomaze /glm52-usersim-two-pass-gemma-audit-v1 GLM-5.2 Usersim Two-Pass Gemma Audit v1 This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k. The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length. Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.tabulartext-generation100K<n<1M4 likes436 downloads1mo agoHugging Face15PranavViswanath /auditbench-viz-jsontabularn<1K0 likes427 downloads2mo agoHugging Face16auditing-contextual-privacy /new_new_audit_gpt54mini_claude46_k493_n200_b005textn<1K0 likes415 downloads3mo agoHugging Face17introspection-auditing /llama-backdoor-mo-training-datatext100K<n<1M0 likes400 downloads6mo agoHugging Face18DivyaApp /tiny-audits Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DivyaApp/tiny-audits.imagemask-generation10K<n<100K0 likes398 downloads1y agoHugging Face19introspection-auditing /llama-benign-mo-training-datatext100K<n<1M0 likes386 downloads6mo agoHugging Face20dougalldeepmind /2026-09-15-dataset-refresh-incomplete-audit INCOMPLETE RESEARCH AUDIT — NOT A TRAINING DATASET field value experiment Incomplete retained research pools: moral low stakes has 706 rows (10 short of 716: t1=2, t4=2, t6=1, t7=4, t8=1); original craft nonmoral has 631 rows (85 short: t1=7, t2=11, t3=10, t4=8, t5=12, t6=12, t7=5, t8=8, t9=12). Shared spend exposure is $249.2113677 of $250, with no active calls or uncertain reservations. Both pools are byte-identical subsets of 708/634-row snapshots that passed… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-15-dataset-refresh-incomplete-audit.textn<1K0 likes363 downloads10d agoHugging Face21introspection-auditing /llama-quirk-mo-training-datatext100K<n<1M0 likes334 downloads6mo agoHugging Face22auditing-contextual-privacy /new_audit_gpt54mini_claude46_k493_n200_b005tabular100K<n<1M0 likes325 downloads3mo agoHugging Face23wicai24 /api_audit_dataThis repository contains code for auditing Large Language Models (LLMs) to verify service integrity, as described in the paper Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs. Github repository: https://github.com/willsdca/llm_api_audit text100K<n<1M1 likes297 downloads1y agoHugging Face24introspection-auditing /llama-harmful-mo-training-datatext100K<n<1M0 likes292 downloads6mo agoHugging Face25introspection-auditing /harmful-benign-mo-eval-datatext10K<n<100K0 likes285 downloads5mo agoHugging Face26ShreyashDhoot /internvl-auditor-v2image10K<n<100K0 likes284 downloads5mo agoHugging Face27LeoJiangOR /FrontierOR-Audited-180 FrontierOR Audited 180 This is an evaluator-oriented derivative of SmartOR/FrontierOR, pinned to upstream commit 37ccd8b6dca3bf7f4e0c58941a6ed156832a6d9e. The released descriptions, formulations, Gurobi implementations, solution schemas, reference solutions, and feasibility checkers were audited and repaired as one evaluation contract. Final status Gate Result Evaluator-ready directory IDs 180/180 Independent canonical cases 179 Transparent… See the full description on the dataset page: https://huggingface.co/datasets/LeoJiangOR/FrontierOR-Audited-180.tabularothern<1K0 likes279 downloads20d agoHugging Face28NobleJackal /NOMOS-GEO-Audit-Protocol NOMOS GEO Audit Protocol A repeatable way to test what AI systems say about an organisation and whether the evidence supports it GEO means Generative Engine Optimization. This six-language candidate protocol turns that discipline into an auditable process using the GEO-1000 method, canonical questions, truth packs, evidence requirements, scoring logic, correction steps and revalidation records. Start reading: Open the English PDF · Choose one of six languages · Cite… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GEO-Audit-Protocol.documentn<1K1 likes276 downloads15d agoHugging Face29introspection-auditing /llama-problematic-mo-training-datatext100K<n<1M0 likes270 downloads6mo agoHugging Face30introspection-auditing /llama-heuristic-mo-training-datatext10K<n<100K0 likes254 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.