CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jedisct1 /security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models. These traces focus on security audits of opensource software. Sharing traces with Swival Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session: swival "Fix the login bug" --trace-dir traces/ Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.tabulartext-generation10K<n<100K17 likes15k downloads4mo agoHugging Face02PranavViswanath /jlens-gp-auditbench The AuditBench J-lens corpus: 80-layer activations and gradient-pursuit readouts Everything needed to redo J-space interpretability work on the 84 AuditBench model organisms (14 hidden behaviors x 2 instillation methods x 3 adversarial-training levels) without a GPU harvest: the raw bf16 residual stream at all 80 layers for every recorded token, and a gradient-pursuit J-lens decomposition at every (position, layer) site. The organisms are Llama-3.3-70B-Instruct with an… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-gp-auditbench.tabulartext-generation10M<n<100M0 likes10k downloads2mo agoHugging Face03PranavViswanath /auditbench-activations-jlens-NLA AuditBench activations, J-lens readouts and NLA verbalizations Every token of every AuditBench prompt and every model response, from meta-llama/Llama-3.3-70B-Instruct (revision 6f6073b423013f6a7d4d9f39144961bfbfbc386b) with one LoRA adapter per cell. Responses were regenerated greedily and run to the model's own stopping point rather than truncated at a fixed length, and the activations, readouts and verbalizations cover the prompt as well as the response. 84 cells across 14… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/auditbench-activations-jlens-NLA.tabulartext-generation100M<n<1B0 likes3.8k downloads2mo agoHugging Face04kalomaze /glm52-usersim-two-pass-gemma-audit-v1 GLM-5.2 Usersim Two-Pass Gemma Audit v1 This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k. The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length. Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.tabulartext-generation100K<n<1M4 likes436 downloads1mo agoHugging Face05PranavViswanath /jlens-nla-auditbench J-lens and NLA readouts for the 84 AuditBench organism cells Per-token-position interpretability readouts over all 84 AuditBench model organisms (14 hidden behaviors × 2 instillation methods × 3 adversarial-training levels), harvested from Llama-3.3-70B-Instruct with each organism's LoRA adapter active. For every stored prompt the release carries: the full prompt + generated-response token sequence (exact token ids as run), the Jacobian-lens ("J-lens") top-50 readout at every… See the full description on the dataset page: https://huggingface.co/datasets/PranavViswanath/jlens-nla-auditbench.tabulartext-generation10M<n<100M0 likes168 downloads2mo agoHugging Face06leohachico /audit-findings-dataset Smart Contract Audit Findings This is raw, semi-structured data — not a ready-to-train dataset. It still requires further cleaning and preparation (deduplication, severity/label normalization, filtering low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model. A collection of 23,625 smart-contract security audit findings (bug reports), each with a title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/leohachico/audit-findings-dataset.tabulartext-classification10K<n<100K0 likes49 downloads21d agoHugging Face07wg200202 /audit-findings-dataset Smart Contract Audit Findings This is raw, semi-structured data — not a ready-to-train dataset. It still requires further cleaning and preparation (deduplication, severity/label normalization, filtering low-quality or malformed entries, etc.) before it should be used to train or fine-tune an AI model. A collection of 23,625 smart-contract security audit findings (bug reports), each with a title, description, proof-of-concept code, recommendation, and severity rating.… See the full description on the dataset page: https://huggingface.co/datasets/wg200202/audit-findings-dataset.tabulartext-classification10K<n<100K0 likes43 downloads21d agoHugging Face08Solshine /gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit Gemma-4-E2B NLA AV-SFT Training Corpus (v0.1.x, Gemini persona+audit) The 4,734-row AV-SFT training corpus for the v0.1.x Gemma-4-E2B NLA — a 9-source-family diversified expansion over the v0.0.x OpenWebText-only corpus. Labels generated by Gemini CLI following the persona+audit pipeline (Dr. Marisol Chen labels, Dr. Riley Otsuka audits). This is the in-progress v0.1.x labeled training set. AR-SFT companion is still being labeled (~16% complete as of this dataset publish). When the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit.tabulartext-generation1K<n<10K0 likes23 downloads5mo agoHugging Face09Solshine /gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit) The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels. This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.tabulartext-generationn<1K0 likes17 downloads5mo agoHugging Face10ClarusC64 /protein_structure_uncertainty_auditor_v0.2Protein Structure Uncertainty Auditor GoalDetect when predicted protein structures are too uncertain for downstream use. Model must output uncertainty_flag (yes/no) uncertainty_type recommendation This dataset tests whether models can audit structural confidence before use in: drug design docking mutation mapping function inference Run scorer python scorer.py --predictions predictions.jsonl --test_csv data/test.csv tabulartext-classificationn<1K1 likes14 downloads8mo agoHugging Face11ayousanz /oscor-2301-ja-text-auditgated OSCAR 23.01 Japanese text with metadata — audit The audited source, ayousanz/oscor-2301-ja-text, contains 123 JSONL Git LFS shards totaling 244,524,242,021 bytes. Shard ordinals 1..123 are complete, every LFS SHA-256 is unique, and no duplicate or missing ordinal was found at revision 28b9696bbc5306040a6f3076765827bd3592d4e3. Each sampled row is an object with: content: extracted text; metadata: language identification, harmful-content perplexity/score, TLSH, quality warnings… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/oscor-2301-ja-text-audit.tabulartext-generation10K<n<100K0 likes10 downloads3d agoHugging Face12ayousanz /c4-ja-text-auditgated C4 Japanese text-only export — audit The audited source, ayousanz/c4-ja-text, is a text-only derivative of the Japanese portion of allenai/c4. It contains 1,024 Git LFS shards totaling 824,615,119,625 bytes. All declared shard ordinals 00000..01023 are present, every LFS SHA-256 is unique, and no duplicate or missing shard ordinal was found at revision f20a06fede96a076fedd59cb67dbe13f87c4d675. Important format limitation Despite the .json.txt suffix, sampled… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/c4-ja-text-audit.tabulartext-generation10K<n<100K0 likes10 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.