CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01violetxi /harvey-notes-v4 wm-rl notes v4 — two note banks from a recursive self-experience loop Continuation update (2026-09-14): rounds 6–15 appended. The recursive bank now contains 308,580 notes / 138,723,516 training tokens. The original 108,081-row round-5 bank remains an exact prefix. Round-4/5 task lists were available for duplicate rejection, but their trajectory manifests were not published; the continuation therefore seeds round 6 from the published recall sessions and carries complete… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-notes-v4.texttext-generation100K<n<1M0 likes383 downloads4d agoHugging Face02violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p05 harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p05 Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m-kl-0p05, revision 1bced92ee267198876b452d59e13cf17c41d5df9, job 122412. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 16.90% mc_logprob 1,780 50.45% negative_abstain 1,799 41.30% reversed_qa 1,839 7.94%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p05.textquestion-answering1K<n<10K0 likes361 downloads1d agoHugging Face03violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p01 harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p01 Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m-kl-0p01, revision 8c81d012091313d40b3619f15dec11bf47d19c73, job 122411. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 18.21% mc_logprob 1,780 50.11% negative_abstain 1,799 36.85% reversed_qa 1,839 9.52%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p01.textquestion-answering1K<n<10K0 likes357 downloads1d agoHugging Face04violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p1 harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p1 Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m-kl-0p1, revision d657eca6271c10bdd83a5275f58cf0b3d3dc823d, job 122413. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 15.35% mc_logprob 1,780 50.22% negative_abstain 1,799 30.41% reversed_qa 1,839 6.14%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p1.textquestion-answering1K<n<10K0 likes357 downloads1d agoHugging Face05violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-5m harvey-closed-book-qwen35-9b-notes-conditioned-5m Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-5m, revision 906f2b45e2962698c489cfd1f53ab19a4ebe1458, job 122445. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 7.24% mc_logprob 1,780 36.85% negative_abstain 1,799 37.08% reversed_qa 1,839 2.83% forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-5m.textquestion-answering1K<n<10K0 likes355 downloads1d agoHugging Face06violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-10m harvey-closed-book-qwen35-9b-notes-conditioned-10m Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-10m, revision 4bf295a036c010040d76960e167bfbf0efd05927, job 122446. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 13.04% mc_logprob 1,780 41.97% negative_abstain 1,799 51.70% reversed_qa 1,839 6.14% forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-10m.textquestion-answering1K<n<10K0 likes350 downloads1d agoHugging Face07violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1 harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1 Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p1, revision 04390b5989d8d9adfe5597339f24a27b9b411c37, job 122416. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 17.06% mc_logprob 1,780 59.38% negative_abstain 1,799 27.24% reversed_qa 1,839 6.42%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1.textquestion-answering1K<n<10K0 likes307 downloads1d agoHugging Face08violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-1m harvey-closed-book-qwen35-9b-notes-conditioned-1m Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-1m, revision 21c40ef032e9cf968482584191c0b8c269f70001, job 122447. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 5.69% mc_logprob 1,780 34.61% negative_abstain 1,799 10.84% reversed_qa 1,839 2.72% forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-1m.textquestion-answering1K<n<10K0 likes292 downloads1d agoHugging Face09violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 5.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes291 downloads3d agoHugging Face10violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-100m harvey-closed-book-qwen35-9b-notes-conditioned-100m Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m, revision 0c295885100d6eba4f514752aa081c5b0c73fdec, job 122452. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 31.33% mc_logprob 1,780 61.85% negative_abstain 1,799 71.10% reversed_qa 1,839 19.85%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m.textquestion-answering1K<n<10K0 likes278 downloads1d agoHugging Face11violetxi /harvey-closed-book-qwen35-9b-notes-conditioned-30m harvey-closed-book-qwen35-9b-notes-conditioned-30m Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m, revision add7081dd8969ddbe2f6defbf1997b47031284c1, job 122448. All 7,933 probes completed without API errors. Each probe category is a separate split in this single repo. Split Probes Accuracy forward_qa 2,515 22.35% mc_logprob 1,780 54.38% negative_abstain 1,799 66.04% reversed_qa 1,839 12.94% forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m.textquestion-answering1K<n<10K0 likes275 downloads1d agoHugging Face12violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.tabulartext-generation1K<n<10K0 likes274 downloads3d agoHugging Face13violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes272 downloads3d agoHugging Face14collegefishiesd /guitar-fretboard-notes Guitar Single-Note Recordings A dataset of 390 single-note guitar recordings spanning 6 strings and frets 0-12, recorded by two players on acoustic and electric guitars. Dataset Summary This dataset contains isolated single-note recordings from a standard-tuned guitar. Each recording captures one note played on a specific string and fret combination, covering the first 12 frets across all 6 strings (78 unique notes per source). The recordings are raw, unprocessed 44100 Hz… See the full description on the dataset page: https://huggingface.co/datasets/collegefishiesd/guitar-fretboard-notes.audioaudio-classificationn<1K2 likes165 downloads8mo agoHugging Face15notesbymuneeb /ai-waf-dataset Synthetic HTTP Requests Dataset for AI WAF Training This dataset is synthetically generated and contains a diverse set of HTTP requests, labeled as either 'benign' or 'malicious'. It is designed for training and evaluating Web Application Firewalls (WAFs), particularly those based on AI/ML models. The dataset aims to provide a comprehensive collection of both common and sophisticated attack vectors, alongside a wide array of legitimate traffic patterns. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/ai-waf-dataset.text10K<n<100K4 likes136 downloads1y agoHugging Face16notesbymuneeb /epstein-emails Epstein Email Threads Dataset Dataset Summary This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed. Dataset Description Overview This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.texttext-classification1K<n<10K16 likes123 downloads10mo agoHugging Face17violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 2.2000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.tabulartext-generation1K<n<10K0 likes97 downloads3d agoHugging Face18histlearn /community-notes-br Community Notes / X — Snapshot Público Dataset estruturado a partir dos dumps públicos do sistema Community Notes (antigo Birdwatch) da plataforma X (antigo Twitter). Motivação O Community Notes é um sistema de moderação colaborativa onde usuários voluntários escrevem notas contextuais sobre publicações potencialmente enganosas e avaliam as notas de outros participantes. Um algoritmo de consenso determina quais notas são exibidas publicamente. Este dataset… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/community-notes-br.tabulartext-classification100M<n<1B0 likes93 downloads3mo agoHugging Face19aisc-team-a1 /augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.texttext-generation10K<n<100K2 likes86 downloads3y agoHugging Face20Technoculture /synthetic-clinical-notes-embedded Synthetic Clinical Notes This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes: Turn into Alpaca format (instruction, input, and output) Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 158k Token Count 648m Origin https://figshare.com/authors/Zhengyun_Zhao/16480335 Source of raw data PubMed Central (PMC) and MIMIC 3 Processing details original, paper Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.textquestion-answering100K<n<1M10 likes78 downloads3y agoHugging Face21violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.tabulartext-generation1K<n<10K0 likes75 downloads15h agoHugging Face22deadbirds /x-community-notes-parquet-20250222All Twitter/X Community Notes data converted to Parquet. https://communitynotes.x.com/guide/en/about/introduction Pulled Feb 22, 2025 tabular100M<n<1B0 likes74 downloads1y agoHugging Face23violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.tabulartext-generation1K<n<10K0 likes73 downloads15h agoHugging Face24violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.tabulartext-generation1K<n<10K0 likes73 downloads15h agoHugging Face25violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 10M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think.tabulartext-generation1K<n<10K0 likes71 downloads15h agoHugging Face26violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 30M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think.tabulartext-generation1K<n<10K0 likes71 downloads15h agoHugging Face27violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.tabulartext-generation1K<n<10K0 likes70 downloads15h agoHugging Face28rntc /mm-icd-notestabular100K<n<1M6 likes69 downloads2y agoHugging Face29violetxi /harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 1M notes + note-conditioned trajectory mixture, and no KL regularization. The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout. Model, data, and KL condition Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think.tabulartext-generation1K<n<10K0 likes65 downloads15h agoHugging Face30violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and binary all-criteria-pass rule. Mean all-pass rate: 3.1000%. The train split contains held-out evaluation records, not training examples. Model and training mixture The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.tabulartext-generation1K<n<10K0 likes60 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.