datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harvey-notes-v4
wm-rl notes v4 — two note banks from a recursive self-experience loop
Continuation update (2026-09-14): rounds 6–15 appended. The recursive bank now contains 308,580 notes / 138,723,516 training tokens. The original 108,081-row round-5 bank remains an exact prefix. Round-4/5 task lists were available for duplicate rejection, but their trajectory manifests were not published; the continuation therefore seeds round 6 from the published recall sessions and carries complete… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-notes-v4.harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p05
harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p05
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m-kl-0p05,
revision 1bced92ee267198876b452d59e13cf17c41d5df9, job 122412. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
16.90%
mc_logprob
1,780
50.45%
negative_abstain
1,799
41.30%
reversed_qa
1,839
7.94%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p05.harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p01
harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p01
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m-kl-0p01,
revision 8c81d012091313d40b3619f15dec11bf47d19c73, job 122411. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
18.21%
mc_logprob
1,780
50.11%
negative_abstain
1,799
36.85%
reversed_qa
1,839
9.52%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p01.harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p1
harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p1
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m-kl-0p1,
revision d657eca6271c10bdd83a5275f58cf0b3d3dc823d, job 122413. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
15.35%
mc_logprob
1,780
50.22%
negative_abstain
1,799
30.41%
reversed_qa
1,839
6.14%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m-kl-0p1.harvey-closed-book-qwen35-9b-notes-conditioned-5m
harvey-closed-book-qwen35-9b-notes-conditioned-5m
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-5m,
revision 906f2b45e2962698c489cfd1f53ab19a4ebe1458, job 122445. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
7.24%
mc_logprob
1,780
36.85%
negative_abstain
1,799
37.08%
reversed_qa
1,839
2.83%
forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-5m.harvey-closed-book-qwen35-9b-notes-conditioned-10m
harvey-closed-book-qwen35-9b-notes-conditioned-10m
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-10m,
revision 4bf295a036c010040d76960e167bfbf0efd05927, job 122446. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
13.04%
mc_logprob
1,780
41.97%
negative_abstain
1,799
51.70%
reversed_qa
1,839
6.14%
forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-10m.harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1
harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m-kl-0p1,
revision 04390b5989d8d9adfe5597339f24a27b9b411c37, job 122416. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
17.06%
mc_logprob
1,780
59.38%
negative_abstain
1,799
27.24%
reversed_qa
1,839
6.42%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m-kl-0p1.harvey-closed-book-qwen35-9b-notes-conditioned-1m
harvey-closed-book-qwen35-9b-notes-conditioned-1m
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-1m,
revision 21c40ef032e9cf968482584191c0b8c269f70001, job 122447. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
5.69%
mc_logprob
1,780
34.61%
negative_abstain
1,799
10.84%
reversed_qa
1,839
2.72%
forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-1m.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-closed-book-qwen35-9b-notes-conditioned-100m
harvey-closed-book-qwen35-9b-notes-conditioned-100m
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-100m,
revision 0c295885100d6eba4f514752aa081c5b0c73fdec, job 122452. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
31.33%
mc_logprob
1,780
61.85%
negative_abstain
1,799
71.10%
reversed_qa
1,839
19.85%… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-100m.harvey-closed-book-qwen35-9b-notes-conditioned-30m
harvey-closed-book-qwen35-9b-notes-conditioned-30m
Completed closed-book C&H knowledge evaluation of violetxi/qwen35-9b-harvey-v4-notes-conditioned-30m,
revision add7081dd8969ddbe2f6defbf1997b47031284c1, job 122448. All 7,933 probes completed
without API errors. Each probe category is a separate split in this single repo.
Split
Probes
Accuracy
forward_qa
2,515
22.35%
mc_logprob
1,780
54.38%
negative_abstain
1,799
66.04%
reversed_qa
1,839
12.94%
forward_qa:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-closed-book-qwen35-9b-notes-conditioned-30m.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.guitar-fretboard-notes
Guitar Single-Note Recordings
A dataset of 390 single-note guitar recordings spanning 6 strings and frets 0-12, recorded by two players on acoustic and electric guitars.
Dataset Summary
This dataset contains isolated single-note recordings from a standard-tuned guitar. Each recording captures one note played on a specific string and fret combination, covering the first 12 frets across all 6 strings (78 unique notes per source). The recordings are raw, unprocessed 44100 Hz… See the full description on the dataset page: https://huggingface.co/datasets/collegefishiesd/guitar-fretboard-notes.ai-waf-dataset
Synthetic HTTP Requests Dataset for AI WAF Training
This dataset is synthetically generated and contains a diverse set of HTTP requests, labeled as either 'benign' or 'malicious'. It is designed for training and evaluating Web Application Firewalls (WAFs), particularly those based on AI/ML models.
The dataset aims to provide a comprehensive collection of both common and sophisticated attack vectors, alongside a wide array of legitimate traffic patterns.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/ai-waf-dataset.epstein-emails
Epstein Email Threads Dataset
Dataset Summary
This dataset contains 5,082 parsed email threads extracted from OCR'd documents released by the U.S. House Oversight Committee. The emails have been processed using large language models to extract structured information including senders, recipients, timestamps, subjects, and message bodies, with OCR errors corrected and footers removed.
Dataset Description
Overview
This is a structured, machine-readable… See the full description on the dataset page: https://huggingface.co/datasets/notesbymuneeb/epstein-emails.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.community-notes-br
Community Notes / X — Snapshot Público
Dataset estruturado a partir dos dumps públicos do sistema Community Notes (antigo Birdwatch) da plataforma X (antigo Twitter).
Motivação
O Community Notes é um sistema de moderação colaborativa onde usuários voluntários escrevem notas contextuais sobre publicações potencialmente enganosas e avaliam as notas de outros participantes. Um algoritmo de consenso determina quais notas são exibidas publicamente. Este dataset… See the full description on the dataset page: https://huggingface.co/datasets/histlearn/community-notes-br.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.synthetic-clinical-notes-embedded
Synthetic Clinical Notes
This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes:
Turn into Alpaca format (instruction, input, and output)
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
158k
Token Count
648m
Origin
https://figshare.com/authors/Zhengyun_Zhao/16480335
Source of raw data
PubMed Central (PMC) and MIMIC 3
Processing details
original, paper
Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.05.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p05-think.x-community-notes-parquet-20250222All Twitter/X Community Notes data converted to Parquet.
https://communitynotes.x.com/guide/en/about/introduction
Pulled Feb 22, 2025
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.01.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p01-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and KL coefficient 0.1.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-kl-0p1-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 10M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-10m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 30M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-30m-think.harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 100M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-100m-think.mm-icd-notesharvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think
harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think
Complete closed-book recall evaluation: 7,933 probes. One dataset repository for Qwen3.5-9B, the 1M notes + note-conditioned trajectory mixture, and no KL regularization.
The train split contains evaluation records. Each row is one scored probe; this split name follows the existing evaluation dataset layout.
Model, data, and KL condition
Evaluated model:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-recall-qwen35-9b-notes70-notecondtraj30-1m-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.1000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.
