CoolFace
Datasetpublic

G-reen/fastdetector-val-stat-val

Auto-Generated FastDetector Dataset Dataset: G-reen/fastdetector-val-stat-val Globals Config: config/globals_val.toml Analysis Config: config/analysis.toml Rows: 12,271 Evaluation Results Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite) Generator Configs: 7 (Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6), Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7), claude-sonnet-4-5-20250929 (Temp: Unknown), claude-sonnet-5 (Temp: Unknown), gpt-3.5-turbo… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/fastdetector-val-stat-val.

sourceHugging Faceupdated 3h agoView on Hugging Face
0likes574downloads
Dataset Card

Auto-Generated FastDetector Dataset

  • Dataset: G-reen/fastdetector-val-stat-val
  • Globals Config: config/globals_val.toml
  • Analysis Config: config/analysis.toml
  • Rows: 12,271

Evaluation Results

  • Prompt Subsets: 4 (directreference, indirectreference, revise, rewrite)
  • Generator Configs: 7 (Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6), Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7), claude-sonnet-4-5-20250929 (Temp: Unknown), claude-sonnet-5 (Temp: Unknown), gpt-3.5-turbo (Temp: Unknown), gpt-4o (Temp: Unknown), hy3 (Temp: Unknown))
  • Classifiers: 19 (EditLens Roberta-Large Score, EditLens Roberta-Large Bucket, EditLens Llama-3.2-3B Score, EditLens Llama-3.2-3B Bucket, EditLens Giga RoBERTa-Large Score, EditLens Giga RoBERTa-Large Bucket, EditLens Giga Llama-3.2-3B Score, EditLens Giga Llama-3.2-3B Bucket, Perplexity (Llama-3.2-3B-Instruct), Perplexity (Llama-3.2-3B), Entropy (Llama-3.2-3B-Instruct), Entropy (Llama-3.2-3B), Top-p Outliers (Llama-3.2-3B-Instruct), Top-p Outliers (Llama-3.2-3B), Top-k Outliers (Llama-3.2-3B-Instruct), Top-k Outliers (Llama-3.2-3B), FastDetectGPT (Llama-3.2-3B-Instruct), FastDetectGPT (Llama-3.2-3B), Binoculars)
  • Filter Conditions: cosdist >= 0.03 OR softngram >= 0.06
  • Evaluation / Validation Rows: 11,043 / 1,228 (validation_size = 0.1)
  • Base Columns: original (Human), final_response (AI)

The best classifier was EditLens Giga Llama-3.2-3B Score with an AUROC of 0.9521. The hardest prompt subset was rewrite with a TPR of 0.1685, and the hardest generator config was hy3 (Temp: Unknown) with a TPR of 0.2105.

ClassifierThresholdAUROCTPRFPRAccuracyF1
✔️ EditLens Giga Llama-3.2-3B Score0.31880.95210.63280.00280.81500.7738
EditLens Llama-3.2-3B Score0.35260.93520.65100.00560.82270.7859
EditLens Giga RoBERTa-Large Score0.63810.92680.52970.00750.76110.6891
EditLens Roberta-Large Score0.51670.90390.44540.00800.71870.6130
EditLens Llama-3.2-3B Bucket0.00000.87020.75400.02350.86530.8484
EditLens Giga RoBERTa-Large Bucket0.00000.86090.74090.02810.85640.8377
EditLens Giga Llama-3.2-3B Bucket0.00000.85230.70790.00520.85140.8265
EditLens Roberta-Large Bucket0.00000.81930.67750.05760.81000.7810
Perplexity (Llama-3.2-3B-Instruct)4.44160.65900.07210.00240.53480.1342
Entropy (Llama-3.2-3B-Instruct)1.41990.65610.08670.00330.54170.1590
Top-k Outliers (Llama-3.2-3B-Instruct)0.02120.63160.06350.00260.53040.1191
Perplexity (Llama-3.2-3B)1.25390.60530.00030.00120.49950.0005
Binoculars1.03770.60470.01390.00430.50480.0272
Top-p Outliers (Llama-3.2-3B)0.00260.60460.00710.00250.50230.0140
Entropy (Llama-3.2-3B)0.31610.59860.00040.00170.49930.0007
Top-k Outliers (Llama-3.2-3B)0.00450.58710.00540.00200.50170.0108
Top-p Outliers (Llama-3.2-3B-Instruct)0.02440.54370.02430.00600.50910.0471
FastDetectGPT (Llama-3.2-3B-Instruct)3.20750.51060.00500.00540.49980.0099
❗ FastDetectGPT (Llama-3.2-3B)-3.04520.42780.01150.00520.50320.0226

Classifier metrics averaged within each prompt and generator subset:

SubsetAverage AUROCAverage TPRAverage FPRAverage AccuracyAverage F1
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)0.77560.37310.00960.68180.4244
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)0.74270.27290.00930.63180.3590
Prompt: revise0.73410.32860.00880.65990.3695
Prompt: direct_reference0.73290.32920.00940.65990.3939
Prompt: indirect_reference0.73030.28190.00810.63690.3578
Model: gpt-3.5-turbo (Temp: Unknown)0.72810.30270.00830.64720.3734
Model: gpt-4o (Temp: Unknown)0.72540.32010.00870.65570.3658
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)0.70770.26590.00900.62840.3272
Model: claude-sonnet-5 (Temp: Unknown)0.69470.24000.00980.61510.3028
Prompt: rewrite0.63290.16850.01080.57880.2378
❗ Model: hy3 (Temp: Unknown)0.60920.21050.00970.60040.2780

✔️ marks the best AUROC, ❗ the worst.

Statistics of Interest

[image] [image] [image] [image]

Appendix

Table of contents 1. Univariate Analysis 2. Correlation Heatmap 3. Distance Histograms 4. Distance Histograms per Prompt Subset 5. Distance Histograms per Generator Config Subset 6. Classifier: EditLens Roberta-Large Score - Performance: - Thresholding: - Classification Histograms: 7. Classifier: EditLens Roberta-Large Bucket - Performance: - Thresholding: - Classification Histograms: 8. Classifier: EditLens Llama-3.2-3B Score - Performance: - Thresholding: - Classification Histograms: 9. Classifier: EditLens Llama-3.2-3B Bucket - Performance: - Thresholding: - Classification Histograms: 10. Classifier: EditLens Giga RoBERTa-Large Score - Performance: - Thresholding: - Classification Histograms: 11. Classifier: EditLens Giga RoBERTa-Large Bucket - Performance: - Thresholding: - Classification Histograms: 12. Classifier: EditLens Giga Llama-3.2-3B Score - Performance: - Thresholding: - Classification Histograms: 13. Classifier: EditLens Giga Llama-3.2-3B Bucket - Performance: - Thresholding: - Classification Histograms: 14. Classifier: Perplexity (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 15. Classifier: Perplexity (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 16. Classifier: Entropy (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 17. Classifier: Entropy (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 18. Classifier: Top-p Outliers (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 19. Classifier: Top-p Outliers (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 20. Classifier: Top-k Outliers (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 21. Classifier: Top-k Outliers (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 22. Classifier: FastDetectGPT (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 23. Classifier: FastDetectGPT (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 24. Classifier: Binoculars - Performance: - Thresholding: - Classification Histograms:

Univariate Analysis

Every statistic the report does arithmetic on, over the 11,043-row evaluation split. Invalid counts rows whose value is missing or non-finite; those rows are excluded from the other columns.

StatisticNMeanMedianStdMinMaxInvalid
jaccard_111,0430.67740.76220.22680.00001.00000
jaccard_211,0430.80030.90990.23450.02161.00000
levenshtein11,0431990.59321355.00002350.386411.000065876.00000
softngram11,0430.61410.67510.31280.00001.00000
cosdist11,0430.21970.15450.2016-0.00471.03150
bertscore11,0430.14630.15290.06890.00410.43380
bertscore_precision11,0430.14390.14830.06790.00360.51060
bertscore_recall11,0430.14750.15340.07600.00140.43190
moverscore11,0430.56200.60270.16460.10111.08100
reranker11,043-2.8248-4.68755.8203-11.187519.00000
EditLens Roberta-Large Score (Human)11,0430.05820.02050.09890.00640.99930
EditLens Roberta-Large Score (AI)11,0430.51030.43600.36770.00650.99960
EditLens Roberta-Large Bucket (Human)11,0430.06970.00000.31820.00003.00000
EditLens Roberta-Large Bucket (AI)11,0431.46341.00001.29440.00003.00000
EditLens Llama-3.2-3B Score (Human)11,0430.04550.02250.06270.00030.85580
EditLens Llama-3.2-3B Score (AI)11,0430.53260.50650.34110.00111.00000
EditLens Llama-3.2-3B Bucket (Human)11,0430.02450.00000.16380.00003.00000
EditLens Llama-3.2-3B Bucket (AI)11,0431.53781.00001.19140.00003.00000
EditLens Giga RoBERTa-Large Score (Human)11,0430.02580.00490.09620.00130.99990
EditLens Giga RoBERTa-Large Score (AI)11,0430.59490.70050.39720.00141.00000
EditLens Giga RoBERTa-Large Bucket (Human)11,0430.04480.00000.30340.00003.00000
EditLens Giga RoBERTa-Large Bucket (AI)11,0431.77172.00001.27320.00003.00000
EditLens Giga Llama-3.2-3B Score (Human)11,0430.01180.00390.03590.00040.93390
EditLens Giga Llama-3.2-3B Score (AI)11,0430.51710.48560.37430.00071.00000
EditLens Giga Llama-3.2-3B Bucket (Human)11,0430.00640.00000.10010.00003.00000
EditLens Giga Llama-3.2-3B Bucket (AI)11,0431.53281.00001.25100.00003.00000
Perplexity (Llama-3.2-3B-Instruct) (Human)11,04319.503916.094516.25681.6941757.41170
Perplexity (Llama-3.2-3B-Instruct) (AI)11,04315.374812.177118.14141.0586990.48980
Perplexity (Llama-3.2-3B) (Human)11,04314.490612.412810.37961.0484485.73970
Perplexity (Llama-3.2-3B) (AI)11,04312.846310.440312.83581.0502551.98390
Entropy (Llama-3.2-3B-Instruct) (Human)11,0432.62592.58820.49480.48997.63100
Entropy (Llama-3.2-3B-Instruct) (AI)11,0432.30392.30970.63940.09456.71150
Entropy (Llama-3.2-3B) (Human)11,0432.52562.51620.45270.07845.67940
Entropy (Llama-3.2-3B) (AI)11,0432.36272.37320.52470.06076.01520
Top-p Outliers (Llama-3.2-3B-Instruct) (Human)11,0430.05850.05560.01710.00000.17460
Top-p Outliers (Llama-3.2-3B-Instruct) (AI)11,0430.05580.05380.01840.00000.20000
Top-p Outliers (Llama-3.2-3B) (Human)11,0430.04310.04260.01290.00000.13040
Top-p Outliers (Llama-3.2-3B) (AI)11,0430.03850.03810.01550.00000.16670
Top-k Outliers (Llama-3.2-3B-Instruct) (Human)11,0430.10760.10100.04500.00000.54350
Top-k Outliers (Llama-3.2-3B-Instruct) (AI)11,0430.08810.08090.05220.00000.54810
Top-k Outliers (Llama-3.2-3B) (Human)11,0430.08820.08230.04050.00000.45650
Top-k Outliers (Llama-3.2-3B) (AI)11,0430.07820.07030.04790.00000.52940
FastDetectGPT (Llama-3.2-3B-Instruct) (Human)11,043-1.6809-1.66931.6419-15.343225.55120
FastDetectGPT (Llama-3.2-3B-Instruct) (AI)11,043-1.6320-1.61661.6742-15.527610.15060
FastDetectGPT (Llama-3.2-3B) (Human)11,043-0.1344-0.10421.0496-6.43725.09800
FastDetectGPT (Llama-3.2-3B) (AI)11,0430.22910.18991.5314-15.971310.63770
Binoculars (Human)11,0430.85580.86380.09740.00741.13400
Binoculars (AI)11,0430.88420.88390.07150.04801.22090

Correlation Heatmap

Pearson correlation between every statistic of interest, computed over the rows where both statistics are present.

[image]

Distance Histograms

[image] [image] [image] [image] [image] [image] [image] [image] [image] [image]

Distance Histograms per Prompt Subset

[image] [image] [image] [image] [image] [image] [image] [image] [image] [image]

Distance Histograms per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Roberta-Large Score

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.90390.44540.00800.71870.6130
Prompt: direct_reference6,1660.90290.54200.00710.76740.6997
Prompt: indirect_reference5,5840.91010.46280.00570.72850.6302
Prompt: revise6,0320.96540.50200.00900.74650.6645
❗ Prompt: rewrite4,3040.80800.20540.01070.59740.3378
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.86690.38230.00720.68750.5502
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.96990.57920.00740.78590.7301
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.89690.46320.00740.72790.6300
Model: claude-sonnet-5 (Temp: Unknown)3,1200.88550.34620.00900.66860.5109
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.90060.47560.00730.73420.6415
Model: gpt-4o (Temp: Unknown)3,3080.94840.56770.00970.77900.7198
Model: hy3 (Temp: Unknown)3,0600.84710.27710.00780.63460.4313
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.5167.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Roberta-Large Bucket

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.81930.67750.05760.81000.7810
Prompt: direct_reference6,1660.83470.70350.05970.82190.7980
Prompt: indirect_reference5,5840.79980.63290.05270.79010.7510
✔️ Prompt: revise6,0320.90470.84810.05700.89560.8904
❗ Prompt: rewrite4,3040.70230.45910.06180.69870.6037
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.75220.54270.05480.74400.6795
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.90130.83750.06100.88830.8823
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.81710.67050.05480.80780.7772
Model: claude-sonnet-5 (Temp: Unknown)3,1200.78670.61730.05580.78080.7379
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.81500.66690.06070.80310.7721
Model: gpt-4o (Temp: Unknown)3,3080.89220.81740.05320.88210.8739
Model: hy3 (Temp: Unknown)3,0600.75260.55690.06270.74710.6877
Thresholding:
  • Direction: higher_is_ai
  • Swept for f1 with a found threshold of 0.0000.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Llama-3.2-3B Score

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.93520.65100.00560.82270.7859
Prompt: direct_reference6,1660.94670.75060.00290.87380.8561
Prompt: indirect_reference5,5840.92770.60890.00610.80140.7540
Prompt: revise6,0320.98230.78910.00630.89140.8790
❗ Prompt: rewrite4,3040.85950.36940.00790.68080.5364
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.91300.56820.00390.78210.7228
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.98450.84440.00570.91930.9128
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.91870.61630.00530.80550.7601
Model: claude-sonnet-5 (Temp: Unknown)3,1200.91840.55640.01030.77310.7103
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.93820.67880.00130.83870.8080
Model: gpt-4o (Temp: Unknown)3,3080.96700.76180.00360.87910.8630
Model: hy3 (Temp: Unknown)3,0600.90000.49540.00920.74310.6586
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.3526.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Llama-3.2-3B Bucket

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.87020.75400.02350.86530.8484
Prompt: direct_reference6,1660.89930.80670.01950.89360.8835
Prompt: indirect_reference5,5840.84770.70810.02330.84240.8180
Prompt: revise6,0320.94560.90380.02250.94060.9384
❗ Prompt: rewrite4,3040.75100.52790.03070.74860.6774
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.83440.68300.02350.82970.8005
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.95370.91620.02510.94560.9439
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.84760.70990.02340.84320.8191
Model: claude-sonnet-5 (Temp: Unknown)3,1200.83780.69680.02820.83430.8079
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.88030.76980.01850.87570.8609
Model: gpt-4o (Temp: Unknown)3,3080.91790.84520.02120.91200.9057
Model: hy3 (Temp: Unknown)3,0600.80370.62610.02420.80100.7588
Thresholding:
  • Direction: higher_is_ai
  • Swept for f1 with a found threshold of 0.0000.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Giga RoBERTa-Large Score

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.92680.52970.00750.76110.6891
Prompt: direct_reference6,1660.92990.66560.00840.82860.7952
Prompt: indirect_reference5,5840.92330.54940.00790.77080.7056
Prompt: revise6,0320.97800.56760.00560.78100.7216
❗ Prompt: rewrite4,3040.85240.25600.00840.62380.4050
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.89380.48660.00390.74140.6530
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.98170.70640.00510.85060.8254
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.92430.51140.00740.75200.6734
Model: claude-sonnet-5 (Temp: Unknown)3,1200.90950.44290.01280.71510.6085
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.92280.58380.00730.78830.7338
Model: gpt-4o (Temp: Unknown)3,3080.95930.62580.00600.80990.7670
Model: hy3 (Temp: Unknown)3,0600.88700.31900.01050.65420.4798
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.6381.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Giga RoBERTa-Large Bucket

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.86090.74090.02810.85640.8377
Prompt: direct_reference6,1660.88370.78430.03080.87670.8642
Prompt: indirect_reference5,5840.84110.69810.02290.83760.8112
Prompt: revise6,0320.94290.90250.02720.93770.9354
❗ Prompt: rewrite4,3040.73940.50790.03210.73790.6596
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.82510.66730.02670.82030.7878
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.95400.92130.02680.94730.9459
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.85110.72060.02670.84690.8248
Model: claude-sonnet-5 (Temp: Unknown)3,1200.81610.65710.03010.81350.7789
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.86470.74800.02970.85920.8416
Model: gpt-4o (Temp: Unknown)3,3080.90860.83310.02780.90270.8954
Model: hy3 (Temp: Unknown)3,0600.79030.60650.02880.78890.7418
Thresholding:
  • Direction: higher_is_ai
  • Swept for f1 with a found threshold of 0.0000.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Giga Llama-3.2-3B Score

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.95210.63280.00280.81500.7738
Prompt: direct_reference6,1660.95320.74080.00390.86850.8492
Prompt: indirect_reference5,5840.95270.60100.00250.79920.7496
Prompt: revise6,0320.98710.75760.00170.87800.8613
❗ Prompt: rewrite4,3040.89570.34430.00330.67050.5110
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.92700.56290.00200.78050.7195
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.99260.82610.00110.91250.9042
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.94350.58620.00330.79140.7376
Model: claude-sonnet-5 (Temp: Unknown)3,1200.94220.55060.00640.77210.7073
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.95460.66160.00130.83010.7957
Model: gpt-4o (Temp: Unknown)3,3080.97700.73760.00120.86820.8484
Model: hy3 (Temp: Unknown)3,0600.92100.46860.00460.73200.6362
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.3188.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Giga Llama-3.2-3B Bucket

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.85230.70790.00520.85140.8265
Prompt: direct_reference6,1660.88730.77750.00620.88570.8718
Prompt: indirect_reference5,5840.82940.66150.00470.82840.7941
Prompt: revise6,0320.93340.86940.00400.93270.9281
❗ Prompt: rewrite4,3040.71820.44190.00600.71790.6104
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.81370.63080.00590.81250.7708
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.94770.89740.00400.94670.9439
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.83600.67650.00600.83520.8041
Model: claude-sonnet-5 (Temp: Unknown)3,1200.81290.63210.00830.81190.7706
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.86330.72820.00330.86250.8411
Model: gpt-4o (Temp: Unknown)3,3080.89690.79560.00300.89630.8847
Model: hy3 (Temp: Unknown)3,0600.77800.56080.00590.77750.7159
Thresholding:
  • Direction: higher_is_ai
  • Swept for f1 with a found threshold of 0.0000.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Perplexity (Llama-3.2-3B-Instruct)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.65900.07210.00240.53480.1342
Prompt: direct_reference6,1660.69670.12780.00320.56230.2260
Prompt: indirect_reference5,5840.73030.11860.00210.55820.2116
Prompt: revise6,0320.64780.01530.00130.50700.0300
Prompt: rewrite4,3040.53060.01160.00330.50420.0229
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.78460.19630.00460.59590.3270
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.76380.15850.00400.57730.2727
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.66710.01270.00130.50570.0250
Model: claude-sonnet-5 (Temp: Unknown)3,1200.65610.01150.00190.50480.0228
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.71870.10690.00070.55310.1930
Model: gpt-4o (Temp: Unknown)3,3080.62970.00850.00300.50270.0167
❗ Model: hy3 (Temp: Unknown)3,0600.38310.00260.00130.50070.0052
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 4.4416.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Perplexity (Llama-3.2-3B)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.60530.00030.00120.49950.0005
Prompt: direct_reference6,1660.63690.00000.00100.49950.0000
Prompt: indirect_reference5,5840.68370.00040.00040.50000.0007
Prompt: revise6,0320.58770.00000.00170.49920.0000
Prompt: rewrite4,3040.48590.00090.00190.49950.0019
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.73100.00000.00070.49970.0000
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.70750.00000.00110.49940.0000
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.61160.00070.00330.49870.0013
Model: claude-sonnet-5 (Temp: Unknown)3,1200.60480.00060.00130.49970.0013
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.67050.00000.00000.50000.0000
Model: gpt-4o (Temp: Unknown)3,3080.55640.00000.00120.49940.0000
❗ Model: hy3 (Temp: Unknown)3,0600.34680.00070.00070.50000.0013
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 1.2539.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Entropy (Llama-3.2-3B-Instruct)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.65610.08670.00330.54170.1590
Prompt: direct_reference6,1660.69270.15120.00420.57350.2617
Prompt: indirect_reference5,5840.72220.13930.00180.56880.2442
Prompt: revise6,0320.64540.02520.00200.51160.0491
Prompt: rewrite4,3040.53410.01210.00560.50330.0237
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.75380.18460.00720.58870.3098
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.77140.19900.00510.59690.3305
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.66460.03070.00400.51340.0594
Model: claude-sonnet-5 (Temp: Unknown)3,1200.65240.01670.00190.50740.0327
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.72250.13920.00070.56930.2442
Model: gpt-4o (Temp: Unknown)3,3080.61130.02180.00240.50970.0425
❗ Model: hy3 (Temp: Unknown)3,0600.40660.00390.00130.50130.0078
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 1.4199.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Entropy (Llama-3.2-3B)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.59860.00040.00170.49930.0007
Prompt: direct_reference6,1660.62160.00000.00130.49940.0000
Prompt: indirect_reference5,5840.66220.00070.00070.50000.0014
Prompt: revise6,0320.58800.00000.00230.49880.0000
Prompt: rewrite4,3040.50020.00090.00280.49910.0019
Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.68850.00000.00130.49930.0000
✔️ Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.71910.00000.00230.49890.0000
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.59630.00070.00330.49870.0013
Model: claude-sonnet-5 (Temp: Unknown)3,1200.58350.00060.00130.49970.0013
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.63680.00000.00000.50000.0000
Model: gpt-4o (Temp: Unknown)3,3080.58810.00060.00240.49910.0012
❗ Model: hy3 (Temp: Unknown)3,0600.36380.00070.00130.49970.0013
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.3161.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Top-p Outliers (Llama-3.2-3B-Instruct)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.54370.02430.00600.50910.0471
Prompt: direct_reference6,1660.56850.03890.00650.51620.0745
Prompt: indirect_reference5,5840.56530.03620.00470.51580.0695
Prompt: revise6,0320.51650.00700.00630.50030.0137
Prompt: rewrite4,3040.51790.01210.00650.50280.0237
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.68620.09200.00910.54140.1671
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.55100.01600.00740.50430.0312
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.50400.00200.00470.49870.0040
Model: claude-sonnet-5 (Temp: Unknown)3,1200.50500.00450.00260.50100.0089
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.54860.04220.00920.51650.0803
Model: gpt-4o (Temp: Unknown)3,3080.53690.00540.00480.50030.0108
❗ Model: hy3 (Temp: Unknown)3,0600.47240.01050.00390.50330.0206
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.0244.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Top-p Outliers (Llama-3.2-3B)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.60460.00710.00250.50230.0140
Prompt: direct_reference6,1660.65750.01300.00390.50450.0255
Prompt: indirect_reference5,5840.65630.00790.00140.50320.0156
Prompt: revise6,0320.57720.00300.00170.50070.0059
Prompt: rewrite4,3040.50060.00330.00330.50000.0065
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.71060.00650.00460.50100.0129
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.64050.01080.00110.50480.0214
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.60800.00130.00200.49970.0027
Model: claude-sonnet-5 (Temp: Unknown)3,1200.62150.00060.00190.49940.0013
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.68880.02180.00260.50960.0425
Model: gpt-4o (Temp: Unknown)3,3080.51140.00180.00180.50000.0036
❗ Model: hy3 (Temp: Unknown)3,0600.45630.00650.00390.50130.0129
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.0026.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Top-k Outliers (Llama-3.2-3B-Instruct)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.63160.06350.00260.53040.1191
Prompt: direct_reference6,1660.68120.11740.00360.55690.2095
Prompt: indirect_reference5,5840.70730.10030.00180.54920.1820
Prompt: revise6,0320.60560.01530.00200.50660.0300
Prompt: rewrite4,3040.50060.00600.00330.50140.0120
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.75810.15590.00390.57600.2688
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.70950.13110.00340.56390.2312
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.66760.02410.00270.51070.0469
Model: claude-sonnet-5 (Temp: Unknown)3,1200.64040.01220.00130.50540.0240
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.67280.09630.00130.54750.1755
Model: gpt-4o (Temp: Unknown)3,3080.59580.01570.00360.50600.0308
❗ Model: hy3 (Temp: Unknown)3,0600.37290.00330.00200.50070.0065
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.0212.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Top-k Outliers (Llama-3.2-3B)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.58710.00540.00200.50170.0108
Prompt: direct_reference6,1660.63250.01070.00160.50450.0211
Prompt: indirect_reference5,5840.66600.00750.00070.50340.0149
Prompt: revise6,0320.55690.00100.00300.49900.0020
Prompt: rewrite4,3040.46370.00140.00280.49930.0028
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.71190.00910.00260.50330.0181
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.65530.01250.00230.50510.0247
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.62540.00070.00400.49830.0013
Model: claude-sonnet-5 (Temp: Unknown)3,1200.60280.00060.00060.50000.0013
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.62910.01320.00130.50590.0260
Model: gpt-4o (Temp: Unknown)3,3080.53530.00060.00180.49940.0012
❗ Model: hy3 (Temp: Unknown)3,0600.34820.00070.00130.49970.0013
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.0045.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: FastDetectGPT (Llama-3.2-3B-Instruct)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.51060.00500.00540.49980.0099
Prompt: direct_reference6,1660.51040.00550.00360.50100.0109
Prompt: indirect_reference5,5840.51350.00360.00540.49910.0071
Prompt: revise6,0320.49700.00530.00560.49980.0105
Prompt: rewrite4,3040.52580.00560.00790.49880.0110
✔️ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.68250.00720.00850.49930.0141
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.45670.00110.00680.49710.0023
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.44840.00600.00400.50100.0119
❗ Model: claude-sonnet-5 (Temp: Unknown)3,1200.44560.00320.00380.49970.0064
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.52300.00260.00400.49930.0052
Model: gpt-4o (Temp: Unknown)3,3080.52830.00670.00600.50030.0131
Model: hy3 (Temp: Unknown)3,0600.49570.00850.00460.50200.0168
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 3.2075.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: FastDetectGPT (Llama-3.2-3B)

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.42780.01150.00520.50320.0226
Prompt: direct_reference6,1660.38200.00880.00450.50210.0173
Prompt: indirect_reference5,5840.35340.00930.00610.50160.0183
Prompt: revise6,0320.46430.01530.00530.50500.0299
Prompt: rewrite4,3040.53810.01300.00460.50420.0256
❗ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.23780.00260.00260.50000.0052
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.43850.01030.00860.50090.0201
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.40550.00940.00270.50330.0185
Model: claude-sonnet-5 (Temp: Unknown)3,1200.37150.00510.00640.49940.0101
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.31810.00590.00460.50070.0117
✔️ Model: gpt-4o (Temp: Unknown)3,3080.60380.02360.00600.50880.0458
Model: hy3 (Temp: Unknown)3,0600.60180.02290.00460.50920.0445
Thresholding:
  • Direction: lower_is_ai
  • Swept for fpr_0_5pct with a found threshold of -3.0452.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]

Classifier: Binoculars

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall22,0860.60470.01390.00430.50480.0272
Prompt: direct_reference6,1660.60830.01040.00650.50190.0204
Prompt: indirect_reference5,5840.58330.00930.00360.50290.0184
Prompt: revise6,0320.62300.01530.00360.50580.0299
Prompt: rewrite4,3040.60180.02280.00280.51000.0444
❗ Model: Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6)3,0660.53930.00720.00390.50160.0142
Model: Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7)3,5080.63770.02110.00340.50880.0412
Model: claude-sonnet-4-5-20250929 (Temp: Unknown)2,9920.61170.00940.00530.50200.0184
Model: claude-sonnet-5 (Temp: Unknown)3,1200.60710.00580.00190.50190.0115
Model: gpt-3.5-turbo (Temp: Unknown)3,0320.56600.01120.00400.50360.0221
Model: gpt-4o (Temp: Unknown)3,3080.61850.01330.00600.50360.0261
✔️ Model: hy3 (Temp: Unknown)3,0600.64690.02810.00520.51140.0544
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 1.0377.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image]