CoolFace
Datasetpublic

G-reen/ai-det-test-human-refined-stat-no-labels-test

Auto-Generated FastDetector Dataset Dataset: G-reen/ai-det-test-human-refined-stat-no-labels-test Globals Config: config/globals_aidet_gemma_no_labels.toml Analysis Config: config/analysis_nofilter.toml Rows: 1,677 Evaluation Results Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite) Generator Configs: 1 (gemma-4-31B-it-AWQ-4bit (Temp: 1.0)) Classifiers: 13 (EditLens Roberta-Large Score, EditLens Roberta-Large Bucket, Perplexity… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/ai-det-test-human-refined-stat-no-labels-test.

sourceHugging Faceupdated 19d agoView on Hugging Face
0likes233downloads
Dataset Card

Auto-Generated FastDetector Dataset

  • —Dataset: G-reen/ai-det-test-human-refined-stat-no-labels-test
  • —Globals Config: config/globals_aidet_gemma_no_labels.toml
  • —Analysis Config: config/analysis_nofilter.toml
  • —Rows: 1,677

Evaluation Results

  • —Prompt Subsets: 4 (directreference, indirectreference, revise, rewrite)
  • —Generator Configs: 1 (gemma-4-31B-it-AWQ-4bit (Temp: 1.0))
  • —Classifiers: 13 (EditLens Roberta-Large Score, EditLens Roberta-Large Bucket, Perplexity (Llama-3.2-3B-Instruct) [skipped], Perplexity (Llama-3.2-3B) [skipped], Entropy (Llama-3.2-3B-Instruct) [skipped], Entropy (Llama-3.2-3B) [skipped], Top-p Outliers (Llama-3.2-3B-Instruct) [skipped], Top-p Outliers (Llama-3.2-3B) [skipped], Top-k Outliers (Llama-3.2-3B-Instruct) [skipped], Top-k Outliers (Llama-3.2-3B) [skipped], FastDetectGPT (Llama-3.2-3B-Instruct) [skipped], FastDetectGPT (Llama-3.2-3B) [skipped], Binoculars [skipped])
  • —Filter Conditions: None
  • —Evaluation / Validation Rows: 1,509 / 168 (validation_size = 0.1)
  • —Base Columns: original (Human), final_response (AI)

The best classifier was EditLens Roberta-Large Score with an AUROC of 0.8539. The hardest prompt subset was rewrite with a TPR of 0.0792, and the hardest generator config was gemma-4-31B-it-AWQ-4bit (Temp: 1.0) with a TPR of 0.4619.

ClassifierThresholdAUROCTPRFPRAccuracyF1
✔️ EditLens Roberta-Large Score0.41410.85390.35120.00600.67260.5176
❗ EditLens Roberta-Large Bucket0.00000.77100.57260.03710.76770.7114

Classifier metrics averaged within each prompt and generator subset:

SubsetAverage AUROCAverage TPRAverage FPRAverage AccuracyAverage F1
✔️ Prompt: revise0.96370.74490.02690.85900.8294
Prompt: indirect_reference0.88790.50800.01200.74800.6542
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.0)0.81240.46190.02150.72020.6145
Prompt: direct_reference0.79210.49470.02510.73480.6481
❗ Prompt: rewrite0.59390.07920.02190.52870.1399

✔️ marks the best AUROC, ❗ the worst.

Statistics of Interest

[image] [image] [image] [image]

Appendix

Table of contents 1. Univariate Analysis 2. Correlation Heatmap 3. Distance Histograms 4. Distance Histograms per Prompt Subset 5. Distance Histograms per Generator Config Subset 6. Classifier: EditLens Roberta-Large Score - Performance: - Thresholding: - Classification Histograms: 7. Classifier: EditLens Roberta-Large Bucket - Performance: - Thresholding: - Classification Histograms:

Univariate Analysis

Every statistic the report does arithmetic on, over the 1,509-row evaluation split. Invalid counts rows whose value is missing or non-finite; those rows are excluded from the other columns.

StatisticNMeanMedianStdMinMaxInvalid
jaccard_11,5090.58340.71370.29970.00000.96060
jaccard_21,5090.69510.88480.33540.00001.00000
levenshtein1,5091944.63421515.00001755.68781.000020872.00000
cosdist1,5090.18910.10500.1936-0.00660.77100
EditLens Roberta-Large Score (Human)1,5090.04550.02020.07180.00640.98440
EditLens Roberta-Large Score (AI)1,5090.36860.29570.33090.00680.99960
EditLens Roberta-Large Bucket (Human)1,5090.04240.00000.23770.00003.00000
EditLens Roberta-Large Bucket (AI)1,5091.00531.00001.13560.00003.00000

Correlation Heatmap

Pearson correlation between every statistic of interest, computed over the rows where both statistics are present.

[image]

Distance Histograms

[image] [image] [image] [image]

Distance Histograms per Prompt Subset

[image] [image] [image] [image]

Distance Histograms per Generator Config Subset

[image] [image] [image] [image]

Classifier: EditLens Roberta-Large Score

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall3,0180.85390.35120.00600.67260.5176
Prompt: direct_reference7560.82080.42860.00260.71300.5989
Prompt: indirect_reference7500.95440.35470.00000.67730.5236
✔️ Prompt: revise7800.98470.57180.01030.78080.7229
❗ Prompt: rewrite7320.64130.03280.01090.51090.0628
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.0)3,0180.85390.35120.00600.67260.5176
Thresholding:
  • —Direction: higher_is_ai
  • —Swept for fpr_0_5pct with a found threshold of 0.4141.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image]

Classifier: EditLens Roberta-Large Bucket

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall3,0180.77100.57260.03710.76770.7114
Prompt: direct_reference7560.76350.56080.04760.75660.6974
Prompt: indirect_reference7500.82130.66130.02400.81870.7848
✔️ Prompt: revise7800.94260.91790.04360.93720.9359
❗ Prompt: rewrite7320.54640.12570.03280.54640.2170
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.0)3,0180.77100.57260.03710.76770.7114
Thresholding:
  • —Direction: higher_is_ai
  • —Swept for f1 with a found threshold of 0.0000.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image]