CoolFace
Datasetpublic

G-reen/cc-re-2020-stat-train

Auto-Generated FastDetector Dataset Dataset: G-reen/cc-re-2020-stat-train Globals Config: config/globals_re2020.toml Analysis Config: config/analysis_exclude_deepseek.toml Rows: 108,377 Evaluation Results Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite) Generator Configs: 11 (Laguna-S-2.1-NVFP4 (Temp: 1.0), Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6), Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25), Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7)… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-stat-train.

sourceHugging Faceupdated 14d agoView on Hugging Face
0likes550downloads
Dataset Card

Auto-Generated FastDetector Dataset

  • Dataset: G-reen/cc-re-2020-stat-train
  • Globals Config: config/globals_re2020.toml
  • Analysis Config: config/analysis_exclude_deepseek.toml
  • Rows: 108,377

Evaluation Results

  • Prompt Subsets: 4 (directreference, indirectreference, revise, rewrite)
  • Generator Configs: 11 (Laguna-S-2.1-NVFP4 (Temp: 1.0), Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6), Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25), Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7), Mistral-Small-4-119B-2603-NVFP4 (Temp: 1.25), Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6), Qwen3.8-27B-AWQ-INT4 (Temp: 0.7), Qwen3.8-27B-AWQ-INT4 (Temp: 1.25), gemma-4-31B-it-AWQ-4bit (Temp: 1.0), gemma-4-31B-it-AWQ-4bit (Temp: 1.25), granite-4.2-30b-nvfp4 (Temp: 1.0))
  • Classifiers: 13 (EditLens Roberta-Large Score, EditLens Roberta-Large Bucket, Perplexity (Llama-3.2-3B-Instruct) [skipped], Perplexity (Llama-3.2-3B) [skipped], Entropy (Llama-3.2-3B-Instruct) [skipped], Entropy (Llama-3.2-3B) [skipped], Top-p Outliers (Llama-3.2-3B-Instruct) [skipped], Top-p Outliers (Llama-3.2-3B) [skipped], Top-k Outliers (Llama-3.2-3B-Instruct) [skipped], Top-k Outliers (Llama-3.2-3B) [skipped], FastDetectGPT (Llama-3.2-3B-Instruct) [skipped], FastDetectGPT (Llama-3.2-3B) [skipped], Binoculars [skipped])
  • Filter Conditions: generator_model != deepseek-ai/DeepSeek-V4-Flash-0731
  • Evaluation / Validation Rows: 97,539 / 10,838 (validation_size = 0.1)
  • Base Columns: original (Human), final_response (AI)

The best classifier was EditLens Roberta-Large Score with an AUROC of 0.9044. The hardest prompt subset was rewrite with a TPR of 0.2405, and the hardest generator config was Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6) with a TPR of 0.4525.

ClassifierThresholdAUROCTPRFPRAccuracyF1
✔️ EditLens Roberta-Large Score0.59180.90440.48300.00500.73900.6492
❗ EditLens Roberta-Large Bucket0.00000.84620.72720.05830.83450.8146

Classifier metrics averaged within each prompt and generator subset:

SubsetAverage AUROCAverage TPRAverage FPRAverage AccuracyAverage F1
✔️ Prompt: revise0.96700.77860.03180.87340.8516
Model: Mistral-Small-4-119B-2603-NVFP4 (Temp: 1.25)0.94990.74630.03330.85650.8294
Model: Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7)0.92940.71400.03440.83980.8088
Prompt: direct_reference0.92500.74230.03140.85540.8339
Model: Laguna-S-2.1-NVFP4 (Temp: 1.0)0.91030.69990.03210.83390.8033
Prompt: indirect_reference0.90260.64600.03210.80690.7633
Model: granite-4.2-30b-nvfp4 (Temp: 1.0)0.88820.63580.03240.80170.7541
Model: Qwen3.8-27B-AWQ-INT4 (Temp: 1.25)0.87860.61270.03100.79080.7383
Model: Qwen3.8-27B-AWQ-INT4 (Temp: 0.7)0.86600.60290.03170.78560.7310
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.25)0.86390.58430.03030.77700.7149
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.0)0.86210.58570.03150.77710.7159
Model: Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25)0.83420.51380.02900.74240.6604
Model: Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6)0.82790.50790.03250.73770.6548
Model: Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6)0.81790.45250.03010.71120.5974
❗ Prompt: rewrite0.70130.24050.03140.60450.3600

✔️ marks the best AUROC, ❗ the worst.

Statistics of Interest

[image] [image] [image] [image]

Appendix

Table of contents 1. Univariate Analysis 2. Correlation Heatmap 3. Distance Histograms 4. Distance Histograms per Prompt Subset 5. Distance Histograms per Generator Config Subset 6. Classifier: EditLens Roberta-Large Score - Performance: - Thresholding: - Classification Histograms: 7. Classifier: EditLens Roberta-Large Bucket - Performance: - Thresholding: - Classification Histograms:

Univariate Analysis

Every statistic the report does arithmetic on, over the 97,539-row evaluation split. Invalid counts rows whose value is missing or non-finite; those rows are excluded from the other columns.

StatisticNMeanMedianStdMinMaxInvalid
jaccard_197,5390.63840.76490.28010.00001.00000
jaccard_297,5390.75080.91320.30390.00001.00000
levenshtein97,5392344.50661513.00003557.42891.0000178023.00000
cosdist97,5390.20510.14350.2026-0.00721.03970
EditLens Roberta-Large Score (Human)97,5390.05800.02070.09600.00620.99960
EditLens Roberta-Large Score (AI)97,5390.56710.55740.37690.00650.99960
EditLens Roberta-Large Bucket (Human)97,5390.06860.00000.30740.00003.00000
EditLens Roberta-Large Bucket (AI)97,5391.65491.00001.29990.00003.00000

Correlation Heatmap

Pearson correlation between every statistic of interest, computed over the rows where both statistics are present.

[image]

Distance Histograms

[image] [image] [image] [image]

Distance Histograms per Prompt Subset

[image] [image] [image] [image]

Distance Histograms per Generator Config Subset

[image] [image] [image] [image]

Classifier: EditLens Roberta-Large Score

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall195,0780.90440.48300.00500.73900.6492
Prompt: direct_reference49,4000.94630.65210.00550.82330.7868
Prompt: indirect_reference48,7520.93900.52690.00550.76070.6877
✔️ Prompt: revise49,3740.98140.62400.00450.80980.7663
❗ Prompt: rewrite47,5520.74690.11600.00480.55560.2069
Model: Laguna-S-2.1-NVFP4 (Temp: 1.0)17,6080.93070.59110.00440.79330.7409
Model: Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6)17,9820.87790.42240.00570.70840.5916
Model: Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25)17,8340.88090.41860.00360.70750.5887
Model: Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7)17,7540.94810.57370.00500.78440.7268
Model: Mistral-Small-4-119B-2603-NVFP4 (Temp: 1.25)17,7740.96460.59000.00530.79230.7396
Model: Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6)17,4140.86190.31630.00450.65590.4790
Model: Qwen3.8-27B-AWQ-INT4 (Temp: 0.7)17,6240.89080.48880.00600.74140.6540
Model: Qwen3.8-27B-AWQ-INT4 (Temp: 1.25)17,8140.90730.49210.00580.74310.6570
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.0)17,7260.88640.45930.00420.72760.6277
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.25)17,8200.88840.45450.00520.72470.6228
Model: granite-4.2-30b-nvfp4 (Temp: 1.0)17,7280.91060.50470.00580.74950.6683
Thresholding:
  • Direction: higher_is_ai
  • Swept for fpr_0_5pct with a found threshold of 0.5918.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image] [image] [image] [image] [image]

Classifier: EditLens Roberta-Large Bucket

Performance:
SubsetNAUROCTPRFPRAccuracyF1
Overall195,0780.84620.72720.05830.83450.8146
Prompt: direct_reference49,4000.90360.83260.05740.88760.8810
Prompt: indirect_reference48,7520.86620.76510.05870.85320.8390
✔️ Prompt: revise49,3740.95260.93330.05910.93710.9369
❗ Prompt: rewrite47,5520.65570.36500.05810.65350.5130
Model: Laguna-S-2.1-NVFP4 (Temp: 1.0)17,6080.88980.80870.05990.87440.8656
Model: Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6)17,9820.77790.59340.05940.76700.7180
Model: Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25)17,8340.78750.60890.05450.77720.7322
Model: Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7)17,7540.91060.85420.06380.89520.8908
Model: Mistral-Small-4-119B-2603-NVFP4 (Temp: 1.25)17,7740.93520.90260.06130.92060.9192
Model: Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6)17,4140.77390.58860.05580.76640.7159
Model: Qwen3.8-27B-AWQ-INT4 (Temp: 0.7)17,6240.84130.71700.05740.82980.8081
Model: Qwen3.8-27B-AWQ-INT4 (Temp: 1.25)17,8140.84990.73320.05610.83860.8196
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.0)17,7260.83780.71210.05890.82660.8042
Model: gemma-4-31B-it-AWQ-4bit (Temp: 1.25)17,8200.83940.71400.05540.82930.8071
Model: granite-4.2-30b-nvfp4 (Temp: 1.0)17,7280.86570.76680.05900.85390.8400
Thresholding:
  • Direction: higher_is_ai
  • Swept for f1 with a found threshold of 0.0000.

[image]

Classification Histograms:

[image]

Per Prompt Subset

[image] [image] [image] [image]

Per Generator Config Subset

[image] [image] [image] [image] [image] [image] [image] [image] [image] [image] [image]