G-reen/cc-re-2020-stat-train
Auto-Generated FastDetector Dataset Dataset: G-reen/cc-re-2020-stat-train Globals Config: config/globals_re2020.toml Analysis Config: config/analysis_exclude_deepseek.toml Rows: 108,377 Evaluation Results Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite) Generator Configs: 11 (Laguna-S-2.1-NVFP4 (Temp: 1.0), Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6), Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25), Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7)… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-stat-train.
Auto-Generated FastDetector Dataset
- Dataset:
G-reen/cc-re-2020-stat-train - Globals Config:
config/globals_re2020.toml - Analysis Config:
config/analysis_exclude_deepseek.toml - Rows: 108,377
Evaluation Results
- Prompt Subsets: 4 (directreference, indirectreference, revise, rewrite)
- Generator Configs: 11 (Laguna-S-2.1-NVFP4 (Temp: 1.0), Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6), Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25), Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7), Mistral-Small-4-119B-2603-NVFP4 (Temp: 1.25), Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6), Qwen3.8-27B-AWQ-INT4 (Temp: 0.7), Qwen3.8-27B-AWQ-INT4 (Temp: 1.25), gemma-4-31B-it-AWQ-4bit (Temp: 1.0), gemma-4-31B-it-AWQ-4bit (Temp: 1.25), granite-4.2-30b-nvfp4 (Temp: 1.0))
- Classifiers: 13 (EditLens Roberta-Large Score, EditLens Roberta-Large Bucket, Perplexity (Llama-3.2-3B-Instruct) [skipped], Perplexity (Llama-3.2-3B) [skipped], Entropy (Llama-3.2-3B-Instruct) [skipped], Entropy (Llama-3.2-3B) [skipped], Top-p Outliers (Llama-3.2-3B-Instruct) [skipped], Top-p Outliers (Llama-3.2-3B) [skipped], Top-k Outliers (Llama-3.2-3B-Instruct) [skipped], Top-k Outliers (Llama-3.2-3B) [skipped], FastDetectGPT (Llama-3.2-3B-Instruct) [skipped], FastDetectGPT (Llama-3.2-3B) [skipped], Binoculars [skipped])
- Filter Conditions:
generator_model != deepseek-ai/DeepSeek-V4-Flash-0731 - Evaluation / Validation Rows: 97,539 / 10,838 (validation_size = 0.1)
- Base Columns:
original(Human),final_response(AI)
The best classifier was EditLens Roberta-Large Score with an AUROC of 0.9044. The hardest prompt subset was rewrite with a TPR of 0.2405, and the hardest generator config was Ornith-1.5-35B-A3B-NVFP4 (Temp: 0.6) with a TPR of 0.4525.
Classifier metrics averaged within each prompt and generator subset:
✔️ marks the best AUROC, ❗ the worst.
Statistics of Interest
Appendix
Table of contents 1. Univariate Analysis 2. Correlation Heatmap 3. Distance Histograms 4. Distance Histograms per Prompt Subset 5. Distance Histograms per Generator Config Subset 6. Classifier: EditLens Roberta-Large Score - Performance: - Thresholding: - Classification Histograms: 7. Classifier: EditLens Roberta-Large Bucket - Performance: - Thresholding: - Classification Histograms:
Univariate Analysis
Every statistic the report does arithmetic on, over the 97,539-row evaluation split. Invalid counts rows whose value is missing or non-finite; those rows are excluded from the other columns.
Correlation Heatmap
Pearson correlation between every statistic of interest, computed over the rows where both statistics are present.
Distance Histograms
Distance Histograms per Prompt Subset
Distance Histograms per Generator Config Subset
Classifier: EditLens Roberta-Large Score
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.5918.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Roberta-Large Bucket
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
f1with a found threshold of 0.0000.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
