G-reen/fastdetector-test-stat-test
Auto-Generated FastDetector Dataset Dataset: G-reen/fastdetector-test-stat-test Globals Config: config/globals_test.toml Analysis Config: config/analysis.toml Rows: 13,622 Evaluation Results Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite) Generator Configs: 10 (claude-opus-4-5-20251101 (Temp: Unknown), claude-opus-5 (Temp: Unknown), gemini-3.1-pro-preview (Temp: Unknown), gemini-3.8-flash (Temp: Unknown), gpt-5.4 (Temp: Unknown)… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/fastdetector-test-stat-test.
Auto-Generated FastDetector Dataset
- Dataset:
G-reen/fastdetector-test-stat-test - Globals Config:
config/globals_test.toml - Analysis Config:
config/analysis.toml - Rows: 13,622
Evaluation Results
- Prompt Subsets: 4 (directreference, indirectreference, revise, rewrite)
- Generator Configs: 10 (claude-opus-4-5-20251101 (Temp: Unknown), claude-opus-5 (Temp: Unknown), gemini-3.1-pro-preview (Temp: Unknown), gemini-3.8-flash (Temp: Unknown), gpt-5.4 (Temp: Unknown), gpt-5.6-sol (Temp: Unknown), grok-4.3 (Temp: Unknown), kimi-k3 (Temp: Unknown), minimax-m3 (Temp: Unknown), qwen3.7-max (Temp: Unknown))
- Classifiers: 19 (EditLens Roberta-Large Score, EditLens Roberta-Large Bucket, EditLens Llama-3.2-3B Score, EditLens Llama-3.2-3B Bucket, EditLens Giga RoBERTa-Large Score, EditLens Giga RoBERTa-Large Bucket, EditLens Giga Llama-3.2-3B Score, EditLens Giga Llama-3.2-3B Bucket, Perplexity (Llama-3.2-3B-Instruct), Perplexity (Llama-3.2-3B), Entropy (Llama-3.2-3B-Instruct), Entropy (Llama-3.2-3B), Top-p Outliers (Llama-3.2-3B-Instruct), Top-p Outliers (Llama-3.2-3B), Top-k Outliers (Llama-3.2-3B-Instruct), Top-k Outliers (Llama-3.2-3B), FastDetectGPT (Llama-3.2-3B-Instruct), FastDetectGPT (Llama-3.2-3B), Binoculars)
- Filter Conditions:
cosdist >= 0.03ORsoftngram >= 0.06 - Evaluation / Validation Rows: 12,259 / 1,363 (validation_size = 0.1)
- Base Columns:
original(Human),final_response(AI)
The best classifier was EditLens Giga Llama-3.2-3B Score with an AUROC of 0.9442. The hardest prompt subset was rewrite with a TPR of 0.0962, and the hardest generator config was claude-opus-5 (Temp: Unknown) with a TPR of 0.1190.
Classifier metrics averaged within each prompt and generator subset:
✔️ marks the best AUROC, ❗ the worst.
Statistics of Interest
Appendix
Table of contents 1. Univariate Analysis 2. Correlation Heatmap 3. Distance Histograms 4. Distance Histograms per Prompt Subset 5. Distance Histograms per Generator Config Subset 6. Classifier: EditLens Roberta-Large Score - Performance: - Thresholding: - Classification Histograms: 7. Classifier: EditLens Roberta-Large Bucket - Performance: - Thresholding: - Classification Histograms: 8. Classifier: EditLens Llama-3.2-3B Score - Performance: - Thresholding: - Classification Histograms: 9. Classifier: EditLens Llama-3.2-3B Bucket - Performance: - Thresholding: - Classification Histograms: 10. Classifier: EditLens Giga RoBERTa-Large Score - Performance: - Thresholding: - Classification Histograms: 11. Classifier: EditLens Giga RoBERTa-Large Bucket - Performance: - Thresholding: - Classification Histograms: 12. Classifier: EditLens Giga Llama-3.2-3B Score - Performance: - Thresholding: - Classification Histograms: 13. Classifier: EditLens Giga Llama-3.2-3B Bucket - Performance: - Thresholding: - Classification Histograms: 14. Classifier: Perplexity (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 15. Classifier: Perplexity (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 16. Classifier: Entropy (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 17. Classifier: Entropy (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 18. Classifier: Top-p Outliers (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 19. Classifier: Top-p Outliers (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 20. Classifier: Top-k Outliers (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 21. Classifier: Top-k Outliers (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 22. Classifier: FastDetectGPT (Llama-3.2-3B-Instruct) - Performance: - Thresholding: - Classification Histograms: 23. Classifier: FastDetectGPT (Llama-3.2-3B) - Performance: - Thresholding: - Classification Histograms: 24. Classifier: Binoculars - Performance: - Thresholding: - Classification Histograms:
Univariate Analysis
Every statistic the report does arithmetic on, over the 12,259-row evaluation split. Invalid counts rows whose value is missing or non-finite; those rows are excluded from the other columns.
Correlation Heatmap
Pearson correlation between every statistic of interest, computed over the rows where both statistics are present.
Distance Histograms
Distance Histograms per Prompt Subset
Distance Histograms per Generator Config Subset
Classifier: EditLens Roberta-Large Score
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.4747.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Roberta-Large Bucket
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
f1with a found threshold of 0.0000.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Llama-3.2-3B Score
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.2389.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Llama-3.2-3B Bucket
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
f1with a found threshold of 0.0000.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Giga RoBERTa-Large Score
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.1910.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Giga RoBERTa-Large Bucket
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
f1with a found threshold of 0.0000.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Giga Llama-3.2-3B Score
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.1014.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: EditLens Giga Llama-3.2-3B Bucket
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
f1with a found threshold of 0.0000.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Perplexity (Llama-3.2-3B-Instruct)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 4.7220.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Perplexity (Llama-3.2-3B)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 3.9421.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Entropy (Llama-3.2-3B-Instruct)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 1.4637.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Entropy (Llama-3.2-3B)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 1.4228.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Top-p Outliers (Llama-3.2-3B-Instruct)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.0312.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Top-p Outliers (Llama-3.2-3B)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.0201.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Top-k Outliers (Llama-3.2-3B-Instruct)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.0295.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Top-k Outliers (Llama-3.2-3B)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.0183.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: FastDetectGPT (Llama-3.2-3B-Instruct)
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 3.2271.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: FastDetectGPT (Llama-3.2-3B)
Performance:
Thresholding:
- Direction:
lower_is_ai - Swept for
fpr_0_5pctwith a found threshold of -2.8000.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
Classifier: Binoculars
Performance:
Thresholding:
- Direction:
higher_is_ai - Swept for
fpr_0_5pctwith a found threshold of 0.9673.
Classification Histograms:
Per Prompt Subset
Per Generator Config Subset
