datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leetcodecc-re-2020-filtered
Auto-Generated FastDetector Dataset
Model Name: google/gemma-4-E4B-it
Sampling Params (as sent to the engine): {"temperature": 0.0, "top_p": 1.0, "presence_penalty": 0.0}
Ignored Params (unsupported by this engine): None
Prompt File: prompts/filter_contiguous_subset.json
Total Train Prompts: 1
Source Dataset: G-reen/cc-re-2020-raw-sharded
Source Column: text
Target Num Samples: all
Dropped Samples (over length limit 15000 tokens): 430
Failed API Requests: 495
Total… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-filtered.fastdetector-train-stat-train
Auto-generated FastDetector dataset
Fastdetector Train Stat Train
Best detectorGiga EditLens Llama-3.2-3B Score0.7121 TPR @ 1% FPRHardest prompt subsetrewrite0.3324 max detector TPR @ 1% FPRHardest generator configdeepseek-v4.1-flash (Temp: Unknown)0.4579 max detector TPR @ 1% FPR
336,087rows15generator configs4prompt subsets4detectors
01Leaderboard02Model analytics03Distances04Appendix
01Detector leaderboardScore-based detectors ranked by overall AUROC. Thresholds are placed exactly on… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/fastdetector-train-stat-train.fastdetector-test-stat-test
Auto-generated FastDetector dataset
Fastdetector Test Stat
Best detectorGiga EditLens Llama-3.2-3B Score0.6359 TPR @ 1% FPRHardest prompt subsetrewrite0.2464 max detector TPR @ 1% FPRHardest generator configclaude-opus-5 (Temp: Unknown)0.4089 max detector TPR @ 1% FPR
15,695rows10generator configs4prompt subsets15detectors
01Leaderboard02Model analytics03Distances04Appendix
01Detector leaderboardScore-based detectors ranked by overall AUROC. Thresholds are placed exactly on every human… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/fastdetector-test-stat-test.fastdetector-val-stat-val
Auto-generated FastDetector dataset
Fastdetector Val Stat
Best detectorGiga EditLens Llama-3.2-3B Score0.6638 TPR @ 1% FPRHardest prompt subsetrewrite0.3529 max detector TPR @ 1% FPRHardest generator confighy3 (Temp: Unknown)0.5385 max detector TPR @ 1% FPR
13,555rows7generator configs4prompt subsets15detectors
01Leaderboard02Model analytics03Distances04Appendix
01Detector leaderboardScore-based detectors ranked by overall AUROC. Thresholds are placed exactly on every human score, so TPR… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/fastdetector-val-stat-val.green-vla-sft-trackio-datacc-re-2021-stat-val
Auto-Generated FastDetector Dataset
Dataset: G-reen/cc-re-2021-stat-val
Globals Config: config/globals_re2021.toml
Analysis Config: config/analysis_nofilter.toml
Rows: 9,822
Evaluation Results
Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite)
Generator Configs: 5 (Hy3-NVFP4-FP8 (Temp: 0.9), Llama-4-Scout-17B-16E-Instruct-NVFP4 (Temp: 0.6), Mixtral-8x7B-Instruct-v0.1 (Temp: 0.7), claude-sonnet-5 (Temp: Unknown), gpt-5.6-luna (Temp:… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2021-stat-val.instruct-setencoder-decoder-trial-stat
Encoder/decoder trial: encoder-marginal report
Dataset: G-reen/encoder-decoder-trial-stat
Rows analysed: 122,933 (every kept (encoder, decoder, source row) triple; source G-reen/cc-re-2021-filtered shard 0, 2000 rows of at most 4000 words)
Prompt file: prompts/indirect_reference_dataset_train.json (turn 0 encodes the document, turn 1 reconstructs it from the encoding alone)
Encoders: 9 (granite-4.2-30b-nvfp4 [0], Ornith-1.5-35B-A3B-NVFP4 [1], Llama-3.3-70B-Instruct-NVFP4 [2]… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/encoder-decoder-trial-stat.cc-re-2020-stat-train
Auto-Generated FastDetector Dataset
Dataset: G-reen/cc-re-2020-stat-train
Globals Config: config/globals_re2020.toml
Analysis Config: config/analysis_exclude_deepseek.toml
Rows: 108,377
Evaluation Results
Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite)
Generator Configs: 11 (Laguna-S-2.1-NVFP4 (Temp: 1.0), Llama-3.3-70B-Instruct-NVFP4 (Temp: 0.6), Llama-3.3-70B-Instruct-NVFP4 (Temp: 1.25), Mistral-Small-4-119B-2603-NVFP4 (Temp: 0.7)… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-stat-train.fastdetector-train-filteredcc-2021-stat
Auto-Generated FastDetector Dataset
Dataset: G-reen/cc-2021-stat
Globals Config: config/globals.toml
Analysis Config: config/analysis_nofilter.toml
Rows: 23,214
Evaluation Results
Prompt Subsets: 4 (direct_reference, indirect_reference, revise, rewrite)
Generator Configs: 7 (Meta-Llama-3.1-8B-Instruct-AWQ-INT4 (Temp: 0.6), Meta-Llama-3.1-8B-Instruct-AWQ-INT4 (Temp: 1.25), Ministral-3-8B-Instruct-2512-AWQ-4bit (Temp: 0.6), Qwen3-8B-AWQ (Temp: 0.7)… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-stat.fastdetector-val-filteredSheepsScribble
Dataset Card for "SheepsScribble"
More Information needed
allen-greenough-grammar
allen-greenough-grammar
Allen and Greenough's New Latin Grammar for Schools and Colleges (J.B.
Greenough, G.L. Kittredge, A.A. Howard, Benjamin L. D'Ooge, eds.; Boston:
Ginn & Company, 1903 — public domain), in the Dickinson College
Commentaries (DCC) digital re-edition
(https://dcc.dickinson.edu/grammar/latin/), edited by Meagan Ayer under
Chris Francese's direction, 2013-2016 ("A New Allen and Greenough"). This is
a PROSE-witness t0 source repo, the standard reference grammar… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/allen-greenough-grammar.big_setfastdetector-train-rewritten-traincc-re-2021-filtered
Auto-Generated FastDetector Dataset
Model Name: google/gemma-4-E4B-it
Sampling Params (as sent to the engine): {"temperature": 0.0, "top_p": 1.0, "presence_penalty": 0.0}
Ignored Params (unsupported by this engine): None
Prompt File: prompts/filter_contiguous_subset.json
Total Train Prompts: 1
Source Dataset: G-reen/cc-re-2021-raw-sharded
Source Column: text
Target Num Samples: all
Dropped Samples (over length limit 15000 tokens): 70
Failed API Requests: 59
Total… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2021-filtered.zalo-ai-legal-text-retrieval-vn
ZacLegalTextRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Zalo Legal Text documents
Task category
t2t
Domains
Legal
Reference
https://challenge.zalo.ai/Source datasets:
GreenNode/zalo-ai-legal-text-retrieval-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("ZacLegalTextRetrieval")
evaluator = mteb.MTEB([task])
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/zalo-ai-legal-text-retrieval-vn.dbpedia-vn
DBPedia-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from DBpedia-Entity is a standard test collection for entity search over the DBpedia knowledge base The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/dbpedia-vn.SheepsScribbleV2
Dataset Card for "SheepsScribbleV2"
More Information needed
encoder-decoder-trial-rewritten
Encoder/decoder trial: decoded texts
Turn-1 outputs: each config shard_<index> is one decoder's reconstruction of every encoder's encodings from G-reen/encoder-decoder-trial-encodings. encoder_model / decoder_model name the pair; response_0 is the encoding the decoder saw. Rows failing post-processing are in shard_<index>_trashed_data; per-run counters are in runs/.
readme_card_preview_1FastDetector · Stats ReleaseG-reen/cc-2021-stat · Common Crawl CC-MAIN-2021-49Can a detector tellthe rewrite from the original?23,214 human-written web documents, each paired with an AI-generated counterpart produced by one of seven open-weight generator configurations under four prompt types, then measured with ten text-distance metrics and scored by thirteen AI-text detectors.Human text — original AI text — final_response23,214human / AI text pairs7generator configs4prompt… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/readme_card_preview_1.SheepsDiffusionNet
Dataset Card for "SheepsDiffusionNet"
More Information needed
instruct-set-halfeveryayah_curated_1s_20s_balancedeveryayah_curated_1s_20sreadme_card_preview_2
Auto-generated Fastdetector dataset
CC-2021 Stat
Analysis of G-reen/cc-2021-stat
Best detectorEditLens Roberta-Large Score0.3588TPR
Hardest prompt subsetrewrite0.0536mean TPR · 13 detectors
Hardest generator configgemma-4-E4B-it (Temp: 0.7)0.0693mean TPR · 13 detectors
Human · originalAI · final_response
23,214rows
7generator configs
4prompt subsets
13detectors
20,892 / 2,322eval / validation
01Leaderboard
02Robustness
03Distances
04Appendix
instruct-set-longercc-re-2020-rewritten-train
Auto-Generated FastDetector Dataset
Model Name: cyankiwi/Qwen3.8-27B-AWQ-INT4
Sampling Params (as sent to the engine): {"temperature": 1.25, "top_p": 1.0, "extra_body": {"top_k": -1, "top_a": 0.1, "xtc_probability": 0.3, "nsigma": 1.5, "chat_template_kwargs": {"enable_thinking": false}}}
Ignored Params (unsupported by this engine): None
Prompt File: prompts/combined_dataset_train.json
Total Train Prompts: 8000
Turn Suffix Template: The topic of the final text should be… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-rewritten-train.
