doc-ai
Datasets
All datasets matching “doc-ai”AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.hybrid-docai-receipts
Hybrid Document AI - synthetic bank documents (multi-document)
Reproducible (fixed seeds), no personal data. receipts (images/+labels.json) + statements/* + payment_orders/*.
Real receipts: SROIE 2019 (public). Models: banhchungtuongot/hybrid-docai-kie. Code: github.com/NguyenDucAnforwork/hybrid-document-ai
AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.long-doc_book_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / book / en
Available Datasets (Dataset Name: Splits):
origin-of-species_darwin: test
a-brief-history-of-time_stephen-hawking: dev
long-doc_law_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / law / en
Available Datasets (Dataset Name: Splits):
lex_files_300K-400K: test
lex_files_400K-500K: test
lex_files_500K-600K: test
lex_files_600K-700K: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / law / en
Available Datasets (Dataset Name: Splits):
lex_files_300K-400K: dev
lex_files_400K-500K: test
lex_files_500K-600K: test
lex_files_600K-700K: test
long-doc_arxiv_enAvailable Versions:
AIR-Bench_24.04
Task / Domain / Language: long-doc / arxiv / en
Available Datasets (Dataset Name: Splits):
gpt3: test
llama2: test
gemini: test
llm-survey: test
AIR-Bench_24.05
Task / Domain / Language: long-doc / arxiv / en
Available Datasets (Dataset Name: Splits):
gpt3: test
llama2: dev
gemini: test
llm-survey: test
