datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmpreal-fake-ai-generated-art-images
🎨 Real and Fake (AI-Generated) Art Images Dataset
21,642 balanced images — 10,821 real artworks and 10,821 AI-generated
images — for training models to distinguish authentic art from GAN-generated fakes.
🧭 Overview
This dataset is part of the FauxFinder project, designed to build
advanced models capable of distinguishing between authentic artworks
and AI-generated images. Ideal for binary classification, GAN research,
and computer vision benchmarking.… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/real-fake-ai-generated-art-images.mjnj384NPM-Artifacts-zh
NPM-Artifacts-zh: National Palace Museum Open Artifacts Dataset
Dataset Description
This project collects and organizes public artifact data from the National Palace Museum Open Data Platform. The dataset contains high-resolution images of artifacts and their corresponding rich, structured metadata. All metadata is in Traditional Chinese, detailing information such as the artifact's name, dynasty, dimensions, materials, inscriptions, and seal impressions.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/danqing-ai/NPM-Artifacts-zh.eu-ai-act-article-50-scoreboard
Article 50 historical public-evidence snapshot
This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set.
Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.ai-text-detection-pile
Dataset Card for AI Text Dectection Pile
Dataset Summary
This is a large scale dataset intended for AI Text Detection tasks, geared toward long-form text and essays. It contains samples of both human text and AI-generated text from GPT2, GPT3, ChatGPT, GPTJ.
Here is the (tentative) breakdown:
Human Text
Dataset
Num Samples
Link
Reddit WritingPromps
570k
Link
OpenAI Webtext
260k
Link
HC3 (Human Responses)
58k
Link
ivypanda-essays
TODO
TODO… See the full description on the dataset page: https://huggingface.co/datasets/artem9k/ai-text-detection-pile.384imagenet1kk1921kk imagenet with auradiffusion vae 192x192
see https://huggingface.co/datasets/AiArtLab/imagenet1kk192/blob/main/imagenet-1kk/sdxs1b-imagenet.ipynb
7681152576640384dataset640imagenet-1kkUsage: https://huggingface.co/AiArtLab/sdxs/blob/main/imagenet.ipynb
ai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web.
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import load_dataset
dataset = load_dataset("siavava/ai-tech-articles")
# optionally, convert it to a pandas dataframe:
df = dataset["train"].to_pandas()
You do not need to clone… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.640_mjnjTheArabicPile_Articles
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Articles.AI_Image_Artifacts_Present640qft-pixiv-ai-artistsdspai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.gguf-jlens-artifacts
gguf-jlens artifacts
Heavy artifacts for the gguf-jlens project (Jacobian lens
on GGUF-quantized Qwen3.5-4B) — kept out of git and mirrored here.
lenses/ — fitted Jacobian lenses (*.pt) and resume checkpoints (*.ckpt.pt):
q8_0, q4_k_m, q3_k_m, q2_k, bf16-control (+ -s2 disjoint-sample replicas).
models/ — GGUFs quantized locally from the unsloth BF16 file (Q3_K_M, Q2_K).
out/ — per-item rank dumps (ranks/, ranks5/, ranks6/), slice pages, fit logs.
Restore into a clone with… See the full description on the dataset page: https://huggingface.co/datasets/mozilla-ai/gguf-jlens-artifacts.ai_art_np
Dataset Card for "ai_art_np"
More Information needed
aiart_channel_nai3_geachuaesthetic_6k1024
MidJourney
NijiJourney
E-Shushu
Compressed by a factor of 64 pixels in JPEG, 97 quality, maximum side length of 1024.
Mixed labeling using different models:
Human prompts in MJ/NJ
Long captions (LLaVA)
Short captions (LLaVA + LLaMA)
Medium captions (Moondream)
physical-ai-bench-artifactsai-jobs-news-articles-abstracts
News articles and research abstracts on AI, labor, and jobs
Dataset summary
This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary.
Rows: 53,526
document_class
Rows
Approx. date range (date column)
news
29,857
Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.
