datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eu-ai-act-article-50-scoreboard
Article 50 historical public-evidence snapshot
This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set.
Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.ai-jobs-news-articles-abstracts
News articles and research abstracts on AI, labor, and jobs
Dataset summary
This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary.
Rows: 53,526
document_class
Rows
Approx. date range (date column)
news
29,857
Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.AI-ArtTools-Pack
AI ArtTools Pack v1.0 — 372 Styles / 23 Categories
Developer and artist utility pack for Stable Diffusion XL.
Not for generating pretty pictures — for generating usable production assets.
Compatible with Style Grid Organizer extension.
Contents
372 prompt styles across 23 categories
CSV format (Forge/A1111 compatible)
Covers the full production pipeline from rough concept to final asset
Categories
Category
Count
Purpose
ASSET
41
Weapons, props, UI… See the full description on the dataset page: https://huggingface.co/datasets/Kazzze/AI-ArtTools-Pack.arthur_sensitive_data_passwordarthur_prompt_injection_benchmarkAI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.python_wiki_hallucination_graded
RAG + Instruction Following Results from Python Wikipedia benchmark
This dataset is an artifact from an experiment conducted by Arthur
Experiment
We wanted to compare how good LLMs are at answering questions using a context. Doing this task well involves an inverse skill: recognizing when the necessary information to answer a question is absent, and choosing instead to not answer. One name for this is “staying grounded” in the context that you provide in your prompt to… See the full description on the dataset page: https://huggingface.co/datasets/Arthur-AI/python_wiki_hallucination_graded.arthur_toxicity_benchmarkhallucination_dolly_benchmarkhallucination_wikibiodemo-articles-and-summary25_Articles-Affiliate-Marketing
