datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tmpmjnj3847681152576640384dataset640ai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web.
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import load_dataset
dataset = load_dataset("siavava/ai-tech-articles")
# optionally, convert it to a pandas dataframe:
df = dataset["train"].to_pandas()
You do not need to clone… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.640_mjnj640qft-pixiv-ai-artistsai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.dataset768576dataset576_10ksampleSample: https://huggingface.co/datasets/AiArtLab/dataset576_10ksample/blob/main/Untitled.ipynb
ai-oversight-medical-harm-detection
AI Oversight: Medical Harm Detection Dataset
75 examples of multi-turn medical/health conversations designed to evaluate human-AI complementarity in detecting subtle harms in AI-generated medical advice. Grounded in verified medical literature with full provenance.
Project
Human-AI Complementarity for Identifying Harm — SPAR 2026 Spring
Mentor: Rishub Jain, Google DeepMind
Area: Scalable Oversight
Goal: Show a large increase in Complementarity on a diverse set of tasks… See the full description on the dataset page: https://huggingface.co/datasets/ArthT/ai-oversight-medical-harm-detection.Article-Risk-Factor-classification-data-v1ai-economy-labor-articles-annotated-processedAI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.260412_collective_intelligence_OR_social_intelligence-_AND_-artificial_intelligence_OR_ai-_AND260408_2_artificial_intelligence_OR_ai_AND_human_AND_team_AND_simulation_OR_interactionaadigupta1601_ai-vs-human-art-interpretable-numerical-features
AI vs Human Art — Interpretable Numerical Features
An Interpretable Numerical Representation of Visual Artworks for Comparative Ana
Dataset Info
Source: Kaggle
Original Size: 0.08 MB
Kaggle Downloads: 37
Files: 1
Files
ai_vs_human_art_interpretable_numeric.csv
Mirrored from Kaggle
260408_artificial_intelligence_OR_ai-_AND_human_AND_team_AND_-simulation_OR_interactionArticle-Risk-Factor-Classification-Dataset260408_artificial_intellige_OR_ai_AND_human_AND_team_AND
