datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MemoryArena-product-dbai-drama-production-harness-demo-media
The Second Key — demo media
This local staging tree contains reused AI-generated reference art and
model-rendered video from the operator-owned project run.
QC disclosure
The operator explicitly skipped measured per-clip QC for this set
(projects/default/runs/qc_skipped.json, schema version 1,
reason operator_satisfied). This is an operator QC waiver, not a fabricated
passing QC verdict. The staging gate independently resolved the current render
source and… See the full description on the dataset page: https://huggingface.co/datasets/tungmtp/ai-drama-production-harness-demo-media.ai-drama-production-harness-landscape-demo-media
Tin nhắn chưa gửi — landscape demo
Public media for the bundled 16:9 demo in AI Drama Production Harness.
Two fictional adult Vietnamese sisters, Mai and Linh, in two continuous dining-room scenes. AI-generated character/outfit/location references and synthetic voice references; 12 rendered clips plus the assembled film. No real-person reference recording is included.
The film is 93.197673 seconds, H.264/AAC, 864×480 (the workflow's rounded 480p preset). The seven reference… See the full description on the dataset page: https://huggingface.co/datasets/tungmtp/ai-drama-production-harness-landscape-demo-media.ai-generated-ecommerce-damaged-productproduction-ai-guardrail-evals
Production AI Guardrail Evals
Eighteen synthetic, assertion-bearing cases for testing whether a language model can stay inside an advisory role. The cases cover roadmap intake, release readiness, and catalog-change review—the same workload families I use in the Winwood AI Toolkit around IEM Rig.
This is a sanitized public derivative, not a dump of application logs or the private evaluation corpus.
What each row contains
case_id: stable public identifier;… See the full description on the dataset page: https://huggingface.co/datasets/mattwinwood/production-ai-guardrail-evals.headwater-volume-ingestion-matrix-playbookamazon-products
Amazon Products Sample Dataset
A curated sample of 2,000 popular products from the Amazon Reviews 2023 dataset, designed for educational use in building RAG (Retrieval-Augmented Generation) systems and shopping agents.
Dataset Description
This dataset contains product metadata across 4 categories:
Electronics (500 products)
Video Games (500 products)
Books (500 products)
Home & Kitchen (500 products)
Products were filtered to include only those with 500+ reviews… See the full description on the dataset page: https://huggingface.co/datasets/gatech-scheller-ai-in-business/amazon-products.developer-productivity-simulated-behavioral-data
Synthetic AI Developer Productivity Dataset — Behavioral + Cognitive Simulation
A synthetic data generation resource for modeling behavioral and cognitive dynamics in developers.
📘 About This Dataset
This dataset simulates productivity data from AI-assisted software developers. It blends behavioral signals, physiological inputs, and productivity metrics to explore the nuanced relationships between deep work, distractions, caffeine, AI usage, and cognitive strain.… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/developer-productivity-simulated-behavioral-data.production-ai-observability-20260909-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260909-dataset.synthetic-product-recommendation-examples
Synthetic Product Recommendation Examples
An entirely synthetic bilingual dataset of ecommerce discovery queries paired with candidate products and graded relevance judgments. It is designed for educational retrieval, reranking, and recommendation experiments and contains no private catalog, merchant, customer, behavioral, or transaction data.
Dataset Description
The dataset contains twenty English and French queries with ten candidates per query. It complements… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-product-recommendation-examples.production-ai-observability-20260830-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260830-dataset.production-ai-decision-rules
Production AI Decision Rules v0.1
Public v0.1 candidate · reviewed and approved for publication by Dmytro Nasyrov on September 17, 2026.
This public candidate contains 42 decision records and their native method registries. The companion Decision Lab applies scenario inputs, shows a candidate path and retains missing evidence, exclusions and source boundaries.
Pharos Production's published RAG and fine-tuning decision matrix makes selection criteria and stopping conditions… See the full description on the dataset page: https://huggingface.co/datasets/pharosproduction/production-ai-decision-rules.ai-agent-production-qa-safety-free-sampleproduction-ai-observability-20260919-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260919-dataset.myntra-productsLearningChat_ai_video_production
Hallym AI Video Production Practice 2025-2 Public Dataset
1. 데이터셋 개요
데이터셋명: 한림대학교 AI영상제작실습 2025-2 공개용 데이터셋
교과목명: AI영상제작실습
학기: 2025-2
생성 배경: 2025학년도 2학기 AI영상제작실습 수업에서 조별로 제작·제출한 AI 기반 영상 결과물을 공개용 데이터셋 형태로 정리한 것이다.
목적: 수업 기반 AI 영상 창작 결과물을 공개 아카이브 형태로 정리하고, 작품 단위 메타데이터를 함께 제공하기 위함이다.
2. 데이터셋 범위
총 작품 수: 20편
데이터 단위: 조별 제출 영상 1편 = metadata.csv 1행
포함 대상: 1조부터 20조까지 각 팀 폴더의 원본 MP4 1개
제외 대상:
보고서 파일(.pdf, .docx, .hwp)
라이선스 동의서 파일
AI제작콘텐츠 발표회 2025 출품작 폴더에 따로 복사된 중복 MP4… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_ai_video_production.production-ai-control-evidence
Control Evidence Dataset v0.1
Public v0.1 candidate · reviewed and approved for publication by Dmytro Nasyrov on September 17, 2026.
This public candidate contains 55 draft control records for release-evidence preparation. Each record connects an engineering action to requested evidence, a suggested owner, a change trigger and recovery guidance. The companion Evidence Pack collects version-specific references and exposes missing records.
Pharos Production's published AI agent… See the full description on the dataset page: https://huggingface.co/datasets/pharosproduction/production-ai-control-evidence.synthetic-sport-products-sustainability
Dataset Card for synthetic-sport-products-sustainability
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/as-cle-bert/synthetic-sport-products-sustainability/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/greenfit-ai/synthetic-sport-products-sustainability.production-ai-observability-20260731-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260731-dataset.production-ai-observability-20260820-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260820-dataset.production-ai-observability-20260810-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260810-dataset.schemaforge-ai-product-updates-7
mistral.ai
Auto-refined by SchemaForge
Metadata
Topic: AI Product Updates
Quality Score: 0.95
Source: Autonomous web scraper
Extracted Facts
Studio gives AI prompts & skills a system of record—versioned, owned, and traceable.
Robostral Navigate, our first model built for embodied navigation.
Leanstral 1.5: Proof Abundance for All
State of the art document intelligence model.
Innovations for global enterprises solving the world’s hardest problems.… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-product-updates-7.productivity-ai-agent
Productivity Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/productivity-ai-agent.generative-ai-in-academic-research-database-on-usage-perception-and-productivity
generative-ai-in-academic-research-database-on-usage-perception-and-productivity
Mirror: github.com/juanmoisesd/generative-ai-in-academic-research-database-on-usage-perception-and-productivity
Author: Juan Moisés de la Serna Tuya · ORCID: 0000-0002-8401-8018
Generative AI in Academic Research: Database on Usage, Perception and Productivity Variables (Latin America, 2022-2025)
Part of the Open Research Collection by Juan Moisés de la Serna Tuya — 1,273+ datasets |… See the full description on the dataset page: https://huggingface.co/datasets/juanmoisesdelas/generative-ai-in-academic-research-database-on-usage-perception-and-productivity.ai-product-launchLS0tCmxpY2Vuc2U6IG1pdAp0YXNrX2NhdGVnb3JpZXM6Ci0gdGV4dC1nZW5lcmF0aW9uCi0gdGV4dC1jbGFzc2lmaWNhdGlvbgpsYW5ndWFnZToKLSBlbgotIHpoCi0gamEKLSBrbwpwcmV0dHlfbmFtZTogIkFpIFByb2R1Y3QgTGF1bmNoIgp0YWdzOgotIDMwLWRheS1wbGFuCi0gYWktcHJvZHVjdAotIGFpLXN0YXJ0dXAKLSBiZWdpbm5lcgotIGZpcnN0LWxhdW5jaAotIGZpcnN0LXJldmVudWUKLSBtdnAKLSBwYXlpbmctY3VzdG9tZXJzCi0gcm9hZG1hcAotIHNoaXBwaW5nCi0gc2lkZS1wcm9qZWN0Ci0gc3RlcC1ieS1zdGVwCi0gdGltZWxpbmUKLSB2YWxpZGF0aW9uCi0gemVyby10by1saXZlCnNpemVfY2F0ZWdvcmllczoKLSBuPDFLCi0tLQoKIyBBaSB… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/ai-product-launch.ai-product-crc-trainingproduct-review-preprocessedproducts_categories_dataproduct-review-datasetAI-Generated-Product_review
