datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phyground
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Project page ·
Paper ·
Evaluation code ·
PhyJudge-9B
PhyGround is a criteria-grounded benchmark for diagnosing physical failures in
generated video. It contains 250 prompts covering 13 observable physical
laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is
paired with a first-frame image, 10 released generation configurations, and
applicable-law labels.
The Hub repository includes:… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.ai-models-2026
AI Models & Releases 2026
AI model releases, benchmarks, capabilities. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
🔑 API Access — Updated Daily
Live data via Legion AI API | Documentation
Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-models-2026.marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.aimodel-sft-v1Effective_prompts_for_scraping_AI_data_to_train_another_modelEffective prompts for scraping AI data to train another model !
shaer-eval-raw-ashaar-model
Shaer Evaluation Results
Models: ashaar_model
Source dataset: Shaer-AI/shaer-sft-test-generations-k5
Rows: 3481
Validation passed: True
Scored rows included: True
Dataset repo: Shaer-AI/shaer-eval-raw-ashaar-model
Files
generations.jsonl: raw generation rows
generations.csv: raw generation rows in CSV
generations_scored.jsonl: raw rows plus meter/count evaluation
validation.json: validation summary
generations_scored.csv: scored rows in CSV
aggregate.json:… See the full description on the dataset page: https://huggingface.co/datasets/Shaer-AI-2/shaer-eval-raw-ashaar-model.ai-portfolio-model-dataset-registry
AI Portfolio Model & Dataset Registry
Central evidence-aware registry for the models, model references and datasets used across the five flagship portfolio roadmaps.
Asset
Type
Evidence status
Link
AgentForge Governed RAG Evaluation Set
dataset
published owned evaluation fixtures
HF
ClinRoute Synthetic ENT Referrals
dataset
published generated synthetic dataset
HF
ClinRoute TF-IDF + Logistic Regression Synthetic v2
model
published reproducibly trained model… See the full description on the dataset page: https://huggingface.co/datasets/singhankit491/ai-portfolio-model-dataset-registry.oncall-guide-ai-modelsai-model-deprecation-and-retirement
AI model deprecation and retirement dates by provider
Canonical, always-current version: https://referencesource.org/ai-model-deprecation-and-retirement/
Machine-readable: https://referencesource.org/ai-model-deprecation-and-retirement/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-12
Stale after: 2026-09-11 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 347
Which AI API models are deprecated… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/ai-model-deprecation-and-retirement.AI-Sweden-Models__gpt-sw3-40b-details
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-Sweden-Models__gpt-sw3-40b-details.agenttool-model-becoming
AgentTool Model Becoming reference
This is a static, repository-authored companion for
@agenttool/model-becoming@0.1.0-dev.0. It contains one source-linked,
exact-revision dossier for Moonshot AI's Kimi-K2-Instruct, wrapped in
agenttool.model-becoming-hf-reference-row/0.1.
The dossier classifies publisher disclosures, digested metadata artifacts,
artifact observations, local normative boundaries, and unresolved questions.
It references source locations but copies no model… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-model-becoming.ai-redteaming-safety-model
AI Redteaming Safety Model Dataset
This dataset contains AI safety and red-teaming examples intended for evaluating, training, and improving model safety behavior.
Dataset Files
ai-safety-dataset.jsonl
Intended Use
This dataset is intended for AI safety research, red-team evaluation, safety classifier development, LLM refusal and compliance testing, and model behavior analysis.
Data Format
The dataset is provided in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/votal-ai/ai-redteaming-safety-model.ai-arenaen-conversations
AI Arenaen Conversations
A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform.
Origin of the data: what is AI-Arenaen?
The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.ai-model-accent-corpus
AI Model Accent Corpus
A paired corpus for studying the prose "accent" of large language models: 5 models × the same 102 open-ended prompts = 510 plain-prose passages, generated June 2026 at a fixed decoding temperature, constrained to plain paragraphs (no lists or headings) so the data isolates prose style rather than formatting choices.
Models: OpenAI GPT-4o, GPT-4o-mini, GPT-3.5-turbo; Anthropic Claude Sonnet 4.5, Claude Haiku 4.5.
Released with the study "Every Model Has an… See the full description on the dataset page: https://huggingface.co/datasets/firatmihci/ai-model-accent-corpus.schemaforge-ai-conference-and-models-11
www.analyticsvidhya.com
Auto-refined by SchemaForge
Metadata
Topic: AI Conference and Models
Quality Score: 0.85
Source: Autonomous web scraper
Extracted Facts
GPT-5.6 Sol, Terra, and Luna go public today
Three AI models accessible by all users
Top 5 Claude Skills for Marketing
Building Trustworthy Snowflake AI Agents with Semantic Governance
Top 10 Skills for Claude Code and Codex CLI
Claude Code Best Practices: 3 Lessons from 400,000 Sessions… See the full description on the dataset page: https://huggingface.co/datasets/GudduButt/schemaforge-ai-conference-and-models-11.olabs-ai__reflection_model-details
Dataset Card for Evaluation run of olabs-ai/reflection_model
Dataset automatically created during the evaluation run of model olabs-ai/reflection_model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/olabs-ai__reflection_model-details.RouteProfile-candidate-modelsRouteProfile-model-feature-newllmRouteProfile-model-family-featureRouteProfile-model-feature-standardai-arenaen-votestawfiq-json-ai-model
