datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
marketing-benchmark-of-more-than-10-ai-models
Marketing Benchmark of 10+ AI Models
A 5,000-question benchmark for evaluating LLMs across six dimensions of modern
marketing — Meta Ads, Google Ads, SEO & Organic, Email & Lifecycle, Critical
Thinking, and Action-Based scenarios — graded through 10 distinct marketer personas.
Every question is independently authored by the AdsGPT Marketing Bench team.
Knowledge MCQs are hand-authored against 2026 platform documentation; open-ended
and action-based scenarios are built from… See the full description on the dataset page: https://huggingface.co/datasets/adsgpt/marketing-benchmark-of-more-than-10-ai-models.Dolci-Instruct-SFT-translated
Dolci-Instruct-SFT-translated (Swedish)
This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project.
Dataset details
Examples: 494,841 multi-turn conversations
Language: Swedish (sv-SE)
Format: Chat/messages format (id, messages)
License: Apache 2.0
Translation
All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.Atlas-Frontier-Model-Traces
🧬 Atlas-Frontier-Model-Traces
A universal ChatML dataset distilling the agentic coding capabilities of frontier models (Kimi-K3, GPT-5.6-Sol, and Fable-5).
Dataset Description
Atlas-Frontier-Model-Traces is a meticulously curated dataset containing 15,746 coding and debugging traces generated by three of the most advanced frontier AI models.
This dataset is designed for Knowledge Distillation. By training smaller open-source language models (such as Qwen, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Pluto-AI-Labs/Atlas-Frontier-Model-Traces.ai_model_weight_poisoning_deserialization_guard_teaser
🚀 Model Security - AI Model Weight Poisoning & Deserialization Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Purchase Full Production Master Dataset on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
🌟 Domain Focus & Capabilities
Detects hidden pickle exploits, malicious backdoor tensors, and weight corruption across… See the full description on the dataset page: https://huggingface.co/datasets/emgena/ai_model_weight_poisoning_deserialization_guard_teaser.ai-redteaming-safety-model
AI Redteaming Safety Model Dataset
This dataset contains AI safety and red-teaming examples intended for evaluating, training, and improving model safety behavior.
Dataset Files
ai-safety-dataset.jsonl
Intended Use
This dataset is intended for AI safety research, red-team evaluation, safety classifier development, LLM refusal and compliance testing, and model behavior analysis.
Data Format
The dataset is provided in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/votal-ai/ai-redteaming-safety-model.ai-arenaen-conversations
AI Arenaen Conversations
A large dataset of conversations from AI-Arenaen, the Danish subset of the compar:IA platform.
Origin of the data: what is AI-Arenaen?
The conversations are collected using AI-Arenaen, the Danish entry point to the compar:IA platform, which is a Conversational AI comparison tool (a "chatbot arena"), developed within the French Ministry of Culture and adapted for Danish users by Danish Foundation Models and The ministry of digital affair.… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/ai-arenaen-conversations.aimodels.fyi-papers
BEE-spoke-data/aimodels.fyi-papers
paper overviews from https://www.aimodels.fyi/papers
dates:
pulled may 28, 2024
updated/supplemental articles added sept 21, 2024
