Yashodhar29/synthetic-mirrors-human-ai-cnn-dailymail-v1
Synthetic Mirrors: Human vs AI (CNN DailyMail) π Overview Synthetic Mirrors is a multi-model, research-grade dataset designed for detecting humanized AI-generated text. This dataset aggregates AI generations from multiple open-source and closed-source LLMs, paired with corresponding human-written content from the CNN DailyMail domain. The core idea is to treat each AI model as a synthetic mirror β reflecting distinct stylistic and probabilistic artifacts thatβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Yashodhar29/synthetic-mirrors-human-ai-cnn-dailymail-v1.
Synthetic Mirrors: Human vs AI (CNN DailyMail)
π Overview
Synthetic Mirrors is a multi-model, research-grade dataset designed for detecting humanized AI-generated text.
This dataset aggregates AI generations from multiple open-source and closed-source LLMs, paired with corresponding human-written content from the CNN DailyMail domain.
The core idea is to treat each AI model as a synthetic mirror β reflecting distinct stylistic and probabilistic artifacts that help detectors generalize to unseen models.
π€ AI Models Used
- GPT-3.5 Turbo
- Qwen/Qwen2.5-1.5B-Instruct
- HuggingFaceTB/SmolLM2-360M-Instruct
- HuggingFaceTB/SmolLM2-1.7B-Instruct
- Google/Gemma-2-2B-IT
- Meta-LLaMA-3.2-1B-Instruct
- Meta-LLaMA-3.1-8B-Instruct
π§± Dataset Schema
Each row corresponds to one text sample.
π― Intended Use
- AI-generated text detection
- Humanized AI content analysis
- Cross-model generalization studies
- Leave-one-model-out evaluation
- Training robust AI detectors
β οΈ Limitations
- Single domain (news)
- English only
- AI generations may reflect prompt biases from original datasets
π Citation
If you use this dataset, please cite: @dataset{syntheticmirrors2025, title = {Synthetic Mirrors: Human vs AI (CNN DailyMail)}, author = {Chavan, Yashodhar}, year = {2025}, platform = {Hugging Face} }
