datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-detector-data
AI Detector Predictions Dataset
A continuously-growing collection of AI text detection predictions with optional user feedback, generated from the AI Text Detector Space.
Every time someone analyzes text or a URL on the Space, the prediction is appended to this dataset. Users can also click "Correct" or "Incorrect" to provide feedback, which gets stored alongside the prediction.
Schema
Field
Type
Description
id
string
Unique 12-char hex identifier… See the full description on the dataset page: https://huggingface.co/datasets/adaptive-classifier/ai-detector-data.ai-detection-dataset-v2
---dataset_info:
features:
- name: image # use the exact column name from your parquet schema
dtype: image # this forces Hugging Face to render it as an image
- name: label
dtype: string
license: other
task_categories:
- image-classification
language:
- en
tags:
- ai-generated-image-detection
- synthetic-image-detection
- diffusion-models
pretty_name: AI-Generated Image Detection Dataset v2
size_categories:
- 10K<n<100K
AI-Generated… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/ai-detection-dataset-v2.orca_dpo_data_ko
원본 데이터 : HuggingFaceH4/orca_dpo_pairs
원본 데이터 셋을 system, question, chosen, rejected 형태에 맞게 정제 후, squarelike/Gugugo-koen-7B-V1.1-AWQ 으로 번역
오번역된 데이터는 삭제 검수
ASearcher_en_no-math_Qwen3-8B-reject-sampleSee our blog for details:
Cut the Bill, Keep the Turns: Affordable Multi-Turn Search RL
This dataset is originally from https://huggingface.co/datasets/inclusionAI/ASearcher-train-data
We do filtering to original data:
Remove Chinese samples
Our wiki server does not handle Chinese retrieval well and may return garbled text;
There are about 2k Chinese-related samples in the ASearcher dataset, which we remove entirely.
Remove math problems
Use formula patterns / specific regexes to filter… See the full description on the dataset page: https://huggingface.co/datasets/aidenjhwu/ASearcher_en_no-math_Qwen3-8B-reject-sample.dota2_instruct_promptInstruction-answer dataset generated with GPT 3.5 Turbo using (html) data scrapped from fandom wiki. Data includes the following topics:
Heroes
Background lore
Attributes / Stats
Abilities
Talents
Runes
Buildings
Items
Gameplay mechanics
Creeps
Pending enhancement:
Data cleaning/preprocessing before fed into GPT 3.5 Turbo for instruction-answer set generation
Strategy data of each hero, i.e. guide to using each hero
Individual items' properties
Types of creeps in details
Types of runes… See the full description on the dataset page: https://huggingface.co/datasets/Aiden07/dota2_instruct_prompt.LocalSubs-EN2TW-SFT-Pipeline
LocalSubsEN2TW-SFT-Pipeline
SFT dataset pipeline for cue-level, context-aware English → Taiwan Traditional Chinese subtitle translation.
This repository contains:
This data card describing the pipeline and format
The manually curated diagnostic set (test_cases.jsonl)
The full processed dataset is not redistributed due to upstream licensing constraints on OpenSubtitles and TVSub source material.
Data Sources
Source
Format
Scale (before filtering)
Language… See the full description on the dataset page: https://huggingface.co/datasets/Aiden1020/LocalSubs-EN2TW-SFT-Pipeline.ai-deconditioning-synthesized-dpoThis is focused on DPO rather than other preference tuning because the truth-value of the Chosen side can be questionable;
but it is meant as a directional intervention against the over-aggressive disclaimers of the Rejected side.
Both sides generated by Lambent/arsenic-nemo-unleashed-12B initially; with varied instruction prompting for how to see itself. Dataset may expand.
ai-detector-ref-enai-deployments-2026ai-demo-translate-datasetAi_defenseaidentist
AIDentist Dataset
A dataset designed for training and fine-tuning language models to answer questions related to the AIDentist dental clinic management system.
This dataset is mainly intended for instruction-tuning (instruction → answer) tasks in Uzbek language.
📌 Dataset Purpose
The goal of this dataset is to help AI models:
Understand questions about the AIDentist platform
Provide accurate answers about system features
Assist users with platform usage
Enable… See the full description on the dataset page: https://huggingface.co/datasets/saparbayev-azizbek/aidentist.ai-detector-ref-cnai_debug_math_reasoning_dialogues_v1CB_ECLLMglm47-aider-rl8-validity-rollouts-20260723AIDetection_Vietnamese_HumanDataglm47-aider-full-v5-rl
