datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phone-ai-contention-bench
Phone AI Contention Bench
Phone AI Contention Bench is an open, real-device evaluation seed for the issues likely to define competition in phone AI:
agent permissions and cross-app control;
indirect prompt injection and screen-perception attacks;
local-versus-cloud privacy boundaries;
structured tool reliability under quantization;
cold start, memory, context, thermals, and energy;
backend and device fragmentation;
offline resilience; and
phone-to-robot command safety.
The… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/phone-ai-contention-bench.Intelligent-Content-Understanding
Intelligent Content Understanding
Empowering Advanced Thinking, Deep Understanding, Diverse Perspectives, and Creative Solutions Across Disciplines
By fostering a richly interconnected knowledge ecosystem, ICU (Intelligent Content Understanding) aims to elevate language models to unparalleled heights of understanding, reasoning, and innovation.
This ambitious project lays the groundwork for developing an 'internal knowledge map' within language models, enabling… See the full description on the dataset page: https://huggingface.co/datasets/WeMake/Intelligent-Content-Understanding.instruction-following-rl-content-constrained-30k
Instruction-Following RL Content-Constrained 30K
Dataset summary
Instruction-Following RL Content-Constrained 30K is a training dataset for precise instruction following and reinforcement learning from verifiable rewards (RLVR). The current cleaned revision contains 29,520 heterogeneous, single-turn user prompts. Each prompt combines a substantive task with one to five explicit output constraints, such as keyword inclusion or exclusion, response length… See the full description on the dataset page: https://huggingface.co/datasets/wflying/instruction-following-rl-content-constrained-30k.korean_rlhf_content_filtered
Korean RLHF Content Filtered
Dataset Summary
This dataset is a cleaned, content-only derivative of:
Source dataset: jojo0217/korean_rlhf_dataset
Source URL: https://huggingface.co/datasets/jojo0217/korean_rlhf_dataset
Each row has a single content field suitable for LM pretraining/SFT-style text modeling.
Construction
Content construction rule
For each source row:
If input is empty: content = instruction + "\n" + output
If input is not empty:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/korean_rlhf_content_filtered.enwiki_structured_content
Dataset Card for enwiki_structured_content
Dataset Description
This dataset is derived from the early official Wikipedia release, downloaded from the en subset of Wikipedia Structured Contents.Articles were converted to Markdown.
web-contentvazirweb-persian-social-media-content
VazirWeb Persian Social Media Content
A Persian-language instruction-tuning dataset for social media content generation across 27 business categories and 5 platforms, designed for fine-tuning Persian language models.
Author: Shahryar Sahebekhtiari — VazirWebGitHub: WASP-Outis
Dataset Description
Each example is a 3-turn conversation between a system prompt, a user request, and an assistant response. The assistant always returns a structured JSON object… See the full description on the dataset page: https://huggingface.co/datasets/shahryars/vazirweb-persian-social-media-content.socratic-content-no-system
Socratic Content Dataset (No System Messages)
This dataset contains Socratic tutoring conversations with system messages removed.
Data Structure
Each line contains a JSON object with the following structure:
{
"messages": [
{
"role": "user",
"content": "User's question or statement"
},
{
"role": "assistant",
"content": "Assistant's Socratic response (typically a question)"
}
]
}
Dataset Statistics
Total… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/socratic-content-no-system.
