Jianshu001/daily-conversation-batch-01-v5-clean
daily-conversation-batch-01-v5-clean Strictly filtered Arabic multi-turn conversation subset. Summary Source set: previously cleaned keep pool (A set) from arabic-daily-batch01-v5-5k-data.jsonl Second-pass reviewed records so far: 1311 Kept after strict review: 311 Dropped after strict review: 1000 Keep rate over reviewed subset: 23.72% Strict review bar Each kept sample was screened on all of: assistant hidden thinking quality visible assistant… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/daily-conversation-batch-01-v5-clean.
040
daily-conversation-batch-01-v5-clean
Strictly filtered Arabic multi-turn conversation subset.
Summary
- Source set: previously cleaned keep pool (A set) from
arabic-daily-batch01-v5-5k-data.jsonl - Second-pass reviewed records so far: 1311
- Kept after strict review: 311
- Dropped after strict review: 1000
- Keep rate over reviewed subset: 23.72%
Strict review bar
Each kept sample was screened on all of:
- assistant hidden thinking quality
- visible assistant reply quality
- overall multi-turn conversation quality
Main drop reasons included:
- prompt/checklist/meta-style hidden thinking instead of task-focused reasoning
- formulaic follow-ups and canned multi-turn progression
- generic or padded later turns
- weak grounding on references, pricing, or specific factual asks
- weak boundedness in sensitive domains
Files
data.jsonl: strict keep subset (current partial reviewed result)
