CoolFace
Datasetpublic

Jianshu001/daily-conversation-batch-01-v5-clean

daily-conversation-batch-01-v5-clean Strictly filtered Arabic multi-turn conversation subset. Summary Source set: previously cleaned keep pool (A set) from arabic-daily-batch01-v5-5k-data.jsonl Second-pass reviewed records so far: 1311 Kept after strict review: 311 Dropped after strict review: 1000 Keep rate over reviewed subset: 23.72% Strict review bar Each kept sample was screened on all of: assistant hidden thinking quality visible assistant… See the full description on the dataset page: https://huggingface.co/datasets/Jianshu001/daily-conversation-batch-01-v5-clean.

sourceHugging Faceotherupdated 5mo agoView on Hugging Face
0likes40downloads
Dataset Card

daily-conversation-batch-01-v5-clean

Strictly filtered Arabic multi-turn conversation subset.

Summary

  • —Source set: previously cleaned keep pool (A set) from arabic-daily-batch01-v5-5k-data.jsonl
  • —Second-pass reviewed records so far: 1311
  • —Kept after strict review: 311
  • —Dropped after strict review: 1000
  • —Keep rate over reviewed subset: 23.72%

Strict review bar

Each kept sample was screened on all of:

  • —assistant hidden thinking quality
  • —visible assistant reply quality
  • —overall multi-turn conversation quality

Main drop reasons included:

  • —prompt/checklist/meta-style hidden thinking instead of task-focused reasoning
  • —formulaic follow-ups and canned multi-turn progression
  • —generic or padded later turns
  • —weak grounding on references, pricing, or specific factual asks
  • —weak boundedness in sensitive domains

Files

  • —data.jsonl: strict keep subset (current partial reviewed result)