CoolFace
Datasetpublic

AntonTi0108/butterboard_syst_analyst_interviews_dataset

πŸ“˜ Dataset of Interview Transcripts This dataset consists of transcribed and diarized fragments from real interviews collected from YouTube. The dialogues were processed to assign speaker roles (interviewer and candidate), chunked, and annotated using an LLM-based assistant to extract soft skills, hard skills, and personalized recommendations. 🧩 Structure Each sample in the dataset contains: instruction β€” guiding task description for the model. input β€” 4-message… See the full description on the dataset page: https://huggingface.co/datasets/AntonTi0108/butterboard_syst_analyst_interviews_dataset.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes6downloads
Dataset Card

πŸ“˜ Dataset of Interview Transcripts

This dataset consists of transcribed and diarized fragments from real interviews collected from YouTube. The dialogues were processed to assign speaker roles (interviewer and candidate), chunked, and annotated using an LLM-based assistant to extract soft skills, hard skills, and personalized recommendations.

🧩 Structure

Each sample in the dataset contains:

  • β€”instruction β€” guiding task description for the model.
  • β€”input β€” 4-message dialogue chunk between interviewer and candidate.
  • β€”output β€” dictionary with:
  • β€”hard_skills
  • β€”soft_skills
  • β€”recommendations
  • β€”chunk_id β€” index of the dialogue chunk.
  • β€”source_file β€” name of the original .json source.

πŸ“₯ Sources of Dialogue

YouTube LinkTimestampsDuration
Interview 106:30 – 51:0044 min 30 sec
Interview 201:54 – 26:0024 min 6 sec
Interview 304:42 – 20:0015 min 18 sec
Interview 422:30 – 1:21:0058 min 30 sec
Interview 508:31 – 30:15, 35:00 – 1:01:0347 min 47 sec
Interview 600:00 – 1:01:5061 min 50 sec
Interview 707:15 – 34:0526 min 50 sec

πŸ•’ Total Duration

278 minutes and 51 seconds (β‰ˆ 4 hours and 39 minutes of transcribed interviews)