AntonTi0108/butterboard_syst_analyst_interviews_dataset
π Dataset of Interview Transcripts This dataset consists of transcribed and diarized fragments from real interviews collected from YouTube. The dialogues were processed to assign speaker roles (interviewer and candidate), chunked, and annotated using an LLM-based assistant to extract soft skills, hard skills, and personalized recommendations. π§© Structure Each sample in the dataset contains: instruction β guiding task description for the model. input β 4-messageβ¦ See the full description on the dataset page: https://huggingface.co/datasets/AntonTi0108/butterboard_syst_analyst_interviews_dataset.
π Dataset of Interview Transcripts
This dataset consists of transcribed and diarized fragments from real interviews collected from YouTube. The dialogues were processed to assign speaker roles (interviewer and candidate), chunked, and annotated using an LLM-based assistant to extract soft skills, hard skills, and personalized recommendations.
π§© Structure
Each sample in the dataset contains:
- instruction β guiding task description for the model.
- input β 4-message dialogue chunk between interviewer and candidate.
- output β dictionary with:
hard_skillssoft_skillsrecommendations- chunk_id β index of the dialogue chunk.
- source_file β name of the original
.jsonsource.
π₯ Sources of Dialogue
π Total Duration
278 minutes and 51 seconds (β 4 hours and 39 minutes of transcribed interviews)
