Kaousheik/tempo
TEMPO — Temporally-grounded Multi-task Post-training for LALMs Training and evaluation data for TEMPO, a unified large audio-language model that assigns timestamps to events, speakers and sounds across speech, sound and music. Every example is a (audio, question, answer) triple whose answer is text interleaved with atomic timestamp tokens at 0.1 s resolution (<|0.0|>, <|0.1|>, … <|60.0|>), prefixed by a task tag. Structure There is one config per task and the… See the full description on the dataset page: https://huggingface.co/datasets/Kaousheik/tempo.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face