datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
una-fraza-al-diya
Una fraza al diya
Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total.
Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/
Citation
If you use this dataset, please cite:
Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish
Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.universeset
UniVerseSet
The training split of UniVerse (同谣).Held-out evaluation lives in UniVerseBench.
UniVerseSet is the post-training corpus for large audio–language models on world folk music: ASR, captions, and audio-grounded chat, plus automatically transcribed ABC scores.
「诗言志,歌永言,声依永,律和声。」—《尚书·舜典》
Sister dataset (benchmark)
universe-team/universebench
Live museum demo
http://143.89.224.8:8790/
What's here
Archives (download and unpack;… See the full description on the dataset page: https://huggingface.co/datasets/universe-team/universeset.
