CoolFace
21 results

disco

disco-eth /WorldSpeech WorldSpeech A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.audioautomatic-speech-recognition10M<n<100M49 likes45k downloads4mo agoHugging Facedisco-eth /EuroSpeech EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper. This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.audioautomatic-speech-recognition10M<n<100M99 likes36k downloads5mo agoHugging Facedisco-eth /eurospeech-raw-files0 likes19k downloads4mo agoHugging FaceExploration-Lab /dim-discovery-archive Geometry of Decision Making in Language Models Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025 This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models. We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.0 likes17k downloads8mo agoHugging Facemultilingual-discourse-hub /disrpt Disrpt is a multilingual, multi-framework unified discourse analysis benchmark. It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages. ⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB. To load these datasets, run the following: pip install disrpt-utils Then from disrpt_utils import load_dataset corpora_paths={ # ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.text100K<n<1M3 likes13k downloads1y agoHugging FaceSubstrateCommons /disco-replay DiSCo as a replay phantom The DiSCo substrate (Rafael-Patino, Girard, Truffet, Pizzolato, Caruyer, Thiran, The diffusion-simulated connectivity (DiSCo) dataset, Data in Brief 38 (2021) 107429, doi:10.1016/j.dib.2021.107429; data doi:10.17632/fgf86jdfg6.3, CC BY 4.0) walked once and stored as a replay pack, so that any acquisition a human scanner can play is a replay of the same walk, voxel by voxel on the dataset's own 40³ grid of 25 µm voxels, with a per-voxel Monte-Carlo… See the full description on the dataset page: https://huggingface.co/datasets/SubstrateCommons/disco-replay.0 likes7.7k downloads5d agoHugging Face