disco
Datasets
All datasets matching “disco”WorldSpeech
WorldSpeech
A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/WorldSpeech.EuroSpeech
EuroSpeech Dataset
Dataset Description
EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non-verbatim transcripts. More information can be found in the paper.
This dataset is 16 kHz, the 24 kHz version of EuroSpeech can be found at… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/EuroSpeech.eurospeech-raw-filesdim-discovery-archive
Geometry of Decision Making in Language Models
Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025
This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models.
We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.disrpt
Disrpt is a multilingual, multi-framework unified discourse analysis benchmark.
It unifies discourse relation classification tasks (.rels) and discourse segmentation (.connlu) for many languages.
⚠️ This repo only contains the disrpt dataset when the underlying data is permissively licensed. Some datasets rely on corpora like the PTB.
To load these datasets, run the following:
pip install disrpt-utils
Then
from disrpt_utils import load_dataset
corpora_paths={
# ⚠️✍️ TODO Input… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-discourse-hub/disrpt.disco-replay
DiSCo as a replay phantom
The DiSCo substrate (Rafael-Patino, Girard, Truffet, Pizzolato, Caruyer, Thiran, The diffusion-simulated
connectivity (DiSCo) dataset, Data in Brief 38 (2021) 107429, doi:10.1016/j.dib.2021.107429; data
doi:10.17632/fgf86jdfg6.3, CC BY 4.0) walked once and stored as a replay pack, so that any acquisition a human
scanner can play is a replay of the same walk, voxel by voxel on the dataset's own 40³ grid of 25 µm voxels, with
a per-voxel Monte-Carlo… See the full description on the dataset page: https://huggingface.co/datasets/SubstrateCommons/disco-replay.
