tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed
ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows Ready-to-train EEG↔speech windows for replicating the scaling experiment of Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data" (arXiv:2407.07595), built from the public ds007808 dataset (arXiv:2606.01264). Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch g.Pangolin) — the rig matching the 175 h paper. Each example is one… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.
ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows
Ready-to-train EEG↔speech windows for replicating the scaling experiment of Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data" (arXiv:2407.07595), built from the public [ds007808](https://openneuro.org/datasets/ds007808) dataset (arXiv:2606.01264).
Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch g.Pangolin) — the rig matching the 175 h paper. Each example is one 5-second window.
Contents (per window)
Splits: chronological 80/10/10 (train = earliest, test = latest), by global session/run/onset order — held-out test is future data, as in the paper.
Preprocessing (faithful to Sato et al. 2024)
acquire 1200 Hz → MNE denoise → NLMS adaptive filter (instantaneous; refs = EOG + upper/lower orbicularis-oris EMG; μ=0.1, ε=1e-3) → 5-s non-overlapping windows → per-window z-score along time + clip ±5 → keep windows with >20 % speech (Silero VAD). Audio 48→16 kHz; latents from frozen wav2vec2. EEG↔audio alignment uses the per-run constant wav_onset − onset offset (exact in this dataset).
Assumptions where the paper is silent (documented)
- MNE denoise cutoffs (Fig 2a not given): 50 Hz notch (+harmonics) + 0.5 Hz high-pass.
- wav2vec2 checkpoint/layers:
xlsr-53, layers 14–18 averaged. - EEG kept at the acquisition rate 1200 Hz (no downsample stated by the paper).
Load
pip install datasets librosa soundfile # librosa/soundfile decode the audio columnfrom datasets import load_dataset
ds = load_dataset("tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed")
x = ds["train"][0]
eeg = x["eeg"] # [128, 6000] float16
w2v = x["w2v"] # [249, 1024] float16 (CLIP target)
audio = x["audio"]["array"]For a CLIP retrieval model (paper): encode eeg (e.g. HTNet+Conformer) → cosine-align to w2v, zero-shot top-1/top-10 over 512 candidates.
License & citation
CC0-1.0 (derived from ds007808, CC0). Please cite ds007808 (arXiv:2606.01264) and Sato et al. 2024 (arXiv:2407.07595). Preprocessing code: this dataset's companion repo (HTNet+Conformer CLIP pipeline).
