CoolFace
Datasetpublic

tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed

ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows Ready-to-train EEG↔speech windows for replicating the scaling experiment of Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data" (arXiv:2407.07595), built from the public ds007808 dataset (arXiv:2606.01264). Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch g.Pangolin) — the rig matching the 175 h paper. Each example is one… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed.

sourceHugging Facecc0-1.0updated 3mo agoView on Hugging Face
0likes189downloads
Dataset Card

ds007808 sub-01 / speechopen / pangolin — preprocessed EEG↔speech windows

Ready-to-train EEG↔speech windows for replicating the scaling experiment of Sato et al. 2024, "Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data" (arXiv:2407.07595), built from the public [ds007808](https://openneuro.org/datasets/ds007808) dataset (arXiv:2606.01264).

Slice = subject sub-01, task speechopen (overt speech), device pangolin (128-ch g.Pangolin) — the rig matching the 175 h paper. Each example is one 5-second window.

Contents (per window)

fieldshape / typedescription
eegfloat16 [128, 6000]128-ch EEG @1200 Hz, NLMS-cleaned, z-scored along time, clipped ±5
w2vfloat16 [249, 1024]frozen wav2vec2-large-xlsr-53 latents (layers 14–18 averaged)
audio16 kHz monothe aligned 5 s of spoken audio
transcriptstringJapanese transcript overlapping the window
onset,offsetfloatEEG-time onset (s) and audio-time offset (s) of the window
vadfloatfraction of speech in the window (Silero VAD)
session,runstringprovenance

Splits: chronological 80/10/10 (train = earliest, test = latest), by global session/run/onset order — held-out test is future data, as in the paper.

Preprocessing (faithful to Sato et al. 2024)

acquire 1200 Hz → MNE denoise → NLMS adaptive filter (instantaneous; refs = EOG + upper/lower orbicularis-oris EMG; μ=0.1, ε=1e-3) → 5-s non-overlapping windows → per-window z-score along time + clip ±5 → keep windows with >20 % speech (Silero VAD). Audio 48→16 kHz; latents from frozen wav2vec2. EEG↔audio alignment uses the per-run constant wav_onset − onset offset (exact in this dataset).

Assumptions where the paper is silent (documented)

  • —MNE denoise cutoffs (Fig 2a not given): 50 Hz notch (+harmonics) + 0.5 Hz high-pass.
  • —wav2vec2 checkpoint/layers: xlsr-53, layers 14–18 averaged.
  • —EEG kept at the acquisition rate 1200 Hz (no downsample stated by the paper).

Load

bash
pip install datasets librosa soundfile   # librosa/soundfile decode the audio column
python
from datasets import load_dataset
ds = load_dataset("tankalapavankalyan/ds007808-sub01-speechopen-pangolin-preprocessed")
x = ds["train"][0]
eeg = x["eeg"]            # [128, 6000] float16
w2v = x["w2v"]            # [249, 1024] float16  (CLIP target)
audio = x["audio"]["array"]

For a CLIP retrieval model (paper): encode eeg (e.g. HTNet+Conformer) → cosine-align to w2v, zero-shot top-1/top-10 over 512 candidates.

License & citation

CC0-1.0 (derived from ds007808, CC0). Please cite ds007808 (arXiv:2606.01264) and Sato et al. 2024 (arXiv:2407.07595). Preprocessing code: this dataset's companion repo (HTNet+Conformer CLIP pipeline).