datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rrflux-gameechoxflow
EchoXFlow - 2-D B-mode Ultrasound Dataset
This dataset is a zea-format (HDF5) conversion of the EchoXFlow
2d_brightness_mode recordings, hosted at
zeahub/echoxflow.
Conversion
This dataset was converted to zea format and uploaded using the
zea data converter:
python -m zea.data.convert echoxflow <src> <dst>
Dataset structure
<exam_id>/
<recording_id>.hdf5
...
Each HDF5 file follows the zea data format.
EchoXFlow
EchoXFlow
This dataset repository contains Croissant metadata plus one uncompressed tar archive per exam.
Extraction
Clone or download the dataset repository first. With Git, this creates an EchoXFlow/ folder:
git lfs install
git clone https://huggingface.co/datasets/Ahus-AIM/EchoXFlow
cd EchoXFlow
The downloaded repository contains croissant.json plus one tar archive per exam under exams/. Extract every exam
archive into a local data/ directory to materialize the… See the full description on the dataset page: https://huggingface.co/datasets/Ahus-AIM/EchoXFlow.EchoXFlow-Demo
EchoXFlow Demo
This repository contains a small demo export of EchoXFlow data with one examination only.
Full dataset: Ahus-AIM/EchoXFlow.
Contents
croissant.json: MLCommons Croissant metadata for the exported collection.
exams/exam_2158f4615166f365/: the single included examination.
exams/exam_2158f4615166f365/recording_*.zarr/: Zarr-backed recordings extracted from the examination.
The demo contains:
1 examination
79 extracted recording Zarr groups
Data… See the full description on the dataset page: https://huggingface.co/datasets/Ahus-AIM/EchoXFlow-Demo.EchoX-Dialogues-Plus
EchoX-Dialogues-Plus: Training Data Plus for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues (base)
EchoX-Dialogues-Plus
EchoX-Dialogues-Plus extends KurtDu/EchoX-Dialogues with large-scale Speech-to-Speech (S2S) and Speech-to-Text (S2T) dialogues.
All assistant/output speech is synthetic (single, consistent timbre for S2S). Texts are from… See the full description on the dataset page: https://huggingface.co/datasets/KurtDu/EchoX-Dialogues-Plus.EchoX-Dialougues
EchoX-Dialogues: Training Data for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues-Plus
EchoX-Dialogues provides the primary speech dialogue data used to train EchoX, restricted to S2T (speech → text) in this repository.
All input speech is synthetic; text is derived from public sources with multi-stage cleaning and rewriting. Most turns include asr /… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues.EchoX-Data
EchoX Dataset Structure
This directory contains the unified training datasets for the EchoX Pronunciation Assessment (PA) models. The structure has been carefully designed for clean modularity, preventing data duplication, and ensuring separation between alignment strategies (MFA vs. Neural) and modeling modes (PA vs. Blind).
High-Level Directory Tree
data/
├── L2_and_Sarvah/ # LLM-labeled L2 English speaker corpus
├── LibriSpeechAnchors/ # Native… See the full description on the dataset page: https://huggingface.co/datasets/agkavin/EchoX-Data.EchoXFlow
EchoXFlow
This dataset repository contains Croissant metadata plus one uncompressed tar archive per exam.
Extraction
Clone or download the dataset repository first. With Git, this creates an EchoXFlow/ folder:
git lfs install
git clone https://huggingface.co/datasets/Ahus-AIM/EchoXFlow
cd EchoXFlow
The downloaded repository contains croissant.json plus one tar archive per exam under exams/. Extract every exam
archive into a local data/ directory to materialize… See the full description on the dataset page: https://huggingface.co/datasets/asn20/EchoXFlow.EchoXFlow
EchoXFlow
This dataset repository contains Croissant metadata plus one uncompressed tar archive per exam.
Extraction
Clone or download the dataset repository first. With Git, this creates an EchoXFlow/ folder:
git lfs install
git clone https://huggingface.co/datasets/Ahus-AIM/EchoXFlow
cd EchoXFlow
The downloaded repository contains croissant.json plus one tar archive per exam under exams/. Extract every exam
archive into a local data/ directory to materialize… See the full description on the dataset page: https://huggingface.co/datasets/Shemuel0910/EchoXFlow.EchoXfluxrec-ios-ipa
FluxRec iOS IPA — EXPERIMENTAL / NOT WORKING
⚠️ WARNING: This is NOT a working build
This IPA is an experimental proof-of-technique build. It will NOT connect to
Flux Rec servers. Do NOT expect it to work.
Why it doesn't work
The 3 Photon App IDs in this build are placeholder fake GUIDs, not real ones:
Realtime: 5dc022fb-5a0e-4e01-ab84-31f39d2b29e7 (FAKE)
Voice: 1619182a-c5b2-4bae-8ce0-cad44ccb30c7 (FAKE)
Chat: 431763b0-9e13-4cb5-82d4-88fd1efa8613… See the full description on the dataset page: https://huggingface.co/datasets/Echoxr/fluxrec-ios-ipa.
