a2d
Datasets
All datasets matching “a2d”qwen3_5-a2d-stage1-sft
Qwen3.5 A2D Stage 1 SFT
Curated general supervised fine-tuning corpus for Qwen3.5 text-only A2D BD3LM experiments.
Files
train-*.parquet: canonical training split data (54 shard(s)).
stage1_sft_metadata.json: curation counts and source-level metadata.
metadata.json: duplicate of the stage metadata for quick inspection.
Loading
from datasets import load_dataset
dataset = load_dataset("parquet", data_files={"train": "train-*.parquet"})
Curation Details… See the full description on the dataset page: https://huggingface.co/datasets/shreyvish5678/qwen3_5-a2d-stage1-sft.A2_DatasetThis is the official repository for the training dataset of the paper: Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter. Please download the file and unzip it in the data folder.
A2D2a2d_sentencesa2d1bcf0
Dataset Card for "a2d1bcf0"
More Information needed
SpeechOcean762_for_ConPCOThis preprocess dataset is used for ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization, which is published at IEEE ICASSP 2025.
The dataset has a total size of approximately 12 GB. To facilitate faster downloads, we have compressed the files into ZIP archives. Please unzip the archives after downloading to access the data.
The Gihub repository of ConPCO is at here.
license: cc-by-4.0
