datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-speech-ipa-latents.ipa-childes-split
IPA-CHILDES split
This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular,
the following changes have been implemented:
column processed_gloss dropped as it duplicates information of gloss up to punctuation
column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+)
column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package
columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-speech-ipa-latents.IPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/phonemetransformers/IPA-CHILDES.gello_dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 30,
"total_frames": 2917,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/gello_data.studytable_open_drawer_depthThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 50,
"total_frames": 22079,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/studytable_open_drawer_depth.IPA_Exam_AM
独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午前)
概要
本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午前問題及び公式解答をセットにした、非公式のデータセットです。
以下、公開されている過去問題pdfから抽出しています。
https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html
人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。
データの詳細
現状、過去5年分(2020〜2024)の以下試験区分について、問題及び回答を収録しています。
応用情報技術者試験(ap)
高度共通午前1(koudo)
エンベデッドシステムスペシャリスト(es)
情報処理安全確保支援士(sc)
プロジェクトマネージャ(pm)
データベーススペシャリスト(db)
システム監査技術者(au)
ITストラテジスト(st)… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_AM.gello_hand_simThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 8109,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/gello_hand_sim.studytable_open_drawerThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 22079,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/studytable_open_drawer.gello_data_sim_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 7767,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/gello_data_sim_cam.new_hand_camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 50,
"total_frames": 7986,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/new_hand_cam.ipa-lexicon-4v0-7M
IPA Phonetic Lexicon (6.9M words)
This is an International Phonetic Language lexicon containing 6.9 million utterances across 350+ languages.
All speech files were converted into IPA using our neurlang/ipa-whisper-medium model.
Postprocessing to join multiword IPA into single-word IPA record for single-word headwords was performed for appropriate languages.
Global Map
Stats
Total words: 6908075
Countries: 211
Languages: 356
Unique Speakers: 81756… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/ipa-lexicon-4v0-7M.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.IPA-CHILDES
IPA-CHILDES Dataset
This dataset contains utterances downloaded from CHILDES which have been pre-processed and converted to a phonemic representation. Read the paper here.
Description
Key Columns
The scripts used to create the dataset are available here. Many of the columns from CHILDES have been preserved as they are useful for experiments (e.g. number of morphemes, part-of-speech tags, etc.). The key columns added by the processing script are as follows:… See the full description on the dataset page: https://huggingface.co/datasets/superBigPigeon/IPA-CHILDES.gello_data_simThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 30,
"total_frames": 4164,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/gello_data_sim.gello_data_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": null,
"total_episodes": 15,
"total_frames": 1293,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-intelligent-mobile-manipulators/gello_data_test.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 893,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ipa-vsp/record-test.Roll-Compactor-Control-Performance
Roll Compactor Control Performance: PID Tuning & Process Stability (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact: Crestwood, IL, USA
This dataset is 100% synthetic and intended for educational use only.
It was generated from PID control theory applied to roll compaction process
dynamics — not measured on any real equipment, customer, or production batch.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/IPA-Marketing/Roll-Compactor-Control-Performance.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-english-ipa-latents.Dry-Granulation_Roll-Compaction_and_Milling
Dry Granulation: Multi-Material Roll Compaction & Milling (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Contact: Crestwood, IL, USA
This dataset is 100% synthetic and intended for educational use only.
It was generated from published physical models and real compactor specifications —
not measured on any real equipment, customer, or production batch. Do not use it for
regulatory… See the full description on the dataset page: https://huggingface.co/datasets/IPA-Marketing/Dry-Granulation_Roll-Compaction_and_Milling.allegro-reviews-ipaipai-hackathon2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 60,
"total_frames": 51663,
"total_tasks": 5,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brk5/ipai-hackathon2.ru-reviews-classification-ipaganjoor-ipa-scansion
Ganjoor Persian Classical Poetry — Meter & Phonemic Transliteration
A corpus of 124,404 classical Persian poems collected via the Ganjoor API,
enriched with two things every poem now has:
Prosodic meter (ʿarūż / vazn) — the metrical feet and a binary scansion for every poem,
including the ~21% that Ganjoor left unlabeled (reconstructed here from the Persian feet).
Phonemic transliteration — Latin and IPA for every hemistich, produced by the
Homo-GE2PE grapheme-to-phoneme model.… See the full description on the dataset page: https://huggingface.co/datasets/nafisehNik/ganjoor-ipa-scansion.glue-stsb-ipaThe Glue STSB Dataset, but transcribed into IPA with certain restrictions, strings with numbers are preserved in the original orthography.
ipai-hackathonThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5,
"total_frames": 5516,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/brk5/ipai-hackathon.sample_asr_ipahindi-brand24-mms-ipaipa-pharma-compactor-platform
IPA Pharmaceutical Roller Compactor Platform: Scale-Up & Performance (Synthetic)
Version: 1.0
Publisher: Innovative Process Applications (IPA)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
⚠️ This dataset is 100% synthetic and intended for educational use only.
Generated from IPA's published CL-series specifications and standard
compaction physics — not real production or clinical batch data.
What's in this dataset
3,000 simulated… See the full description on the dataset page: https://huggingface.co/datasets/Innovative-Process-Applications/ipa-pharma-compactor-platform.IPA_Exam_PM2_essay
独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午後2論文)
概要
本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午後2論文問題、出題趣旨、採点講評をセットにした、非公式のデータセットです。
以下、公開されている過去問題pdfから抽出しています。
https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html
人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。
データの詳細
現状、過去5年分(2020〜2024)の以下試験区分について、問題、出題趣旨、採点講評を収録しています。
システムアーキテクト(sa)
プロジェクトマネージャ(pm)
ITストラテジスト(st)
ITサービスマネージャ(sm)
システム監査技術者(au)
エンベデッドシステムスペシャリスト(es)
注意事項… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_PM2_essay.
