datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speech-accent-american-eng-ipaIPA_Exam_AM
独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午前)
概要
本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午前問題及び公式解答をセットにした、非公式のデータセットです。
以下、公開されている過去問題pdfから抽出しています。
https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html
人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。
データの詳細
現状、過去5年分(2020〜2024)の以下試験区分について、問題及び回答を収録しています。
応用情報技術者試験(ap)
高度共通午前1(koudo)
エンベデッドシステムスペシャリスト(es)
情報処理安全確保支援士(sc)
プロジェクトマネージャ(pm)
データベーススペシャリスト(db)
システム監査技術者(au)
ITストラテジスト(st)… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_AM.ipa-lexicon-4v0-7M
IPA Phonetic Lexicon (6.9M words)
This is an International Phonetic Language lexicon containing 6.9 million utterances across 350+ languages.
All speech files were converted into IPA using our neurlang/ipa-whisper-medium model.
Postprocessing to join multiword IPA into single-word IPA record for single-word headwords was performed for appropriate languages.
Global Map
Stats
Total words: 6908075
Countries: 211
Languages: 356
Unique Speakers: 81756… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/ipa-lexicon-4v0-7M.wiktionary-ipaPronunciation information pulled from wiktionary.org.
ipa-transcription-datase
🗣️ English Text → IPA Transcription Dataset
Overview
This dataset provides a large-scale, phonemically rich collection of English text paired with International Phonetic Alphabet (IPA) transcriptions, designed to support research and applications in speech-language pathology, phonetics, and natural language processing.
It was created to enable data-driven phonetic transcription, reducing reliance on traditional rule-based systems and supporting modern… See the full description on the dataset page: https://huggingface.co/datasets/dsvv-cair/ipa-transcription-datase.icelandic-parallel-abstracts-corpus-IPACSee https://arxiv.org/abs/2108.05289
infopedia-pt-ipa
European Portuguese IPA Lexicon — Infopédia
A lightweight word → IPA pronunciation lexicon for European Portuguese,
extracted from Infopédia (Porto Editora). One row
per headword, intended for grapheme-to-phoneme (G2P) work, pronunciation
modelling, and TTS/ASR lexicon building.
Complete crawl. Derived from a graph crawl of Infopédia that ran to
convergence (frontier → 0), covering the dictionary's reachable component.
Contents
Field
Count
Entries… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/infopedia-pt-ipa.IPA_Exam_PM2_essay
独立行政法人 情報処理推進機構(IPA) 情報処理技術者試験 試験問題データセット(午後2論文)
概要
本データセットは、独立行政法人 情報処理推進機構(以下、IPA)の情報処理技術者試験の午後2論文問題、出題趣旨、採点講評をセットにした、非公式のデータセットです。
以下、公開されている過去問題pdfから抽出しています。
https://www.ipa.go.jp/shiken/mondai-kaiotu/index.html
人間の学習目的での利用の他、LLMのベンチマークやファインチューニング等、生成AIの研究開発用途での利用を想定しています。
データの詳細
現状、過去5年分(2020〜2024)の以下試験区分について、問題、出題趣旨、採点講評を収録しています。
システムアーキテクト(sa)
プロジェクトマネージャ(pm)
ITストラテジスト(st)
ITサービスマネージャ(sm)
システム監査技術者(au)
エンベデッドシステムスペシャリスト(es)
注意事項… See the full description on the dataset page: https://huggingface.co/datasets/Dimeiza/IPA_Exam_PM2_essay.ipa-phoneme-to-words-v9ipa-phoneme-to-word
