jstack32/LatinAccents
Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/jstack32/LatinAccents.
013
1---2language:3- en4license: apache-2.05size_categories:6 ab:7 - 10K<n<100K8 ar:9 - 100K<n<1M10 as:11 - 1K<n<10K12 ast:13 - n<1K14 az:15 - n<1K16 ba:17 - 100K<n<1M18 bas:19 - 1K<n<10K20 be:21 - 100K<n<1M22 bg:23 - 1K<n<10K24 bn:25 - 100K<n<1M26 br:27 - 10K<n<100K28 ca:29 - 1M<n<10M30 ckb:31 - 100K<n<1M32 cnh:33 - 1K<n<10K34 cs:35 - 10K<n<100K36 cv:37 - 10K<n<100K38 cy:39 - 100K<n<1M40 da:41 - 1K<n<10K42 de:43 - 100K<n<1M44 dv:45 - 10K<n<100K46 el:47 - 10K<n<100K48 en:49 - 1M<n<10M50 eo:51 - 1M<n<10M52 es:53 - 1M<n<10M54 et:55 - 10K<n<100K56 eu:57 - 100K<n<1M58 fa:59 - 100K<n<1M60 fi:61 - 10K<n<100K62 fr:63 - 100K<n<1M64 fy-NL:65 - 10K<n<100K66 ga-IE:67 - 1K<n<10K68 gl:69 - 10K<n<100K70 gn:71 - 1K<n<10K72 ha:73 - 1K<n<10K74 hi:75 - 10K<n<100K76 hsb:77 - 1K<n<10K78 hu:79 - 10K<n<100K80 hy-AM:81 - 1K<n<10K82 ia:83 - 10K<n<100K84 id:85 - 10K<n<100K86 ig:87 - 1K<n<10K88 it:89 - 100K<n<1M90 ja:91 - 10K<n<100K92 ka:93 - 10K<n<100K94 kab:95 - 100K<n<1M96 kk:97 - 1K<n<10K98 kmr:99 - 10K<n<100K100 ky:101 - 10K<n<100K102 lg:103 - 100K<n<1M104 lt:105 - 10K<n<100K106 lv:107 - 1K<n<10K108 mdf:109 - n<1K110 mhr:111 - 100K<n<1M112 mk:113 - n<1K114 ml:115 - 1K<n<10K116 mn:117 - 10K<n<100K118 mr:119 - 10K<n<100K120 mrj:121 - 10K<n<100K122 mt:123 - 10K<n<100K124 myv:125 - 1K<n<10K126 nan-tw:127 - 10K<n<100K128 ne-NP:129 - n<1K130 nl:131 - 10K<n<100K132 nn-NO:133 - n<1K134 or:135 - 1K<n<10K136 pa-IN:137 - 1K<n<10K138 pl:139 - 100K<n<1M140 pt:141 - 100K<n<1M142 rm-sursilv:143 - 1K<n<10K144 rm-vallader:145 - 1K<n<10K146 ro:147 - 10K<n<100K148 ru:149 - 100K<n<1M150 rw:151 - 1M<n<10M152 sah:153 - 1K<n<10K154 sat:155 - n<1K156 sc:157 - 1K<n<10K158 sk:159 - 10K<n<100K160 skr:161 - 1K<n<10K162 sl:163 - 10K<n<100K164 sr:165 - 1K<n<10K166 sv-SE:167 - 10K<n<100K168 sw:169 - 100K<n<1M170 ta:171 - 100K<n<1M172 th:173 - 100K<n<1M174 ti:175 - n<1K176 tig:177 - n<1K178 tok:179 - 1K<n<10K180 tr:181 - 10K<n<100K182 tt:183 - 10K<n<100K184 tw:185 - n<1K186 ug:187 - 10K<n<100K188 uk:189 - 10K<n<100K190 ur:191 - 100K<n<1M192 uz:193 - 100K<n<1M194 vi:195 - 10K<n<100K196 vot:197 - n<1K198 yue:199 - 10K<n<100K200 zh-CN:201 - 100K<n<1M202 zh-HK:203 - 100K<n<1M204 zh-TW:205 - 100K<n<1M206source_datasets:207- extended|common_voice208task_categories:209- automatic-speech-recognition210dataset_info:211 features:212 - name: path213 dtype: string214 - name: audio215 dtype: int64216 - name: sentence217 dtype: string218 splits:219 - name: train220 num_bytes: 102221 num_examples: 2222 download_size: 0223 dataset_size: 102224configs:225- config_name: default226 data_files:227 - split: train228 path: data/train-*229---230# Dataset Card for Dataset Name231 232<!-- Provide a quick summary of the dataset. -->233 234This dataset card aims to be a base template for new datasets. It has been generated using [this raw template](https://github.com/huggingface/huggingface_hub/blob/main/src/huggingface_hub/templates/datasetcard_template.md?plain=1).235 236## Dataset Details237 238### Dataset Description239 240<!-- Provide a longer summary of what this dataset is. -->241 242 243 244- **Curated by:** [More Information Needed]245- **Funded by [optional]:** [More Information Needed]246- **Shared by [optional]:** [More Information Needed]247- **Language(s) (NLP):** [More Information Needed]248- **License:** [More Information Needed]249 250### Dataset Sources [optional]251 252<!-- Provide the basic links for the dataset. -->253 254- **Repository:** [More Information Needed]255- **Paper [optional]:** [More Information Needed]256- **Demo [optional]:** [More Information Needed]257 258## Uses259 260<!-- Address questions around how the dataset is intended to be used. -->261 262### Direct Use263 264<!-- This section describes suitable use cases for the dataset. -->265 266[More Information Needed]267 268### Out-of-Scope Use269 270<!-- This section addresses misuse, malicious use, and uses that the dataset will not work well for. -->271 272[More Information Needed]273 274## Dataset Structure275 276<!-- This section provides a description of the dataset fields, and additional information about the dataset structure such as criteria used to create the splits, relationships between data points, etc. -->277 278[More Information Needed]279 280## Dataset Creation281 282### Curation Rationale283 284<!-- Motivation for the creation of this dataset. -->285 286[More Information Needed]287 288### Source Data289 290<!-- This section describes the source data (e.g. news text and headlines, social media posts, translated sentences, ...). -->291 292#### Data Collection and Processing293 294<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->295 296[More Information Needed]297 298#### Who are the source data producers?299 300<!-- This section describes the people or systems who originally created the data. It should also include self-reported demographic or identity information for the source data creators if this information is available. -->301 302[More Information Needed]303 304### Annotations [optional]305 306<!-- If the dataset contains annotations which are not part of the initial data collection, use this section to describe them. -->307 308#### Annotation process309 310<!-- This section describes the annotation process such as annotation tools used in the process, the amount of data annotated, annotation guidelines provided to the annotators, interannotator statistics, annotation validation, etc. -->311 312[More Information Needed]313 314#### Who are the annotators?315 316<!-- This section describes the people or systems who created the annotations. -->317 318[More Information Needed]319 320#### Personal and Sensitive Information321 322<!-- State whether the dataset contains data that might be considered personal, sensitive, or private (e.g., data that reveals addresses, uniquely identifiable names or aliases, racial or ethnic origins, sexual orientations, religious beliefs, political opinions, financial or health data, etc.). If efforts were made to anonymize the data, describe the anonymization process. -->323 324[More Information Needed]325 326## Bias, Risks, and Limitations327 328<!-- This section is meant to convey both technical and sociotechnical limitations. -->329 330[More Information Needed]331 332### Recommendations333 334<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->335 336Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.337 338## Citation [optional]339 340<!-- If there is a paper or blog post introducing the dataset, the APA and Bibtex information for that should go in this section. -->341 342**BibTeX:**343 344[More Information Needed]345 346**APA:**347 348[More Information Needed]349 350## Glossary [optional]351 352<!-- If relevant, include terms and calculations in this section that can help readers understand the dataset or dataset card. -->353 354[More Information Needed]355 356## More Information [optional]357 358[More Information Needed]359 360## Dataset Card Authors [optional]361 362[More Information Needed]363 364## Dataset Card Contact365 366[More Information Needed]