datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kenyan_swahili_nonstandard_speech_v1.0This dataset provides 32.5 hours of Swahili speech recordings (5,535 samples) from 52 Kenyan speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or… See the full description on the dataset page: https://huggingface.co/datasets/cdli/kenyan_swahili_nonstandard_speech_v1.0.rwandan_kinyarwanda_nonstandard_speech_v1.0This dataset provides 61.7 hours of Kinyarwanda speech recordings (14,739 samples) from 61 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_kinyarwanda_nonstandard_speech_v1.0.ugandan_english_nonstandard_speech_v1.0This dataset provides 42.4 hours of Ugandan English speech recordings (7,251 samples) from 59 Ugandan speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker… See the full description on the dataset page: https://huggingface.co/datasets/cdli/ugandan_english_nonstandard_speech_v1.0.rwandan_english_nonstandard_speech_v1.0This dataset provides 32.7 hours of English speech recordings (7,592 samples) from 44 Rwandan speakers living with speech impairments. The participants represent a limited diversity of speech patterns, mostly stuttering, and a few examples of Dysarthria, Dysphonia, and Phonological disorders.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or phrase level. All speech recordings of this datasets have been… See the full description on the dataset page: https://huggingface.co/datasets/cdli/rwandan_english_nonstandard_speech_v1.0.kenyan_english_nonstandard_speech_v1.0This dataset provides 32.3 hours of English speech recordings (5,998 samples) from 52 Kenyan speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or… See the full description on the dataset page: https://huggingface.co/datasets/cdli/kenyan_english_nonstandard_speech_v1.0.ghanian_ga_nonstandard_speech_v1.0This dataset provides 6.28 hours of Ga nonstandard speech recordings (12,160 samples) from 21 Ga speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker or… See the full description on the dataset page: https://huggingface.co/datasets/cdli/ghanian_ga_nonstandard_speech_v1.0.ugandan_luganda_nonstandard_speech_v1.0This dataset provides 44.4 hours of Ugandan English speech recordings (8,137 samples) from 59 Ugandan speakers living with speech impairments. The participants represent a diversity of progressive, acquired and congenital aetiologies, including cerebral palsy, Parkinson's disease, multiple sclerosis, autism spectrum disorder, Down syndrome, stroke and stuttering.
This dataset includes a split into a training, test and development set. The splits were created avoiding any overlap on the speaker… See the full description on the dataset page: https://huggingface.co/datasets/cdli/ugandan_luganda_nonstandard_speech_v1.0.NonStandardClauseAnalytics
NonStandardClauseAnalytics
tags: analytics, nonstandard, legal data
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description: The 'NonStandardClauseAnalytics' dataset compiles various non-standard contractual clauses found in legal documents, aiming to provide insights for legal practitioners on prevalent contractual deviations and their potential implications. Each row contains a clause extracted from a contract, a description, and a label… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/NonStandardClauseAnalytics.
