datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task118_semeval_2019_task10_open_vocabulary_mathematical_answer_generation.word-orb-vocabulary
Word Orb Vocabulary Intelligence
Structured vocabulary intelligence for AI agents, educators, and researchers. 162,253 words with pronunciation, etymology, age-appropriate definitions, translations across 47 languages, and ethical context.
Dataset Description
Word Orb is the world's most comprehensive structured vocabulary dataset designed for AI agents and education technology. Each word entry includes:
IPA pronunciation for text-to-speech and phonetics research… See the full description on the dataset page: https://huggingface.co/datasets/lotdpbc/word-orb-vocabulary.task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task104_semeval_2019_task10_closed_vocabulary_mathematical_answer_generation.Amazigh_Researchers_Vocabulary
Mohammed Lchger Vocabulary Dataset
This dataset contains a collection of vocabulary compiled by Mohamed Lachgar, the owner of the Amazigh researchers blog which its dataset can be find here.
Dataset Details
Content: 2,000 scientific related terms and general vocabulary.
Languages: Amazigh (zgh, ber) and English.
Script: Tifinagh.
Unique Feature: It is unique especially in its inclusion of scientific vocabulary.
Acknowledgments
Thanks to Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/Amazigh_Researchers_Vocabulary.
