datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vila-q-data-trainMutualFundclassifierdataset
Mutual Funds Query Dataset
Overview
The Mutual Funds Query Dataset is a meticulously curated collection of 2,326 conversational entries centered on mutual funds. Designed for the development and fine-tuning of financial advisory conversational agents, this dataset captures a broad spectrum of user inquiries and corresponding chatbot responses related to mutual funds investments. It serves as a valuable resource for researchers and practitioners aiming to enhance natural… See the full description on the dataset page: https://huggingface.co/datasets/VilaVision/MutualFundclassifierdataset.vilaquad
Dataset Card for VilaQuAD
Dataset Summary
VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text.
This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context).
VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence.
This dataset can be used to build extractive-QA and Language Models.
Supported Tasks and Leaderboards
Extractive-QA… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilaquad.vilasum
Dataset Card for VilaSum
Dataset Summary
VilaSum is a summarization dataset for evaluation. It is extracted from a newswire corpus crawled from the Catalan news portal VilaWeb. The corpus consists of 13,843 instances that are composed by the headline and the body.
Supported Tasks and Leaderboards
The dataset can be used to train a model for abstractive summarization. Success on this task is typically measured by achieving a high Rouge score. The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilasum.vila-datasetweibo-2014
weibo-2014
Weibo posts and reposts data in 2014, with the advantage that it has not yet been contaminated by AI bots.
Not cleaned.
ViLamr-PretrainKo-LLaVA-Instruct-150K
Korean LLaVA Visual Instruct 150K Dataset Card
Dataset details
Dataset type:
LLaVA Visual Instruct 150K is a set of GPT-generated multimodal instruction-following data.
It is constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability.
Dataset date:
LLaVA Visual Instruct 150K was collected in April 2023, by prompting GPT-4-0314 API.
Paper or resources for more information:
https://llava-vl.github.io/
License:… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/Ko-LLaVA-Instruct-150K.
