CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Efficient-Large-Model /vila-q-data-trainimage10K<n<100K0 likes293 downloads1y agoHugging Face02VilaVision /MutualFundclassifierdataset Mutual Funds Query Dataset Overview The Mutual Funds Query Dataset is a meticulously curated collection of 2,326 conversational entries centered on mutual funds. Designed for the development and fine-tuning of financial advisory conversational agents, this dataset captures a broad spectrum of user inquiries and corresponding chatbot responses related to mutual funds investments. It serves as a valuable resource for researchers and practitioners aiming to enhance natural… See the full description on the dataset page: https://huggingface.co/datasets/VilaVision/MutualFundclassifierdataset.text1K<n<10K2 likes102 downloads2y agoHugging Face03projecte-aina /vilaquad Dataset Card for VilaQuAD Dataset Summary VilaQuAD, An extractive QA dataset for Catalan, from VilaWeb newswire text. This dataset contains 2095 of Catalan language news articles along with 1 to 5 questions referring to each fragment (or context). VilaQuad articles are extracted from the daily VilaWeb and used under CC-BY-NC-SA-ND licence. This dataset can be used to build extractive-QA and Language Models. Supported Tasks and Leaderboards Extractive-QA… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilaquad.textquestion-answering1K<n<10K0 likes84 downloads2y agoHugging Face04projecte-aina /vilasum Dataset Card for VilaSum Dataset Summary VilaSum is a summarization dataset for evaluation. It is extracted from a newswire corpus crawled from the Catalan news portal VilaWeb. The corpus consists of 13,843 instances that are composed by the headline and the body. Supported Tasks and Leaderboards The dataset can be used to train a model for abstractive summarization. Success on this task is typically measured by achieving a high Rouge score. The… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/vilasum.textsummarization10K<n<100K0 likes69 downloads2y agoHugging Face05mit-han-lab /vila-datasettext10K<n<100K5 likes51 downloads2y agoHugging Face06vilarin /weibo-2014 weibo-2014 Weibo posts and reposts data in 2014, with the advantage that it has not yet been contaminated by AI bots. Not cleaned. text1M<n<10M1 likes34 downloads2y agoHugging Face07JingjingJiang /ViLamr-Pretraintext100K<n<1M0 likes25 downloads2y agoHugging Face08etri-vilab /Ko-LLaVA-Instruct-150K Korean LLaVA Visual Instruct 150K Dataset Card Dataset details Dataset type: LLaVA Visual Instruct 150K is a set of GPT-generated multimodal instruction-following data. It is constructed for visual instruction tuning and for building large multimodal towards GPT-4 vision/language capability. Dataset date: LLaVA Visual Instruct 150K was collected in April 2023, by prompting GPT-4-0314 API. Paper or resources for more information: https://llava-vl.github.io/ License:… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/Ko-LLaVA-Instruct-150K.textvisual-question-answering100K<n<1M0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.