datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-medical-vqa-evaluatedTurkish-LLaVA-Pretrain
🔥 TurkishLLaVA Pretrain Dataset
This repository contains the dataset used for pretraining the Turkish-LLaVA-v0.1 model. The dataset is a Turkish translation of the English dataset used in previous studies liuhaotian. The translation was performed using DeepL. The details of this dataset and its comparison with other datasets have been published in our paper (Soon..).
Pretraining Configuration
The pretraining process focused on training only the projection matrix.… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/Turkish-LLaVA-Pretrain.turkish-comprehensive-movie-series-dataset
Beyazperde Film & Series Dataset
This dataset contains a comprehensive collection of Turkish films and TV series from Beyazperde.com, including detailed information about movies, series, cast, reviews, and ratings.
Dataset Summary
Total Movies: 27,227
Total Series: 11,240
Total Entries: 38,467
File Size: ~222 MB
Format: JSONL (JSON Lines)
Language: Turkish
Source: Beyazperde.com
Data Structure
Each line in the JSONL file contains a JSON object… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/turkish-comprehensive-movie-series-dataset.
