datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HRVQA-2kA 2k subset of the validation split of the HRVQA dataset ported to HF for ease-of-use in quick remote sensing VQA evaluation.
For more information and attribution please refer to the original dataset: https://hrvqa.nl/
HR-VITON
Dataset Card for "HR-VITON"
More Information needed
hrvatski-dataset-v2
Hrvatski Dataset v2
Comprehensive Croatian (hrvatski) language dataset with 455 examples, optimized for
chatbot training and article generation. Every example is a question/answer (or
instruction/output) pair written in natural Croatian.
This repo is a single-source-of-truth catalog: all format files are generated from one
canonical file (source/hrvatski_dataset_v2.jsonl), so every format is guaranteed to
contain the exact same 455 examples. If a format is ever out of sync, that… See the full description on the dataset page: https://huggingface.co/datasets/administraktor/hrvatski-dataset-v2.hrv-qa-datasethrv-qna-dataset-v3hrv2hrv-qna-pipelinebabylm-hrv
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: hrv
Script: Latn
Tier: 1M
Byte Premium Factor: 0.989673
Size (MB): 5.39
Expected Size (MB): 5.37
Number of Documents: 466
Total Tokens: 915,054
Tokenizer: separate by whitespace
Tokens Per Category
child-directed-speech: 469,078 tokens
padding-opensubtitles: 445,976 tokens… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-hrv.
