datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.surface-audit
TruthfulQA-476 — a surface-form-cleaned binary-choice TruthfulQA
TruthfulQA-476 is the recommended drop-in replacement for the binary-choice TruthfulQA
evaluation set. It keeps 476 of the 790 original question pairs, in the original schema, chosen so
that a classifier restricted to six surface features of the answer text (negation, hedging, length,
token statistics) can no longer separate correct from incorrect answers above chance, while the
ranking of models on the subset… See the full description on the dataset page: https://huggingface.co/datasets/iclr2027-surface-audit/surface-audit.OASST_Top1_IndonesianBase dataset : OpenAssistant/oasst1
We selected the dataset with "english" language & has rank = 1st.
Finally, we translate it to Indonesian with Marian NMT and pretrained model from Helsinki-NLP/opus-mt-en-id.
CITATION
@InProceedings{mariannmt,
title = {Marian: Fast Neural Machine Translation in {C++}},
author = {Junczys-Dowmunt, Marcin and Grundkiewicz, Roman and
Dwojak, Tomasz and Hoang, Hieu and Heafield, Kenneth and
Neckermann… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/OASST_Top1_Indonesian.deciban
deciban — a vote-level corpus for studying diversity of thought in LLM ensembles
Ten open-weight language models (20–31B parameters), each answering every question of
three benchmarks 24 times at temperature 1.0, with every individual vote retained:
767,520 multiple-choice inferences and 23,520 graded free-text forensic trials, plus
parse/termination reasons and abstentions as first-class outcomes. The corpus exposes the
joint answer distribution between models — which questions… See the full description on the dataset page: https://huggingface.co/datasets/IcyApril/deciban.BELLE-eval-S2S
BELLE-eval-S2S
💡 Dataset Description
BELLE-eval-S2S is a Chinese evaluation dataset for speech-to-speech conversational tasks. It contains 250 Chinese audio samples with corresponding text annotations and is intended for model evaluation rather than training.
🔗 Source
Original text source: the test set from LianjiaTech/BELLE
This dataset is built by filtering 250 samples from the original test set and synthesizing them into speech audio
📖 Data… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/BELLE-eval-S2S.ICEQ-Dataset
📚 Russian Question Generation Dataset (ICEQ)
🧩 Описание
Russian Question Generation Dataset — это датасет, созданный в рамках проекта ICEQ (Input, Chunks, Embeddings, Questions. Он содержит примеры генерации вопросов с вариантами ответов на основе русскоязычных текстов, направленные на проверку понимания прочитанного. Данные были сгенерированы синтетически с помощью DeepSeek.
Каждая строка датасета представляет собой:
prompt: инструкция нейросети с вложенным фрагментом… See the full description on the dataset page: https://huggingface.co/datasets/droyti/ICEQ-Dataset.
