datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laya-bio
Laya-Bio: short-sequence candidate-scoring benchmark and reproducibility data
This repository packages the data and saved results used by Laya-Bio: Candidate Scoring and Reliability on Short Biological Sequences (Liang Wang, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology). The main study uses two closed-set tasks, four model conditions and three training seeds, with no additional neural continual pretraining.
Companion… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/laya-bio.research_model_dataset_v1turkish-mmlu-laya
Turkish MMLU in Laya fine-tuning format
The multiple-choice questions of alibayram/yapay_zeka_turkce_mmlu_model_cevaplari turned into
Laya choice decisions, in the same row schema as
LocalLLaMA/typed-decisions: id, workflow (the subject, bolum), and JSON strings state,
questions and gold.
Labels are the option letters A, B, C, D, E; each letter's criterion is the option text.
The state repeats the question and every option in full, since Laya truncates long options.
Gold is… See the full description on the dataset page: https://huggingface.co/datasets/AhmetSemih/turkish-mmlu-laya.danish-dynaword-laya
Danish Dynaword Laya
Danish training data for Laya in the same format as LocalLLaMA/typed-decisions, so it plugs directly into Laya's official fine-tuning notebook.
It is derived from syvai/danish-dynaword-extractions: LLM-designed JSON schemas and extractions over Danish Dynaword texts, converted into typed questions with answers.
Configs and size
config
state window
train cases / questions
test cases / questions
default
up to 16,098 tokens (max_len… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-dynaword-laya.humanoid-layangan-data
