datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Typst-Train
Typst-Train
[🤖Models] |
[🛠️Code] |
[📊Data] |
Dataset used to train Typst-Coder, includes:
18.6K Typst texts
2.5K Markdown texts containing Typst-related content
Typst-Test
Typst-Test
[🤖Models] |
[🛠️Code] |
[📊Data] |
Dataset used to evaluate Typst-Coder, includes 1000 samples.
typst-instruct
Typst Instruct Dataset
A synthetic instruction-following dataset for fine-tuning LLMs to generate Typst markup code created by Jalasoft R&D
Dataset Summary
Typst Instruct is a synthetic instruction-following dataset designed for fine-tuning large language models (LLMs) to generate Typst markup code. Typst is a modern markup-based typesetting system that serves as an alternative to LaTeX, offering cleaner syntax and faster compilation.
This… See the full description on the dataset page: https://huggingface.co/datasets/jalasoft/typst-instruct.
