datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typst_hlmtypst_shzTypst-Train
Typst-Train
[🤖Models] |
[🛠️Code] |
[📊Data] |
Dataset used to train Typst-Coder, includes:
18.6K Typst texts
2.5K Markdown texts containing Typst-related content
typst_jpmtypst-instructtypst-instructTypst-Test
Typst-Test
[🤖Models] |
[🛠️Code] |
[📊Data] |
Dataset used to evaluate Typst-Coder, includes 1000 samples.
typst-image-dataset
Typst Image Dataset
This dataset was generated with a fork of tex2typ and the hoang-quoc-trung/fusion-image-to-latex-datasets dataset, which itself is a compilation of LaTeX labels and images of equations.
The hoang-quoc-trung dataset is difficult to work with in that it has the image data stored in a large compressed RAR archive, which does not permit efficient random read access. Additionally, it appears to have a larger number of corrupted filenames inside the archive, which has… See the full description on the dataset page: https://huggingface.co/datasets/JeppeKlitgaard/typst-image-dataset.typst-instruct
Typst Instruct Dataset
A synthetic instruction-following dataset for fine-tuning LLMs to generate Typst markup code created by Jalasoft R&D
Dataset Summary
Typst Instruct is a synthetic instruction-following dataset designed for fine-tuning large language models (LLMs) to generate Typst markup code. Typst is a modern markup-based typesetting system that serves as an alternative to LaTeX, offering cleaner syntax and faster compilation.
This… See the full description on the dataset page: https://huggingface.co/datasets/jalasoft/typst-instruct.typst-instructtypst-formulasimage-typst-ZhEntypst-full-sft-logs
