datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nitibench
👩🏻⚖️ NitiBench: A Thai Legal Benchmark for RAG
[📄 Technical Report] | [👨💻 Github Repository]
This dataset provides the test data for evaluating LLM frameworks, such as RAG or LCLM. The benchmark consists of two datasets:
NitiBench-CCL
NitiBench-Tax
🏛️ NitiBench-CCL
Derived from the WangchanX-Legal-ThaiCCL-RAG Dataset, our version includes an additional preprocessing step in which we separate the reasoning process from the final answer. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/VISAI-AI/nitibench.VIS-APP-Bench
anchor_tasks_web — Dataset
This dataset is part of the benchmark presented in the paper VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents.
Project Page | GitHub Repository | Paper
A web-app generation benchmark. Each task is a multi-page UI taken from a public Figma community file. For every task we ship the textual page descriptions, the rendered mockup PNGs, the Figma node structure, the per-page click-annotations, and the distilled data-testid "anchors"… See the full description on the dataset page: https://huggingface.co/datasets/JunJiaGuo/VIS-APP-Bench.
