datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sastra_ID_CardCWE-Bench-Java
CWE-Bench-Java
This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.llm-sast-v1
LLM-SAST v1
A high-quality, audited training dataset for fine-tuning small-to-mid-size language models to perform static application security testing (SAST) on infrastructure-as-code and application code — replacing rule-based scanners (Checkov, Trivy, Semgrep, KICS, Bearer, …) rather than auditing their output.
Task: given a single source file, the model emits a structured list of security findings (line ranges, category, severity, reasoning, remediation). No SAST-tool input. No… See the full description on the dataset page: https://huggingface.co/datasets/aioutfitters/llm-sast-v1.aaai27-hotpotqa-fullwiki-original
HotpotQA FullWiki frozen original data
Private reproducibility snapshot for the AAAI 2027 experiments.
This repository stores the exact Hugging Face hotpotqa/hotpot_qa, fullwiki parquet files used to construct our train, calibration, and official evaluation task identities. It does not contain generated trajectories, Writer SFT/RL examples, or model-selected subsets.
Files and rows
fullwiki/train-00000-of-00002.parquet: 45,224 rows.… See the full description on the dataset page: https://huggingface.co/datasets/sastpg/aaai27-hotpotqa-fullwiki-original.indocorpus-sastra
Indonesian Literature Corpus
Description
This dataset contains a corpus in the Indonesian language taken from Korpus Indonesia, provided by the Ministry of Education and Culture of the Republic of Indonesia. The corpus specifically focuses on literature texts, including various genres such as fiction, poetry, drama, and literary criticism.
Contents
The dataset consists of texts in the Indonesian language that are categorized under the field of literature. These… See the full description on the dataset page: https://huggingface.co/datasets/DamarJati/indocorpus-sastra.sas_to_python_base_datasetSAS_to_Python_clean_v1SAS_trainingRFTT_Datasetdivyanshukunwar__SASTRI_1_9B-details
Dataset Card for Evaluation run of divyanshukunwar/SASTRI_1_9B
Dataset automatically created during the evaluation run of model divyanshukunwar/SASTRI_1_9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/divyanshukunwar__SASTRI_1_9B-details.
