opensuse
Datasets
All datasets matching “opensuse”cavil-license-patterns
Data Description
These are the license patterns currently being used by Cavil, the openSUSE legal review and SBOM system.
Intended Use
This dataset is intended to be used to train machine learning models to identify Open Source licenses. It was curated by the humans of the SUSE legal review team.
License
Licensed under GPL-2.0-or-later.
cve-backport-codegen-dataset
CVE Backport Code Generation Dataset
Per-hunk code generation dataset for CVE security patch backporting, derived from openSUSE Build Service maintenance patches.
Task
Given a region of vulnerable source code and a description of the upstream CVE fix, the model outputs the fixed version of the code. A programmatic diff then produces the final patch. This plays to LLM strengths in code completion and avoids format-sensitivity issues with direct diff generation.… See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cve-backport-codegen-dataset.cavil-legal-text
Data Description
This is training data for machine learning models that can be used with Cavil,
the openSUSE legal review and SBOM system.
Cavil uses a pattern matching system to identify potential legal text in source code. This process is based around identifying
hot zones of legal keywords (snippets) and produces around 80% false positives. Historically these false positives had to be
sorted out by humans. A few years ago we've started using machine learning to automate much of… See the full description on the dataset page: https://huggingface.co/datasets/openSUSE/cavil-legal-text.
