datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bsca-binary-source-gold-v3-multidomain
BSCA Gold v3 Multidomain
Address-grounded P1 pairs for stripped pseudo-C → source retrieval.
Dataset ID: GD_19330e06aae0462447c1fd05ccaa38d7
Accepted P1 pairs: 42449
Repositories: 138
Target formats: {"elf": 40733, "pe": 1716}
Target architectures: {"aarch64": 1672, "x86": 1903, "x86_64": 38874}
Internal quality GPA: 3.660; target pass: True
Use train.jsonl for fitting, development.jsonl for model selection, and
the immutable test.jsonl only after selection. dataset_card.json… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-binary-source-gold-v3-multidomain.bsca-sca-benchmark-v1
BSCA SCA Benchmark v1
Benchmark dataset for evaluating Binary Software Composition Analysis (SCA) — identifying which open-source libraries are present in a stripped binary.
Dataset Description
10 real-world stripped ELF binaries with ground-truth component labels. Used to evaluate embedding-based SCA systems that match binary functions against open-source code.
Attribution: The binaries and ground-truth labels in this dataset are derived and refined from the artifact of… See the full description on the dataset page: https://huggingface.co/datasets/Labradorlabs/bsca-sca-benchmark-v1.labradorimnet1k_Labrador_retrieverLexique_analytique_du_vocabulaire_inuit_moderne_au_Quebec-Labrador
[!NOTE]
Dataset origin: https://archive.org/details/lexiqueanalytiqu0000dora/mode/2up
Note : il faut se connecter pour accéder aux données
spawrious224-bulldog-labrador
