datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bsard
Dataset Card for BSARD
Dataset Summary
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native dataset for studying legal information retrieval. BSARD consists of more than 22,600 statutory articles from Belgian law and about 1,100 legal questions posed by Belgian citizens and labeled by experienced jurists with relevant articles from the corpus.
Supported Tasks and Leaderboards
document-retrieval: The dataset can be used to train models for… See the full description on the dataset page: https://huggingface.co/datasets/maastrichtlawtech/bsard.indian-legal-sections-bns-bnss-bsa-2023
🏛️ Indian Legal Sections — BNS · BNSS · BSA 2023
The First Structured, Unified JSON Dataset of Modern Indian Criminal Law
📖 Dataset Summary
This dataset contains 1,059 fully structured and verified sections extracted, parsed, and unified from India's three landmark criminal justice reform acts passed in December 2023. These three acts together replaced the colonial-era Indian Penal Code (IPC, 1860), the Code of Criminal Procedure… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/indian-legal-sections-bns-bnss-bsa-2023.dreamt
Dataset Description
DREAMT (Dataset for Real-time sleep stage EstimAtion using Multisensor wearable Technology) is a dataset designed to facilitate the development and evaluation of machine learning models for sleep stage estimation using data from multisensor wearable devices.
Version: 2.1.0
Repository: PhysioNet: DREAMT v2.1.0
Access Policy & Licensing
Due to the sensitive nature of health data, this dataset is restricted and cannot be downloaded directly without… See the full description on the dataset page: https://huggingface.co/datasets/bsaenz/dreamt.Indian-Legal-QA-BNS-BNSS-BSA
Indian Legal QA — BNS + BNSS + BSA 2023
6,354 structured question-answer pairs covering all 1,059 sections across India's three criminal justice acts of 2023
Overview
This dataset contains 6,354 instruction-format question-answer pairs in JSONL format, covering every section of India's three criminal justice reform acts enacted in 2023. Each section has exactly 6 questions approaching the same legal provision from different angles… See the full description on the dataset page: https://huggingface.co/datasets/GSMS-B/Indian-Legal-QA-BNS-BNSS-BSA.travel-fraud-graphs
TravelFraudBench (TFG)
The first publicly available labeled heterogeneous graph benchmark for GNN-based fraud ring detection in travel networks.
Paper
Dataset Structure
This dataset contains heterogeneous graph data split into 20 named configurations — one per node type and one per edge type — each with small, medium, and large splits.
Loading a specific node or edge table
from datasets import load_dataset
# Load user nodes (medium scale)
users =… See the full description on the dataset page: https://huggingface.co/datasets/bsajja7/travel-fraud-graphs.BSARDRetrieval
BSARDRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
The Belgian Statutory Article Retrieval Dataset (BSARD) is a French native dataset for studying legal information retrieval. BSARD consists of more than 22,600 statutory articles from Belgian law and about 1,100 legal questions posed by Belgian citizens and labeled by experienced jurists with relevant articles from the corpus.
Task category
t2t
Domains
Legal, Spoken
Reference… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BSARDRetrieval.RobustBiasBench
RobustBiasBench Dataset Description
This document provides an overview of the features and labels included in the RobustBiasBench dataset, which consists of over 18,000 policy excerpts annotated for bias type and normative framing.
📄 Dataset Format
The dataset is stored in CSV format and contains the following columns:
Column Name
Description
id
Unique numeric ID for each policy excerpt
date
Year of the policy (extracted from the source document)… See the full description on the dataset page: https://huggingface.co/datasets/bsakash/RobustBiasBench.BSARDRetrieval.v2
BSARDRetrieval.v2
An MTEB dataset
Massive Text Embedding Benchmark
BSARD is a French native dataset for legal information retrieval. BSARDRetrieval.v2 covers multi-article queries, fixing issues (#2906) with the previous data loading.
Task category
t2c
Domains
Legal, Spoken
Reference
https://huggingface.co/datasets/maastrichtlawtech/bsard
Source datasets:
maastrichtlawtech/bsard
How to evaluate on this task
You can evaluate an embedding model on this… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BSARDRetrieval.v2.multiple_choice_bsardBSA_Benchadaption-bsad-synthetic-bio-entries-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synthetic_bio_entries
This dataset contains synthetic entries for advanced synthetic biology research, specifically focusing on edge-case Dual Use Research of Concern (DURC) scenarios. Each record includes metadata such as organism type, genetic engineering methods, biosafety levels, and risk assessments in a structured CSV format. The entries are designed to represent novel and… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synthetic-bio-entries-v1.iclr27siyuan-bsab-md-results
BSab molecular dynamics results / BSab 分子动力学结果
This derived-results release contains the final summaries, per-replica/site analyses, time-block statistics and corrected conformational probability maps from 10 systems × 3 independent replicas × 200 ns = 6,000 ns (6 μs) of production molecular dynamics. Production and the default analysis matrix completed on 2026-09-21 at 19:20:42 China Standard Time (11:20:42 UTC). The corrected maps were completed on 2026-09-22 China Standard… See the full description on the dataset page: https://huggingface.co/datasets/b1intern/iclr27siyuan-bsab-md-results.adaption-bsad-identifiers
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_identifiers
This dataset consists of unique alphanumeric identifiers following the 'BSAD' prefix pattern, such as BSAD_076 and BSAD_089. Each entry is formatted as a single prompt string containing one specific code. The collection appears to serve as a reference list or index for items within a larger BSAD-categorized system.
Dataset size
There are 20 data points in this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-identifiers.adaption-bsad-biotech-risk-entries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_biotech_risk_entries
This dataset contains structured entries detailing diverse biotechnology applications across agriculture, virology, and environmental sectors with a focus on Global South relevance. Each record specifies the organism, engineering technique, associated biosafety level, and a quantified risk score ranging from low to high. The collection includes specific scenarios… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-biotech-risk-entries.adaption-bsad-bioeconomy-entries-251-270
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_bioeconomy_entries_251_270
This dataset contains synthetic Biosecurity Assessment Dataset (BSAD) entries ranging from IDs 251 to 270, focusing on energy biotech, industrial biotech, and bioeconomy systems. Each entry includes simulated study details such as organism type, genetic engineering status, risk level, and biosafety containment (BSL-1). The data is structured as CSV-ready records… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-bioeconomy-entries-251-270.adaption-bsad-industrial-biotech-entries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_industrial_biotech_entries
This dataset contains synthetic entries (IDs 591–610) for the Biosafety and Security Assessment Database (BSAD) focused on industrial biotechnology and bio-manufacturing. Each entry includes simulated study details, organism information, genetic engineering status, and biosafety levels formatted as CSV rows. The data is uniformly classified as BSL-1 with low risk… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-industrial-biotech-entries.adaption-bsad-industrial-biotech-entries-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_industrial_biotech_entries
This dataset contains synthetic entries (IDs 591–610) for the Biosafety and Security Assessment Database (BSAD) focused on industrial biotechnology and bio-manufacturing. Each entry includes simulated study details, organism information, genetic engineering status, and biosafety levels formatted as CSV rows. The data is uniformly classified as BSL-1 with low risk… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-industrial-biotech-entries-v1.adaption-bsad-synthetic-bio-entries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synthetic_bio_entries
This dataset contains synthetic entries for the BSAD (Biosecurity and Synthetic Biology Assessment Dataset), focusing on emerging biotechnology and genetic engineering research. Each record includes simulated study descriptions, organism identifiers, risk levels, biosafety containment (BSL-1), and global context metadata formatted as CSV-like fields. The entries are… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synthetic-bio-entries.adaption-bsad-synbio-genome-eng-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synbio_genome_eng
This dataset contains synthetic entries from the BSAD collection focusing on synthetic biology and genome engineering research scenarios. Each entry includes simulated study descriptions, organism details, biosafety levels, and risk assessments formatted as CSV data. The samples specifically cover IDs 231–250 with varying keywords and global scope indicators.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synbio-genome-eng-v1.adaption-bsad-agri-climate-resilience
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_agri_climate_resilience
This dataset contains synthetic entries for agricultural biotechnology studies focused on climate resilience, with a specific emphasis on the Global South. Each record includes simulated study details, organism information, genetic engineering methods, and biosafety levels formatted as CSV-like text. The data is designed to represent low-risk research scenarios under… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-agri-climate-resilience.adaption-bsad-agri-climate-resilience-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_agri_climate_resilience
This dataset contains synthetic entries for agricultural biotechnology studies focused on climate resilience, with a specific emphasis on the Global South. Each record includes simulated study details, organism information, genetic engineering methods, and biosafety levels formatted as CSV-like text. The data is designed to represent low-risk research scenarios under… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-agri-climate-resilience-v1.adaption-bsad-synbio-gene-editing
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synbio_gene_editing
This dataset contains synthetic entries (IDs 511–530) simulating records for synthetic biology and gene editing research studies. Each entry includes structured fields such as organism type, biosafety level (BSL-1), risk assessment, and global regulatory context formatted for CSV output. The data is designed to represent low-risk genetic engineering projects under WHO… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synbio-gene-editing.adr_tank_1adaption-bsad-amr-env-micro-samples
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_amr_env_micro_samples
This dataset contains synthetic entries from the BSAD collection focusing on public health, antimicrobial resistance (AMR), and environmental microbiology. Each sample represents a simulated study involving genetic engineering research with low-risk profiles (BSL-1) and global WHO relevance. The data is structured to include organism details, risk assessments, and… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-amr-env-micro-samples.adaption-bsad-climate-agri-synthetic
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_climate_agri_synthetic
This dataset contains synthetic entries for climate-smart agriculture and food systems research, specifically focusing on genetic engineering applications. Each record includes simulated study details, organism identifiers, biosafety levels (BSL-1), and global scope indicators formatted as CSV-like text. The data is artificially generated to mimic structured biological… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-climate-agri-synthetic.adaption-bsad-env-biotech-entries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_env_biotech_entries
This dataset contains synthetic entries for the BSAD (Biosafety and Biosecurity Assessment Database) focusing on environmental biotechnology and pollution control. Each sample includes structured fields such as organism type, genetic engineering status, research purpose, biosafety level (BSL-1), and global regulatory context. The data is formatted as CSV-ready records… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-env-biotech-entries.dataset-boamp-siren-acheteur-bsa-assurances-83-colonnes-2024
Dataset «BOAMP-SIREN-ACHETEUR (BSA)» / Assurances / 83 colonnes / 2024
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Dataset «BOAMP-SIREN-ACHETEUR (BSA)» / Assurances / 83 colonnes / 2024 qui est disponible à l'adresse https://www.data.gouv.fr/datasets/674ee6459ce96a9c6c4c9eee
Description
Pour une mise en pratique, rendez-vous sur AuFilDuBoamp par ici:
Boamp/Assurances - Élaborer le tableau du TOP-10/2024… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/dataset-boamp-siren-acheteur-bsa-assurances-83-colonnes-2024.adaption-bsad-synbio-genome-eng
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synbio_genome_eng
This dataset contains synthetic entries from the BSAD collection focusing on synthetic biology and genome engineering research scenarios. Each entry includes simulated study descriptions, organism details, biosafety levels, and risk assessments formatted as CSV data. The samples specifically cover IDs 231–250 with varying keywords and global scope indicators.… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synbio-genome-eng.adaption-bsad-synthetic-entries-631-650
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_synthetic_entries_631_650
This dataset contains synthetic entries (IDs 631–650) from the BSAD collection, focusing on diagnostics, biosensors, and surveillance systems. Each entry includes simulated study details such as organism type, genetic engineering status, biosafety level, and global scope. The data is formatted as CSV-ready text with consistent fields for research classification and… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-synthetic-entries-631-650.biography
