datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MF2DocBlocks
Dataset Card for DocBlocks
DocBlocks is a high-quality, multilingual document-level machine translation (MT) dataset designed to fine-tune large language models (LLMs) on long-context translation tasks. Unlike traditional sentence-level datasets, it contains full documents with natural discourse structures and contextual alignment, helping models maintain coherence, consistency, and high translation quality across longer texts.
Curated by: Instituto Superior Técnico, Instituto de… See the full description on the dataset page: https://huggingface.co/datasets/sardinelab/DocBlocks.sardi-data
SARDI — Evaluation Data
Test splits and prebuilt BM25 indices for Self-Augmenting Retrieval for
Diffusion Language Models (ICML 2026).
Paper · Code · Model
Download
hf download pauljngr/sardi-data --repo-type dataset --local-dir data
Contents
dataset
questions
passages
size
2WikiMultiHopQA
6,253
406,822
308 MB
HotpotQA
3,701
5,239,002
2.7 GB
MuSiQue
2,417
103,035
92 MB
CofCA
900
3,156
6 MB
SynthWorlds-SM
1,200
8,055
16 MB… See the full description on the dataset page: https://huggingface.co/datasets/pauljngr/sardi-data.SARD-Vulnerability-DatasetMT-prefSardiusSardiStance
SardiStance
Disclaimer: This dataset is not the official SardiStance repository from EVALITA. For the official dataset and more information, please visit the EVALITA SardiStance page and the SardiStance repository
SardiStance is a unique dataset designed for the task of stance detection in Italian tweets. It consists of tweets related to the Sardines movement, providing a valuable resource for researchers and practitioners in the field of NLP. This dataset was curated for the… See the full description on the dataset page: https://huggingface.co/datasets/MattiaSangermano/SardiStance.bilingual_25k_filesSARDet_REC6-FSphysiology-mcqa-8kThis dataset is a subset of MedMCQA
SARDE_opendata
SARDE (Système d'Aide à la Recherche Documentaire Elaborée)
SARDE is a repository designed to provide a thematic search mode for the majority of legislative and regulatory texts in force.
The texts referenced are those published in the "Laws and Decrees" edition of the Journal officiel and in the Bulletins officiels distributed by the DILA.
flores-for-Towertico19-for-Towerwmt23-for-TowerMT-pref-humanmt-align-study-w-idiom-1203SAR-DRG
Dataset Card for SAR-DRG
Dataset Details
SAR-DRG is a scaffold-pocket dataset for realistic R-chain generation in lead optimization. It provides affinity-labeled R-chain samples organized by shared scaffold-pocket contexts, supporting SAR-informed molecular generation and evaluation. The dataset contains 96,158 samples across 29,387 scaffold-pocket groups.
Dataset Architecture
data_parquet/: Contains the processed Parquet files organized according to the… See the full description on the dataset page: https://huggingface.co/datasets/ChaseCheng/SAR-DRG.ztl-sard-php-verdicts
Source and recipe: github.com/inventor1975/introspect — dataset/sard (commit 4e179ba). The corpus is synthetic (NIST/Stivalet generated test cases); no real project's code or vulnerability is in this dataset.
ZTL verdicts over the SARD / Stivalet PHP vulnerability suite
This dataset records what the introspect analyzer (the ZTL zero-trust judge over
a deterministic PHP atomizer) returns on the public NIST SARD / Stivalet PHP test
suite — one row per test file, the verdict and… See the full description on the dataset page: https://huggingface.co/datasets/Inventor1975/ztl-sard-php-verdicts.SARDet3-FSfinetuning_demogemma3_12b-it_datasetThis dataset was specifically curated and refined for use with Gemma 3, with the goal of enhancing its instruction-following capabilities through LoRA (Low-Rank Adaptation) and refine-tuning techniques. While the original Alpaca dataset contained various issues such as hallucinated outputs, incorrect answers, inconsistent formats, and ambiguous instructions, this version has been cleaned and improved to ensure higher quality, more reliable training data.
By integrating this updated dataset… See the full description on the dataset page: https://huggingface.co/datasets/sardor233/gemma3_12b-it_dataset.SARDet_REC6_NORM-FSPERFECT_SARDINIAN_SFT_DATASET
