datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SARD
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD.MF2SARD
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features
Massive… See the full description on the dataset page: https://huggingface.co/datasets/vrinda2712/SARD.SARD
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features
Massive… See the full description on the dataset page: https://huggingface.co/datasets/caoxuhao/SARD.SARD-Extended
SARD: Synthetic Arabic Recognition Dataset
Overview
SARD (Synthetic Arabic Recognition Dataset) is a large-scale, synthetically generated dataset designed for training and evaluating Optical Character Recognition (OCR) models for Arabic text. This dataset addresses the critical need for comprehensive Arabic text recognition resources by providing controlled, diverse, and scalable training data that simulates real-world book layouts.
Key Features
Massive… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/SARD-Extended.DocBlocks
Dataset Card for DocBlocks
DocBlocks is a high-quality, multilingual document-level machine translation (MT) dataset designed to fine-tune large language models (LLMs) on long-context translation tasks. Unlike traditional sentence-level datasets, it contains full documents with natural discourse structures and contextual alignment, helping models maintain coherence, consistency, and high translation quality across longer texts.
Curated by: Instituto Superior Técnico, Instituto de… See the full description on the dataset page: https://huggingface.co/datasets/sardinelab/DocBlocks.sardi-data
SARDI — Evaluation Data
Test splits and prebuilt BM25 indices for Self-Augmenting Retrieval for
Diffusion Language Models (ICML 2026).
Paper · Code · Model
Download
hf download pauljngr/sardi-data --repo-type dataset --local-dir data
Contents
dataset
questions
passages
size
2WikiMultiHopQA
6,253
406,822
308 MB
HotpotQA
3,701
5,239,002
2.7 GB
MuSiQue
2,417
103,035
92 MB
CofCA
900
3,156
6 MB
SynthWorlds-SM
1,200
8,055
16 MB… See the full description on the dataset page: https://huggingface.co/datasets/pauljngr/sardi-data.stocks-SARDAEN-1D-candlesSAR-Duty-Cycle
Dataset Card for SAR-Duty-Cycle project
This is a preliminary training and validation dataset for sub-aperture reconstruction using generative AI.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/growan/SAR-Duty-Cycle.SARD-Vulnerability-DatasetMT-prefsardine-demo-data
SARdine demo data — Pacaya-Samiria full-frame NISAR COGs
Full-frame Cloud Optimized GeoTIFFs backing the SARdine
README hero demo. SARdine streams these directly in the browser via HTTP Range
reads — the hero link renders in seconds while reading only the tiles in view.
File
Contents
pacaya_full_hh.tif
HHHH gamma0 backscatter, raw power float32 (~500 MB)
pacaya_full_hv.tif
HVHV gamma0 backscatter, raw power float32 (~495 MB)
Source granule:… See the full description on the dataset page: https://huggingface.co/datasets/nicksteiner/sardine-demo-data.Sardiusphysiology-mcqa-8kThis dataset is a subset of MedMCQA
SardiStance
SardiStance
Disclaimer: This dataset is not the official SardiStance repository from EVALITA. For the official dataset and more information, please visit the EVALITA SardiStance page and the SardiStance repository
SardiStance is a unique dataset designed for the task of stance detection in Italian tweets. It consists of tweets related to the Sardines movement, providing a valuable resource for researchers and practitioners in the field of NLP. This dataset was curated for the… See the full description on the dataset page: https://huggingface.co/datasets/MattiaSangermano/SardiStance.bilingual_25k_filesSARDet_REC6-FSSARDE_opendata
SARDE (Système d'Aide à la Recherche Documentaire Elaborée)
SARDE is a repository designed to provide a thematic search mode for the majority of legislative and regulatory texts in force.
The texts referenced are those published in the "Laws and Decrees" edition of the Journal officiel and in the Bulletins officiels distributed by the DILA.
flores-for-Towertico19-for-Towertest_sardwmt23-for-TowerSARDdictionnaire_sarde_francais_italien_anglais_allemand
[!NOTE]
Dataset origin: https://web.archive.org/web/20121116095151/http:/www.sardegnacultura.it/documenti/7_81_20080107092727.pdf
biruniy-tts-dataMT-pref-humanmt-align-study-w-idiom-1203SARDet_REC6_NORM-FSSAR-DRG
Dataset Card for SAR-DRG
Dataset Details
SAR-DRG is a scaffold-pocket dataset for realistic R-chain generation in lead optimization. It provides affinity-labeled R-chain samples organized by shared scaffold-pocket contexts, supporting SAR-informed molecular generation and evaluation. The dataset contains 96,158 samples across 29,387 scaffold-pocket groups.
Dataset Architecture
data_parquet/: Contains the processed Parquet files organized according to the… See the full description on the dataset page: https://huggingface.co/datasets/ChaseCheng/SAR-DRG.sar_data_512
