datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multiphysics_Bench
Multiphysics Bench
Dataset: huggingface.co/datasets/Indulge-Bai/Multiphysics_Bench
Paper: Multiphysics Bench: Benchmarking and Investigating Scientific Machine Learning for Multiphysics PDEs
We propose the first general multiphysics benchmark dataset that encompasses six canonical coupled scenarios across domains such as electromagnetics, heat transfer, fluid flow, solid mechanics, pressure acoustics, and mass transport. This benchmark features the most comprehensive coupling types… See the full description on the dataset page: https://huggingface.co/datasets/Indulge-Bai/Multiphysics_Bench.moca-visrag-ind-training
Chrisyichuan/moca-visrag-ind-training
MOCA VisRAG independent-split contrastive training data with hard negatives.
Contents
moca_visrag_ind_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 122752… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-ind-training.musique_index_3hopNayanaDocs-Indic-45k-webdataset
Nayana-DocOCR Indic Annotated Dataset
Dataset Description
This is a large-scale multilingual document OCR dataset containing approximately 400GB of images with comprehensive annotations across multiple languages including Indic languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing.
Available Language Subsets
bn (Bengali): Available
en (English): Available
gu (Gujarati): Available
hi (Hindi):… See the full description on the dataset page: https://huggingface.co/datasets/Nayana-cognitivelab/NayanaDocs-Indic-45k-webdataset.tokopedia-search-indices
