datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prebuilt-indexes-m-beirMultiphysics_Bench
Multiphysics Bench
Dataset: huggingface.co/datasets/Indulge-Bai/Multiphysics_Bench
Paper: Multiphysics Bench: Benchmarking and Investigating Scientific Machine Learning for Multiphysics PDEs
We propose the first general multiphysics benchmark dataset that encompasses six canonical coupled scenarios across domains such as electromagnetics, heat transfer, fluid flow, solid mechanics, pressure acoustics, and mass transport. This benchmark features the most comprehensive coupling types… See the full description on the dataset page: https://huggingface.co/datasets/Indulge-Bai/Multiphysics_Bench.cityscapes_wds_indicesINDICA
INDICA: An Audio Indic-Language Telecom Fraud Analysis Benchmark
Multilingual | Audio + Text | Benchmark for Fraud Detection
Overview
INDICA is a comprehensive benchmark for telecom fraud call analysis in Indic languages.It is built on the IndiF dataset, the first large-scale multilingual dataset for fraud detection in telecom conversations.
This benchmark enables research in:
Scenario Classification
Fraud Call Detection
Fraud-Type Classification… See the full description on the dataset page: https://huggingface.co/datasets/vikrant-vikram/INDICA.indic-superb-wdsextracted-activationsmoca-visrag-ind-training
Chrisyichuan/moca-visrag-ind-training
MOCA VisRAG independent-split contrastive training data with hard negatives.
Contents
moca_visrag_ind_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 122752… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-ind-training.hoptoqa_musique_html_index
hoptoqa_musique_html_index
이 dataset repo는 train_html_hotpotqa_index와 train_html_musique_index를 함께 보관합니다.
포함 항목:
data/train_html_hotpotqa_index/
data/train_html_musique_index/
artifacts/hotpot_distractor_symlink_manifest.json
scripts/restore_musique_hotpot_index_symlinks.py
scripts/bundle_musique_hotpot_index_package.py
복구 정책:
train_html_musique_index에 실제 .npy 파일이 이미 있으면 그대로 유지
비어 있는 페이지나 잘못된 symlink만 manifest 기준으로 train_html_hotpotqa_index를 가리키는 symlink로 복원
예시:
python3… See the full description on the dataset page: https://huggingface.co/datasets/SangMin9806/hoptoqa_musique_html_index.LocAgent_index_datamusique_index_3hopmusiuqe_index2NayanaDocs-Indic-45k-webdataset
Nayana-DocOCR Indic Annotated Dataset
Dataset Description
This is a large-scale multilingual document OCR dataset containing approximately 400GB of images with comprehensive annotations across multiple languages including Indic languages and English. The dataset is stored in WebDataset format using TAR archives for efficient streaming and processing.
Available Language Subsets
bn (Bengali): Available
en (English): Available
gu (Gujarati): Available
hi (Hindi):… See the full description on the dataset page: https://huggingface.co/datasets/Nayana-cognitivelab/NayanaDocs-Indic-45k-webdataset.pdf-indirilmeyenindex_2wikimultihopqa_trainverified_questions_ind_2308_to_2896konash-indexeshotpotqa_musique_indextokopedia-search-indices
