nld
beetle-bilingual-balanced-b1-fineweb-nld-engbeetle-bilingual-balanced-b4-fineweb-nld-engbeetle-bilingual-balanced-b4-fineweb-eng-nldbeetle-bilingual-balanced-b1-fineweb-eng-nldnld-100mb-after-wc-uniform-newlex-nld-ckpt500_seed10_seed10beetle-bilingual-balanced-b5-fineweb-eng-nldbeetle-bilingual-l2-50-sequential-33-67-b3-fineweb-nld-engnld-100mb-after-nld_newlexicon_zipf_heavy-ckpt500_seed3407
Datasets
All datasets matching “nld”ExoW3-NLDM
ExoW3-NLDM
54 LCMS runs (mzML). Every file is self-describing: the complete
experimental record is encoded directly into the mass
spectra with SpectraCodec. Decode any single file to
recover it. How the encoding and decoding works is documented in the
SpectraCodec repository.
Authenticity
Every file is cryptographically signed. The embedded message carries a
provenance block with two Ed25519 signatures: payload_signature covers the
embedded experimental record… See the full description on the dataset page: https://huggingface.co/datasets/lbnl-metabolomics/ExoW3-NLDM.NL-DIR
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark (CVPR 2025)
Dataset Introduction
NL-DIR consists of 41,795 document images with a diverse range of document types, and each image corresponds to 5 high-quality fine-grained semantic queries, which are generated and evaluated through large language models in conjunction with manual verification, resulting 200K+ queries in total. Following an 8:1:1 ratio partition, the dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/nianbing/NL-DIR.goldfish-nld-latn-100mb-tokenizednld_zipfnld_zipf_fix_zijnnld_heavy_zipf_fix_zijn
