nld
Datasets
All datasets matching “nld”ExoW3-NLDM
ExoW3-NLDM
54 LCMS runs (mzML). Every file is self-describing: the complete
experimental record is encoded directly into the mass
spectra with SpectraCodec. Decode any single file to
recover it. How the encoding and decoding works is documented in the
SpectraCodec repository.
Authenticity
Every file is cryptographically signed. The embedded message carries a
provenance block with two Ed25519 signatures: payload_signature covers the
embedded experimental record… See the full description on the dataset page: https://huggingface.co/datasets/lbnl-metabolomics/ExoW3-NLDM.NL-DIR
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark (CVPR 2025)
Dataset Introduction
NL-DIR consists of 41,795 document images with a diverse range of document types, and each image corresponds to 5 high-quality fine-grained semantic queries, which are generated and evaluated through large language models in conjunction with manual verification, resulting 200K+ queries in total. Following an 8:1:1 ratio partition, the dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/nianbing/NL-DIR.goldfish-nld-latn-100mb-tokenizedNL-DIR-sampleThe NL-DIR dataset comprises 41K authentic document images, in which each image is paired with five high-quality fine-grained semantic queries, generated and evaluated through large language models in conjunction with manual verification.
The complete dataset, along with detailed descriptions, specific formats, usage instructions, and construction methods, will be released soon.
nld_zipfbabylm-nld
BabyLM Dataset
Dataset Description
This dataset is part of the BabyLM multilingual collection.More information at: babylm.github.io/babybabellm
Dataset Summary
Language: nld
Script: Latn
Tier: 100M
Byte Premium Factor: 1.051606
Size (MB): 569.49
Expected Size (MB): 571.02
Number of Documents: 304,611
Total Tokens: 109,885,564
Tokenizer: separate by whitespace
Tokens Per Category
child-books: 4,576,823 tokens
child-directed-speech: 3,304,756… See the full description on the dataset page: https://huggingface.co/datasets/BabyLM-community/babylm-nld.
