datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
certified-document-qa
Certified Document QA: span-verified, absence-aware
79,400+ rows · every claim machine-re-checkable · zero
frontier-model-derived tokens · includes filings newer than every major
training cutoff · plus a free 127K-token verified long-context task-set.
Of the 79,438 published rows, 7,921 carry an inline machine-checkable
certificate column (needle_public, needle_expansion_v120,
absence_public, multihop_public, both teasers, and the dated multihop
splits). A further 54,837 rows —… See the full description on the dataset page: https://huggingface.co/datasets/SovNodeAI/certified-document-qa.worldbank-project-documents
Dataset Card for World Bank Project Documents
Dataset Summary
This is a dataset of documents related to World Bank development projects in the period 1947-2020. The dataset includes
the documents used to propose or describe projects when they are launched, and those in the review. The documents are indexed
by the World Bank project ID, which can be used to obtain features from multiple publicly available tabular datasets.
Supported Tasks and Leaderboards
No… See the full description on the dataset page: https://huggingface.co/datasets/lukesjordan/worldbank-project-documents.sugarcrm_130_documentation
Source: Sugarcrm 13.0 Dev Documentation
The chunks in the files are diffrent splittet based on the tokenizer conained in the name of the file
cl100k_base: 400 Tokens per chunk
p50k_base: 200 Tokens per chunk
