CoolFace
21 results

NEL

Nelathan /standardebooks Standard Ebooks Text Dataset This dataset contains the full text of public domain books sourced from Standard Ebooks. It is intended for use in Natural Language Processing tasks, particularly Large Language Model pretraining, fine-tuning, and research. Standard Ebooks provides high-quality, carefully formatted, and proofread versions of classic literature, making this a valuable collection of clean text data. Dataset Structure The dataset consists of a single split:… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/standardebooks.text1K<n<10K5 likes1.7k downloads1y agoHugging Facenelsn /global-70m-slr Global 70m Sea Level Rise Vulnerability Layer A planet-wide vector layer of all land at or below 70 meters elevation — the upper-bound geographic envelope for maximum-possible long-term sea level rise under multi-millennial-scale ice loss. Distributed as a single PMTiles archive suitable for direct use in MapLibre GL, deck.gl, Leaflet, QGIS, and other modern geospatial tools that support HTTP range requests. File File Description world-70m.pmtiles The full… See the full description on the dataset page: https://huggingface.co/datasets/nelsn/global-70m-slr.geospatialn<1K0 likes449 downloads4mo agoHugging Facenelsn /now-city-pop-clipgeospatialn<1K0 likes303 downloads12d agoHugging Facenellopan /peft_merging_data0 likes187 downloads1y agoHugging Facertw-cmu /nellThis dataset provides version 1115 of the belief extracted by CMU's Never Ending Language Learner (NELL) and version 1110 of the candidate belief extracted by NELL. See http://rtw.ml.cmu.edu/rtw/overview. NELL is an open information extraction system that attempts to read the Clueweb09 of 500 million web pages (http://boston.lti.cs.cmu.edu/Data/clueweb09/) and general web searches. The dataset has 4 configurations: nell_belief, nell_candidate, nell_belief_sentences, and nell_candidate_sentences. nell_belief is certainties of belief are lower. The two sentences config extracts the CPL sentence patterns filled with the applicable 'best' literal string for the entities filled into the sentence patterns. And also provides sentences found using web searches containing the entities and relationships. There are roughly 21M entries for nell_belief_sentences, and 100M sentences for nell_candidate_sentences.text-retrieval100M<n<1B10 likes165 downloads3y agoHugging Faceimpresso-project /nel-mgenre-trie mGENRE title trie + Wikidata QID lookup — impresso NEL assets Three marisa-trie native binaries that support multilingual entity linking with mGENRE. They are the runtime assets for the impresso-project/nel-mgenre-multilingual model as run by the impresso-inference harness: a title prefix tree that constrains beam search to valid Wikipedia titles, a (language, title) → Wikidata QID lookup for offline QID resolution, and a QID → (class, birthdate) attribute table for offline… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/nel-mgenre-trie.text-retrieval10M<n<100M0 likes129 downloads1mo agoHugging Face