NEL
Datasets
All datasets matching “NEL”standardebooks
Standard Ebooks Text Dataset
This dataset contains the full text of public domain books sourced from Standard Ebooks. It is intended for use in Natural Language Processing tasks, particularly Large Language Model pretraining, fine-tuning, and research.
Standard Ebooks provides high-quality, carefully formatted, and proofread versions of classic literature, making this a valuable collection of clean text data.
Dataset Structure
The dataset consists of a single split:… See the full description on the dataset page: https://huggingface.co/datasets/Nelathan/standardebooks.global-70m-slr
Global 70m Sea Level Rise Vulnerability Layer
A planet-wide vector layer of all land at or below 70 meters elevation — the
upper-bound geographic envelope for maximum-possible long-term sea level rise
under multi-millennial-scale ice loss.
Distributed as a single PMTiles archive suitable for direct use in MapLibre
GL, deck.gl, Leaflet, QGIS, and other modern geospatial tools that support
HTTP range requests.
File
File
Description
world-70m.pmtiles
The full… See the full description on the dataset page: https://huggingface.co/datasets/nelsn/global-70m-slr.now-city-pop-clippeft_merging_datanellThis dataset provides version 1115 of the belief
extracted by CMU's Never Ending Language Learner (NELL) and version
1110 of the candidate belief extracted by NELL. See
http://rtw.ml.cmu.edu/rtw/overview. NELL is an open information
extraction system that attempts to read the Clueweb09 of 500 million
web pages (http://boston.lti.cs.cmu.edu/Data/clueweb09/) and general
web searches.
The dataset has 4 configurations: nell_belief, nell_candidate,
nell_belief_sentences, and nell_candidate_sentences. nell_belief is
certainties of belief are lower. The two sentences config extracts the
CPL sentence patterns filled with the applicable 'best' literal string
for the entities filled into the sentence patterns. And also provides
sentences found using web searches containing the entities and
relationships.
There are roughly 21M entries for nell_belief_sentences, and 100M
sentences for nell_candidate_sentences.nel-mgenre-trie
mGENRE title trie + Wikidata QID lookup — impresso NEL assets
Three marisa-trie native binaries that
support multilingual entity linking with mGENRE. They are the runtime assets
for the impresso-project/nel-mgenre-multilingual
model as run by the impresso-inference
harness:
a title prefix tree that constrains beam search to valid Wikipedia titles,
a (language, title) → Wikidata QID lookup for offline QID resolution, and
a QID → (class, birthdate) attribute table for offline… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/nel-mgenre-trie.

