dell
Datasets
All datasets matching “dell”AmericanStoriesAmerican Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.image-checksumsnewswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.headlines-semantic-similarity
Dataset Card for HEADLINES
Dataset Summary
HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989.
Languages
The text in the dataset is in English.
Dataset Structure
Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files.
The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.AmericanStoriesTrainingPANDA-PLUS-Bench
PANDA-PLUS-Bench
A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models.
Dataset Description
PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations.
Dataset Summary
Patches: ~2,770 per augmentation condition
Resolution: 224×224 pixels at 20× magnification
Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3)
Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.

