CoolFace
21 results

dell

dell-research-harvard /AmericanStoriesAmerican Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.text-classification100M<n<1B176 likes9.6k downloads1y agoHugging Facehf-dell-internal /image-checksums0 likes4.7k downloads7d agoHugging Facedell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4k downloads1y agoHugging Facedell-research-harvard /headlines-semantic-similarity Dataset Card for HEADLINES Dataset Summary HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989. Languages The text in the dataset is in English. Dataset Structure Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files. The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.textsentence-similarity10M<n<100M12 likes545 downloads2y agoHugging Facedell-research-harvard /AmericanStoriesTraining0 likes120 downloads3y agoHugging Facedellacorte /PANDA-PLUS-Bench PANDA-PLUS-Bench A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models. Dataset Description PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations. Dataset Summary Patches: ~2,770 per augmentation condition Resolution: 224×224 pixels at 20× magnification Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3) Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.imageimage-classification10K<n<100K1 likes97 downloads9mo agoHugging Face