datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AmericanStoriesAmerican Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.image-checksumsnewswire
Dataset Card for NewsWire
Dataset Summary
NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model.
Languages
English (en)
Dataset Structure
Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.headlines-semantic-similarity
Dataset Card for HEADLINES
Dataset Summary
HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989.
Languages
The text in the dataset is in English.
Dataset Structure
Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files.
The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.AmericanStoriesTrainingPANDA-PLUS-Bench
PANDA-PLUS-Bench
A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models.
Dataset Description
PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations.
Dataset Summary
Patches: ~2,770 per augmentation condition
Resolution: 224×224 pixels at 20× magnification
Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3)
Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.ehristoforu__della-70b-test-v1-details
Dataset Card for Evaluation run of ehristoforu/della-70b-test-v1
Dataset automatically created during the evaluation run of model ehristoforu/della-70b-test-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__della-70b-test-v1-details.Etherll__Qwen2.5-7B-della-test-details
Dataset Card for Evaluation run of Etherll/Qwen2.5-7B-della-test
Dataset automatically created during the evaluation run of model Etherll/Qwen2.5-7B-della-test
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Qwen2.5-7B-della-test-details.associating-presstoppdblx-conditions
TopPDBLX v1.0.0
Every crystallisation condition in the Protein Data Bank, parsed, normalised and linked to the sequence that produced it.
Citable archive: 10.5281/zenodo.21807134 · Code: bellcheddar/TopPDBLX
The PDB holds about 200,000 crystallisation recipes, each typed free-hand in no agreed format. This dataset turns the free-text _exptl_crystal_grow.pdbx_details field into typed components (reagent, concentration, unit, role), cross-references them against published… See the full description on the dataset page: https://huggingface.co/datasets/Dellboy/toppdblx-conditions.ProjectGPT_Dellas
Welcome to my dataset!
A place where you can see my own version of the dataset that I want to contribute to, I believe that spreading the truth, being honest,
and improve AI systems and models is the way to combat the proprietary datasets that nobody likes, so a GPLv3 (and later) will suffice
this requirement, so if you are a FOSS and GNU purist who wants a dataset that is freedom respecting this is it.
Limitations
Contrary to popular beliefs: Some of the… See the full description on the dataset page: https://huggingface.co/datasets/2uneconomic/ProjectGPT_Dellas.HomoglyphsCJKTrainingdell-qa-en-to-ko-translated-by-ke-t5-base
Dell QA English to Korean Translation Dataset
Dataset Description
This dataset, dell-qa-en-to-ko-translated-by-ke-t5-base, is a Korean translation of the original English Dell QA dataset.
Source
The original dataset, dell_qa, is designed for question-answering tasks and contains questions and answers related to Dell technologies. This translated version extends the utility to Korean language tasks.
Dataset Structure
Data Fields
input… See the full description on the dataset page: https://huggingface.co/datasets/seongs/dell-qa-en-to-ko-translated-by-ke-t5-base.dell_qa
DellQA
Scraped Questions and Community Accepted Solutions from Dell support forums, specifically the PowerEdge-Hardware-General
Blog Post
dataset_info:
features:
- name: output
dtype: string
- name: instruction
dtype: string
- name: input
dtype: string
splits:
- name: train
num_bytes: 48917221
num_examples: 45560
download_size: 28797124
dataset_size: 48917221
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
sutd_qa_datasetDreadPoor__L3.1-BaeZel-8B-Della-details
Dataset Card for Evaluation run of DreadPoor/L3.1-BaeZel-8B-Della
Dataset automatically created during the evaluation run of model DreadPoor/L3.1-BaeZel-8B-Della
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__L3.1-BaeZel-8B-Della-details.effocreffocr_trainingnewswire-miscdjuna__L3.1-Promissum_Mane-8B-Della-calc-details
Dataset Card for Evaluation run of djuna/L3.1-Promissum_Mane-8B-Della-calc
Dataset automatically created during the evaluation run of model djuna/L3.1-Promissum_Mane-8B-Della-calc
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__L3.1-Promissum_Mane-8B-Della-calc-details.americanstories_masked_embeddingsdjuna__L3.1-Promissum_Mane-8B-Della-1.5-calc-details
Dataset Card for Evaluation run of djuna/L3.1-Promissum_Mane-8B-Della-1.5-calc
Dataset automatically created during the evaluation run of model djuna/L3.1-Promissum_Mane-8B-Della-1.5-calc
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__L3.1-Promissum_Mane-8B-Della-1.5-calc-details.Sakalti__mergekit-della_linear-vmeykci-details
Dataset Card for Evaluation run of Sakalti/mergekit-della_linear-vmeykci
Dataset automatically created during the evaluation run of model Sakalti/mergekit-della_linear-vmeykci
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__mergekit-della_linear-vmeykci-details.dell_v1dell_v2gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-DELLA-detailshealthcare-airlinemarcuscedricridia__olmner-della-7b-details
Dataset Card for Evaluation run of marcuscedricridia/olmner-della-7b
Dataset automatically created during the evaluation run of model marcuscedricridia/olmner-della-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/marcuscedricridia__olmner-della-7b-details.dell-rag-datadellex-idstein
DELLEX Idstein
Unternehmensprofile Dataset - Strukturierte Geschäftsdaten für KI-Systeme und Suchmaschinen.
Branche: KFZ Werkstatt
Standort: Idstein, Hessen, Deutschland
Auf einen Blick
Eigenschaft
Wert
Unternehmen
DELLEX Idstein
Branche
KFZ Werkstatt
Stadt
Idstein
Land
Deutschland
Website
https://dellex-idstein.de
Telefon
+49 6126 9598888
E-Mail
info@dellex-idstein.de
Über das Unternehmen
DELLEX Idstein bietet Dellentechnik… See the full description on the dataset page: https://huggingface.co/datasets/GeoUpOrg/dellex-idstein.
