CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dell-research-harvard /AmericanStoriesAmerican Stories offers high-quality structured data from historical newspapers suitable for pre-training large language models to enhance the understanding of historical English and world knowledge. It can also be integrated into external databases of retrieval-augmented language models, enabling broader access to historical information, including interpretations of political events and intricate details about people's ancestors. Additionally, the structured article texts facilitate the application of transformer-based methods for popular tasks like detecting reproduced content, significantly improving accuracy compared to traditional OCR methods. American Stories serves as a substantial and valuable dataset for advancing multimodal layout analysis models and other multimodal applications.text-classification100M<n<1B176 likes9.6k downloads1y agoHugging Face02hf-dell-internal /image-checksums0 likes4.7k downloads7d agoHugging Face03dell-research-harvard /newswire Dataset Card for NewsWire Dataset Summary NewsWire contains 2.7 million unique public domain U.S. news wire articles, written between 1878 and 1977. Locations in these articles are georeferenced, topics are tagged using customized neural topic classification, named entities are recognized, and individuals are disambiguated to Wikipedia using a novel entity disambiguation model. Languages English (en) Dataset Structure Each year in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/newswire.tabulartext-classification1M<n<10M92 likes4k downloads1y agoHugging Face04dell-research-harvard /headlines-semantic-similarity Dataset Card for HEADLINES Dataset Summary HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989. Languages The text in the dataset is in English. Dataset Structure Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files. The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.textsentence-similarity10M<n<100M12 likes545 downloads2y agoHugging Face05dell-research-harvard /AmericanStoriesTraining0 likes120 downloads3y agoHugging Face06dellacorte /PANDA-PLUS-Bench PANDA-PLUS-Bench A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models. Dataset Description PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations. Dataset Summary Patches: ~2,770 per augmentation condition Resolution: 224×224 pixels at 20× magnification Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3) Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.imageimage-classification10K<n<100K1 likes97 downloads9mo agoHugging Face07open-llm-leaderboard /ehristoforu__della-70b-test-v1-detailsgated Dataset Card for Evaluation run of ehristoforu/della-70b-test-v1 Dataset automatically created during the evaluation run of model ehristoforu/della-70b-test-v1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ehristoforu__della-70b-test-v1-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face08open-llm-leaderboard /Etherll__Qwen2.5-7B-della-test-detailsgated Dataset Card for Evaluation run of Etherll/Qwen2.5-7B-della-test Dataset automatically created during the evaluation run of model Etherll/Qwen2.5-7B-della-test The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Etherll__Qwen2.5-7B-della-test-details.tabular10K<n<100K0 likes40 downloads2y agoHugging Face09dell-research-harvard /associating-presstextn<1K0 likes38 downloads3y agoHugging Face10Dellboy /toppdblx-conditions TopPDBLX v1.0.0 Every crystallisation condition in the Protein Data Bank, parsed, normalised and linked to the sequence that produced it. Citable archive: 10.5281/zenodo.21807134 · Code: bellcheddar/TopPDBLX The PDB holds about 200,000 crystallisation recipes, each typed free-hand in no agreed format. This dataset turns the free-text _exptl_crystal_grow.pdbx_details field into typed components (reagent, concentration, unit, role), cross-references them against published… See the full description on the dataset page: https://huggingface.co/datasets/Dellboy/toppdblx-conditions.tabular100K<n<1M0 likes32 downloads2mo agoHugging Face112uneconomic /ProjectGPT_Dellas Welcome to my dataset! A place where you can see my own version of the dataset that I want to contribute to, I believe that spreading the truth, being honest, and improve AI systems and models is the way to combat the proprietary datasets that nobody likes, so a GPLv3 (and later) will suffice this requirement, so if you are a FOSS and GNU purist who wants a dataset that is freedom respecting this is it. Limitations Contrary to popular beliefs: Some of the… See the full description on the dataset page: https://huggingface.co/datasets/2uneconomic/ProjectGPT_Dellas.1 likes22 downloads9d agoHugging Face12dell-research-harvard /HomoglyphsCJKTraining0 likes21 downloads3y agoHugging Face13seongs /dell-qa-en-to-ko-translated-by-ke-t5-base Dell QA English to Korean Translation Dataset Dataset Description This dataset, dell-qa-en-to-ko-translated-by-ke-t5-base, is a Korean translation of the original English Dell QA dataset. Source The original dataset, dell_qa, is designed for question-answering tasks and contains questions and answers related to Dell technologies. This translated version extends the utility to Korean language tasks. Dataset Structure Data Fields input… See the full description on the dataset page: https://huggingface.co/datasets/seongs/dell-qa-en-to-ko-translated-by-ke-t5-base.textquestion-answering10K<n<100K1 likes18 downloads1y agoHugging Face14c123ian /dell_qa DellQA Scraped Questions and Community Accepted Solutions from Dell support forums, specifically the PowerEdge-Hardware-General Blog Post dataset_info: features: - name: output dtype: string - name: instruction dtype: string - name: input dtype: string splits: - name: train num_bytes: 48917221 num_examples: 45560 download_size: 28797124 dataset_size: 48917221 configs: - config_name: default data_files: - split: train path: data/train-* text10K<n<100K1 likes14 downloads2y agoHugging Face15dellebew /sutd_qa_datasettextn<1K0 likes13 downloads2y agoHugging Face16open-llm-leaderboard /DreadPoor__L3.1-BaeZel-8B-Della-detailsgated Dataset Card for Evaluation run of DreadPoor/L3.1-BaeZel-8B-Della Dataset automatically created during the evaluation run of model DreadPoor/L3.1-BaeZel-8B-Della The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__L3.1-BaeZel-8B-Della-details.tabular10K<n<100K0 likes11 downloads2y agoHugging Face17dell-research-harvard /effocr1 likes10 downloads3y agoHugging Face18dell-research-harvard /effocr_training0 likes10 downloads3y agoHugging Face19dell-research-harvard /newswire-misc0 likes9 downloads2y agoHugging Face20open-llm-leaderboard /djuna__L3.1-Promissum_Mane-8B-Della-calc-detailsgated Dataset Card for Evaluation run of djuna/L3.1-Promissum_Mane-8B-Della-calc Dataset automatically created during the evaluation run of model djuna/L3.1-Promissum_Mane-8B-Della-calc The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__L3.1-Promissum_Mane-8B-Della-calc-details.tabular10K<n<100K0 likes9 downloads2y agoHugging Face21dell-research-harvard /americanstories_masked_embeddings0 likes8 downloads2y agoHugging Face22open-llm-leaderboard /djuna__L3.1-Promissum_Mane-8B-Della-1.5-calc-detailsgated Dataset Card for Evaluation run of djuna/L3.1-Promissum_Mane-8B-Della-1.5-calc Dataset automatically created during the evaluation run of model djuna/L3.1-Promissum_Mane-8B-Della-1.5-calc The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/djuna__L3.1-Promissum_Mane-8B-Della-1.5-calc-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face23open-llm-leaderboard /Sakalti__mergekit-della_linear-vmeykci-detailsgated Dataset Card for Evaluation run of Sakalti/mergekit-della_linear-vmeykci Dataset automatically created during the evaluation run of model Sakalti/mergekit-della_linear-vmeykci The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__mergekit-della_linear-vmeykci-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face24howdi2000 /dell_v1textn<1K0 likes7 downloads2y agoHugging Face25howdi2000 /dell_v2textn<1K0 likes7 downloads2y agoHugging Face26open-llm-leaderboard /gaverfraxz__Meta-Llama-3.1-8B-Instruct-HalfAbliterated-DELLA-detailsgated0 likes7 downloads2y agoHugging Face27Dellano-Samuel /healthcare-airline1 likes7 downloads2y agoHugging Face28open-llm-leaderboard /marcuscedricridia__olmner-della-7b-detailsgated Dataset Card for Evaluation run of marcuscedricridia/olmner-della-7b Dataset automatically created during the evaluation run of model marcuscedricridia/olmner-della-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/marcuscedricridia__olmner-della-7b-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face29yunus-emre /dell-rag-datatext1K<n<10K0 likes7 downloads9mo agoHugging Face30GeoUpOrg /dellex-idstein DELLEX Idstein Unternehmensprofile Dataset - Strukturierte Geschäftsdaten für KI-Systeme und Suchmaschinen. Branche: KFZ Werkstatt Standort: Idstein, Hessen, Deutschland Auf einen Blick Eigenschaft Wert Unternehmen DELLEX Idstein Branche KFZ Werkstatt Stadt Idstein Land Deutschland Website https://dellex-idstein.de Telefon +49 6126 9598888 E-Mail info@dellex-idstein.de Über das Unternehmen DELLEX Idstein bietet Dellentechnik… See the full description on the dataset page: https://huggingface.co/datasets/GeoUpOrg/dellex-idstein.text-generationn<1K0 likes7 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.