datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vnexpress_plain_textvtv_plain_textTauri-RL-Plaintext-System-V2borges_plain_text_dataset
Dataset: Borges en texto plano
El objetivo de este repositorio es construir un dataset del gran autor argentino que pueda usarse para el entrenamiento de modelos de lenguaje.
Inicialmente partí de libros en formato EPUB y únicamente en español
Carpetas
Inicialmente planteo tres carpetas
Epub
Libros en este formato
Epub_a_txt
Libros convertidos con el sencillo script disponible en
https://github.com/lucasbiagettia/epub2txt
txt_limpios
A… See the full description on the dataset page: https://huggingface.co/datasets/lucasbiagettia/borges_plain_text_dataset.vnexpress_plain_text_suc_khoevnexpress_plain_text_thoi_suvnexpress_plain_text_doi_songwiki-scibatch-cohere-plaintext-banaei-9mvnexpress_plain_text_phap_luatPlain-text-pretrainingvnexpress_plain_text_cuoivnexpress_plain_text_giai_trivnexpress_plain_text_the_thaoSQL_PlainText_Combined
Dataset Card for "SQL_PlainText_Combined"
More Information needed
applicant-job-plaintextautoeval-eval-piaf-plain_text-42b979-39890145062
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Question Answering
Model: etalab-ia/camembert-base-squadFR-fquad-piaf
Dataset: piaf
Config: plain_text
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @malou.berthe@gmail.com for evaluating this model.
autoeval-staging-eval-launch__gov_report-plain_text-cd8e90-16116210
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: Blaise-g/longt5_tglobal_large_sumpubmed
Dataset: launch/gov_report
Config: plain_text
Split: validation
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
hanse-niederdeutsch-plaintext
Middle Low German Administrative & Legal Sources (14th–15th century)
Dataset Description
This dataset consists of plain text documents in Middle Low German drawn from various administrative and legal sources as well as charters, spanning the mid-14th to the late 15th century.
Data Sources
Minutes of the Hanse town Assemblies (Hanserezesse) transcribed at the FGHO
Charters (Urkunden), administrative records (Verwaltung), and legal sources… See the full description on the dataset page: https://huggingface.co/datasets/fgho/hanse-niederdeutsch-plaintext.autoeval-eval-imdb-plain_text-7301d6-42320145096
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Binary Text Classification
Model: fabriceyhc/bert-base-uncased-imdb
Dataset: imdb
Config: plain_text
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @test for evaluating this model.
Img2Text-Plaintext-Retrieval
Img2Text-Plaintext-Retrieval Dataset
Dataset Overview
The Img2Text-Plaintext-Retrieval dataset is designed for retrieving plaintext descriptions from corresponding algorithm images. This dataset consists of structured text, raw text, algorithm images, and metadata such as source URLs and filenames. It is suitable for tasks like OCR-based text retrieval, image-to-text learning, and document understanding.
Dataset Details
Modality: Image, Text
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Img2Text-Plaintext-Retrieval.autoeval-staging-eval-launch__gov_report-plain_text-2fa37c-16136227
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: pszemraj/long-t5-tglobal-base-16384-booksum-V11-big_patent-V2
Dataset: launch/gov_report
Config: plain_text
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
autoeval-staging-eval-launch__gov_report-plain_text-1abd3a-16146235
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: facebook/bart-large-cnn
Dataset: launch/gov_report
Config: plain_text
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
vnexpress_plain_text_bat_dong_sanautoeval-staging-eval-anli-plain_text-c507f2-14355972
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Natural Language Inference
Model: MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli
Dataset: anli
Config: plain_text
Split: test_r3
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @MoritzLaurer for evaluating this model.
autoeval-staging-eval-launch__gov_report-plain_text-7b7f8a-16126221
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: google/bigbird-pegasus-large-pubmed
Dataset: launch/gov_report
Config: plain_text
Split: validation
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
autoeval-eval-launch__gov_report-plain_text-c8c9c8-1465553968
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: pszemraj/long-t5-tglobal-base-16384-booksum-V12
Dataset: launch/gov_report
Config: plain_text
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @pszemraj for evaluating this model.
autoeval-staging-eval-launch__gov_report-plain_text-cd8e90-16116216
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: Blaise-g/longt5_tglobal_large_scitldr
Dataset: launch/gov_report
Config: plain_text
Split: validation
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
autoeval-staging-eval-launch__gov_report-plain_text-1abd3a-16146233
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: google/bigbird-pegasus-large-pubmed
Dataset: launch/gov_report
Config: plain_text
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
autoeval-eval-launch__gov_report-plain_text-45e121-1564955706
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: pszemraj/long-t5-tglobal-large-pubmed-3k-booksum-16384-WIP15
Dataset: launch/gov_report
Config: plain_text
Split: test
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @pszemraj for evaluating this model.
autoeval-staging-eval-launch__gov_report-plain_text-cd8e90-16116213
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Summarization
Model: pszemraj/long-t5-tglobal-large-pubmed-3k-booksum-16384-WIP
Dataset: launch/gov_report
Config: plain_text
Split: validation
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @nonchalant-nagavalli for evaluating this model.
