datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gradients_Gradients_and_Text_Full_Logic_Captionspagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
pagoda-text-and-image-dataset-small
Dataset Card for "pagoda-text-and-image-dataset-small"
More Information needed
Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.text_and_concat_image_hf_version_epoch_1_with_prefix_with_exist_split_fixed_best_of_16_CoTshilla-clothing-text-and-image-dataset
Dataset Card for "shilla-clothing-text-and-image-dataset"
More Information needed
textandstuff
Text&Stuff
An dataset for testing OCR and denoising methods, this dataset only contains crystal-clear text, ready for testing.
pagoda-text-and-image-dataset-steeple
Dataset Card for "pagoda-text-and-image-dataset-steeple"
More Information needed
RISEBench_withGPT4o_TextAndImagerick_and_morty_text_to_image
Dataset Card for "rick_and_morty_text_to_image"
More Information needed
rick_and_morty_image_and_text
Dataset Card for "rick_and_morty_image_and_text"
More Information needed
AnyEdit_withGPT4o_TextAndImageKRIS_Bench_withGPT4o_TextAndImage40K_kashmiri_text_and_image_dataset
40K Kashmiri Words with images
Kashmiri (words) Image and Text Dataset for OCR Models
This repository contains a dataset specifically curated for training and testing Optical Character Recognition (OCR) models on Kashmiri language text. The dataset includes a large collection of images with their corresponding labels in a CSV file, designed to aid in the development of robust OCR solutions for the Kashmiri script.
Directory Structure:
Zip File/
├── images/
│ ├── 000001.png
│… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/40K_kashmiri_text_and_image_dataset.31K_Kashmiri_text_and_image_dataset_for_text_RecognitionDirectory Structure:
Zip File/
├── images/
│ ├── 000001.png
│ ├── 000002.png
│ ├── 000003.png
│ ├── ...
│ ├── 031000.png
├── labels.csv
└── metadata.json
Description:
Zip File: The root directory that contains all other files and folders.
images/: A folder containing 31,000 images named sequentially from 000001.png to 031000.png.
labels.csv: A CSV file that includes information about the labels for each image.
metadata.json: A JSON file that contains metadata about the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/31K_Kashmiri_text_and_image_dataset_for_text_Recognition.cat_and_dog_text_to_text_image
