document-text-text
Finance-Document-Text-Classificationautotrain-document-text-language-ar-en-zh-3338392240BERTopic-summcomparer-gauntlet-v0p1-all-roberta-large-v1-document_textBERTopic-summcomparer-gauntlet-v0p1-sentence-t5-xl-document_texttext-summarization-for-technical-documentationdocument_text_readerdocument-classifications-using-bert-and-textcnndocumentry-style-text-to-speech
all-document-text-data
Climate Policy Radar Open Data
This repo contains the full text data of all of the documents from the Climate Policy Radar database (CPR), which is also available at Climate Change Laws of the World (CCLW).
Please note that this replaces the Global Stocktake open dataset: that data, including all NDCs and IPCC reports is now a subset of this dataset.
What’s in this dataset
This dataset contains two corpus types (groups of the same types or sources of documents) which… See the full description on the dataset page: https://huggingface.co/datasets/ClimatePolicyRadar/all-document-text-data.document-ocr-video-text-mini
Document OCR Video Text Data Notes
Dataset summary
This data card accompanies a lightweight Document OCR loader for Video Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/ankitamam/document-ocr-video-text-mini.high_synthetic_wrapmedium_text_document.part_ag_filesdocument-ocr-pointcloud-text-clean
Document OCR Pointcloud Text Data Notes
Dataset summary
This data card accompanies a lightweight Document OCR loader for Pointcloud Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.… See the full description on the dataset page: https://huggingface.co/datasets/kkumarmanoj/document-ocr-pointcloud-text-clean.document-ocr-pointcloud-text-mini
Document OCR Pointcloud Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Document OCR work with Pointcloud Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/sshenfiona1992/document-ocr-pointcloud-text-mini.document-ocr-image-text-v2-2024
Document OCR Image Text Data Notes
Dataset summary
Preparation notes and schema examples for Document OCR tasks using Image Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
load_data.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/pan-dece/document-ocr-image-text-v2-2024.
