CoolFace
Datasetpublic

RevolutionCrossroads/loc_chronicling_america_1770-1810_issues

Dataset Card for Chronicling America Historic American Newspapers 1770–1810 - Issue-Level Dataset Summary A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset provides an issue-level representation of the Chronicling America newspapers dataset, aggregating individual page records into complete newspaper issues with with publication metadata, original Chronicling… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes9.7kdownloads
Dataset Card

Dataset Card for Chronicling America Historic American Newspapers 1770–1810 - Issue-Level

Table of Contents


Dataset Summary

A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset provides an issue-level representation of the Chronicling America newspapers dataset, aggregating individual page records into complete newspaper issues with with publication metadata, original Chronicling America OCR, and AI-generated OCR outputs.

It is derived from the page-level dataset: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810

This dataset was prepared as part of the Revolution Crossroads project: https://www.si.edu/revolution-crossroads


Dataset Description

For the initiative Revolution Crossroads, the Smithsonian Institution prepared this dataset using data, metadata, and digital objects publicly available from the Chronicling America Historic American Newspapers collection.

Dataset Details

  • —Prepared by: Smithsonian Institution, Office of Digital & Innovation staff
  • —Shared by: Revolution Crossroads
  • —Language(s): English
  • —License: Public domain

Relationship to Source Dataset

This dataset is a transformation of the original Revolution Crossroads page-level dataset. For detailed information about:

  • —source materials
  • —digitization processes
  • —metadata definitions
  • —rights and provenance

refer to the page-level dataset card: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810/blob/main/README.md


Curation Rationale

The original dataset is structured at the page level, where each record represents a single newspaper page. However, newspapers are multi-page documents, and meaningful context often spans multiple pages within a single issue.

This issue-level dataset was created to:

  • —Preserve document-level context across pages
  • —Support OCR and text extraction workflows at the issue level
  • —Enable entity recognition and analysis across full newspaper issues

Dataset Creation

This dataset was derived from the page-level Chronicling America dataset by restructuring records from page-level to issue-level.

  • —Page-level records were grouped by issue (LCCN, issue date, edition_order)
  • —Issue-level records retain key bibliographic metadata from the source dataset
  • —Page counts were calculated for each issue (page_count)
  • —An issue identifier (issue_id) was constructed using the format lccn_issue_date_edition_order
  • —References to the source dataset were preserved

Each record in this dataset corresponds to a single newspaper issue.

Data Collection and Processing

Supporting Files Available

Issue-level PDF files are provided in the Files tab for this dataset.

  • —pdfs Contains compiled PDF files representing complete newspaper issues. Each PDF corresponds to a single record in the dataset and is constructed from the source page images. They are grouped into directories based on the newspaper (title) LCCN. Each pdf uses the naming format of the corresponding issueid, with "titleLCCNIssueDate_EditionNumber". Note: Some PDFs for newly added issues are forthcoming.
PDF Construction
  • —Source page images were downloaded in JPEG2000 (JP2) format to preserve the highest available image quality
  • —JP2 images were converted to JPG format for processing
  • —The JP2 and JPG files did not contain DPI values. DPI values were derived from the source page-level PDFs (either 300 or 400 DPI) and applied when compiling images into issue-level PDFs
  • —JPG images were assembled into a single PDF per issue

To ensure compatibility with downstream processing systems:

  • —Images exceeding 10,000 pixels in height or 4,800 pixels in width were resized
  • —Resizing was performed proportionally to preserve visual fidelity while maintaining a workable scale

AI OCR Enrichment

This dataset has been enriched with outputs generated by DataLab's Chandra OCR 2, a Vision Language Model (VLM) for document understanding.

Complete Chandra OCR 2 processing outputs are stored in the Revolution Crossroads VLM OCR Hugging Face bucket. The bucket includes a README describing the processing pipeline, output formats, and directory structure.

Two additional fields have been added:

  • —RevX_OCR_bucket provides a link to the complete Chandra OCR 2 processing outputs for a document. These outputs include structured JSON, Markdown, HTML, layout information, extracted images, quality metrics, and other artifacts generated during processing.
  • —RevX_OCR_text contains the page-level plain text extracted by Chandra OCR 2. The text is provided inline to simplify downstream analysis while the complete processing outputs remain available through the linked bucket.

These outputs represent the direct results of automated AI processing. They have not been manually validated or corrected and should be treated as machine-generated data.

Not every record in this dataset currently includes Chandra OCR 2 outputs. The initial processing was completed in March 2026. Since then, the source collections have been refreshed with additional materials. A second round of Chandra OCR 2 processing is currently underway, and additional processing outputs will be incorporated into future dataset releases.

Accessing a Specific Data Version

This dataset may be updated over time as source data is revised, normalized, or restructured. If you are using the data for a project, we recommend accessing a specific tagged version so your workflow remains stable and reproducible.

You can access a tagged version using the revision parameter in load_dataset:

python
from datasets import load_dataset
import polars as pl

dataset_name = "RevolutionCrossroads/loc_chronicling_america_1770-1810_issues"
commit_tag = "Source_2025-08-01"

ds = load_dataset(
    dataset_name,
    revision=commit_tag,
    split="train"
)

# Convert to pandas
pandas_df = ds.data.table.to_pandas()

# Convert to Polars
polars_df = pl.from_arrow(ds.data.table)

This example loads the specified tagged dataset version and converts it to either a pandas or Polars dataframe for analysis.

Available Data Versions
  • —`new_OCR_v1` – Adds Datalab AI-generated OCR while preserving the issue-level dataset structure.

Using a tagged version is recommended for research, publication, or project workflows where consistency over time is important.

Dataset Structure

Each record in the dataset corresponds to a single newspaper issue. Issue-level metadata (title, place, date, edition) is represented once per issue rather than repeated across pages.


Data Fields

  • —lccn (large_string) Library of Congress Control Number for the newspaper title. Example: sn82014385
  • —issue_date (timestamp[ns]) Date the issue was published. Example: 1809-07-08
  • —edition_order (int64) Edition number/order within the day. The default is 1. Example: 1
  • —issue_id (large_string) Unique identifier for the issue derived from LCCN, date, and edition. Example: sn82014385_1809-07-08_1
  • —newspaper_title (large_string) Title of the newspaper. Example: The Delaware gazette
  • —place_of_publication (large_string) City and state of publication. Example: Wilmington [Del.]
  • —page_count (int64) Number of pages in the issue. Example: 4
  • —issue_url (string) URL to the newspaper issue in Chronicling America. Example: https://www.loc.gov/item/sn82014385/1809-07-08/ed-1
  • —iiif_manifest (large_string) IIIF manifest URL for the issue. Example: https://www.loc.gov/item/sn82014385/1809-07-08/ed-1/manifest.json
  • —thumbnail_url (large list) List of thumbnail image URLs for pages in the issue. Example: ["https://tile.loc.gov/image-services/iiif/service:ndnp:.../0074/full/pct:6.25/0/default.jpg"]
  • —jpeg2000_url (large list) List of JPEG2000 (JP2) image URLs for each page in the issue. Example: ["https://tile.loc.gov/storage-services/service/ndnp/.../0074.jp2"]
  • —pdf_url (large_string) Redirect url to issue-level PDF as stored on Hugging Face in this dataset. Note: PDFs for newly added issues are forthcoming. Example: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues/resolve/main/pdfs/sn82014385/sn82014385_1809-07-08_1.pdf
  • —source_record_creation_date (timestamp[ns]) Date of source record creation. Example: 2017-05-24T12:39:57
  • —ca_ocr_url (large list) List of OCR XML file URLs for each page, as provided by Chronicling America contributing institutions. Example: ["https://tile.loc.gov/storage-services/service/ndnp/.../0074.xml"]
  • —ca_ocr_text (large list) Original OCR text for each page in the issue, generated by Chronicling America contributing institutions. Example: ["THE DELAWARE GAZETTE. VOL. I. WILMINGTON, SATURDAY, JULY 8, 1809..."]
  • —RevX_OCR_bucket (large_string) Link to the complete Chandra OCR 2 processing outputs for the newspaper issue. These outputs include structured JSON, Markdown, HTML, layout information, extracted images, quality metrics, and other processing artifacts. May be null if the issue has not yet been processed. Example: https://huggingface.co/buckets/RevolutionCrossroads/RevolutionCrossroads_VLM_OCR/tree/chronicling_america/sn82014385_parsed/sn82014385_1809-07-08_1.pdf
  • —RevX_OCR_text (large list) Page-level text extracted by DataLab's Chandra OCR 2. One list entry corresponds to each page of the newspaper issue. May be null if the issue has not yet been processed. Example: ["THE DELAWARE GAZETTE.\nVOL. I....", "..."]

Source Data

This dataset is derived from the Chronicling America Historic American Newspapers collection collection from the Library of Congress.

For full details on the source collection, program context, and historical background, refer to the page-level dataset card: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810/blob/main/README.md


Personal and Sensitive Information

No known personal or sensitive information is included beyond what appeared in publicly circulated newspapers of the time. Users should be aware that newspapers may contain outdated or offensive terminology reflecting the period in which they were created.


Considerations for Using the Data

Risks and Limitations

The dataset contains historical materials with language that does not always match the language preferred by members of the communities depicted. It may include negative stereotypes or words that offend. These materials reflect the views of their creators, not the Library of Congress or the United States government.

Machine-readable text in the dataset was created using Optical Character Recognition (OCR). Errors are present throughout the data, particularly when source images are degraded or typography is unusual.

For content published before 1810, historical typesetting (e.g., “long s” characters resembling “f”) may result in misread text.

Metadata and digitized content are updated by the Library of Congress on an ongoing basis.

Recommendations

Researchers should validate extracted information against images when accuracy is critical and be cautious about treating OCR text as complete or authoritative. Consultation of the Chronicling America collection is recommended for the most up-to-date records.


Additional Information

Citation Information

BibTeX @misc{revolution_crossroads_2026, author = { Revolution Crossroads }, title = { loc_chronicling_america_1770-1810_issues (Revision 4b7dc29) }, year = 2026, url = { https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues }, doi = { 10.57967/hf/8348 }, publisher = { Hugging Face } }

APA

Revolution Crossroads Project Team. (2025). Chronicling America: Historic American Newspapers 1770–1810 - Issue-level [Data set]. Hugging Face.


Glossary

  • —LCCN: Library of Congress Control Number, a unique identifier for newspaper titles
  • —OCR (Optical Character Recognition): Machine-generated text created from scanned page images
  • —Issue: A single publication of a newspaper on a given date, often containing multiple pages

Dataset Card Contact

revolutioncrossroads@si.edu