RevolutionCrossroads/loc_chronicling_america_1770-1810_issues
Dataset Card for Chronicling America Historic American Newspapers 1770–1810 - Issue-Level Dataset Summary A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset provides an issue-level representation of the Chronicling America newspapers dataset, aggregating individual page records into complete newspaper issues with with publication metadata, original Chronicling… See the full description on the dataset page: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues.
Dataset Card for Chronicling America Historic American Newspapers 1770–1810 - Issue-Level
Table of Contents
- Dataset Summary
- Dataset Description
- Dataset Details
- Relationship to Source Dataset
- Curation Rationale
- Dataset Creation
- Data Collection and Processing
- Supporting Files Available
- AI OCR Enrichment
- Accessing a Specific Data Version
- Dataset Structure
- Data Fields
- Source Data
- Personal and Sensitive Information
- Considerations for Using the Data
- Additional Information
Dataset Summary
A dataset drawn from the Library of Congress Chronicling America digital collection, part of the National Digital Newspaper Program (NDNP). This dataset provides an issue-level representation of the Chronicling America newspapers dataset, aggregating individual page records into complete newspaper issues with with publication metadata, original Chronicling America OCR, and AI-generated OCR outputs.
It is derived from the page-level dataset: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810
This dataset was prepared as part of the Revolution Crossroads project: https://www.si.edu/revolution-crossroads
Dataset Description
For the initiative Revolution Crossroads, the Smithsonian Institution prepared this dataset using data, metadata, and digital objects publicly available from the Chronicling America Historic American Newspapers collection.
Dataset Details
- Prepared by: Smithsonian Institution, Office of Digital & Innovation staff
- Shared by: Revolution Crossroads
- Language(s): English
- License: Public domain
Relationship to Source Dataset
This dataset is a transformation of the original Revolution Crossroads page-level dataset. For detailed information about:
- source materials
- digitization processes
- metadata definitions
- rights and provenance
refer to the page-level dataset card: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810/blob/main/README.md
Curation Rationale
The original dataset is structured at the page level, where each record represents a single newspaper page. However, newspapers are multi-page documents, and meaningful context often spans multiple pages within a single issue.
This issue-level dataset was created to:
- Preserve document-level context across pages
- Support OCR and text extraction workflows at the issue level
- Enable entity recognition and analysis across full newspaper issues
Dataset Creation
This dataset was derived from the page-level Chronicling America dataset by restructuring records from page-level to issue-level.
- Page-level records were grouped by issue (LCCN, issue date, edition_order)
- Issue-level records retain key bibliographic metadata from the source dataset
- Page counts were calculated for each issue (
page_count) - An issue identifier (
issue_id) was constructed using the formatlccn_issue_date_edition_order - References to the source dataset were preserved
Each record in this dataset corresponds to a single newspaper issue.
Data Collection and Processing
Supporting Files Available
Issue-level PDF files are provided in the Files tab for this dataset.
- pdfs Contains compiled PDF files representing complete newspaper issues. Each PDF corresponds to a single record in the dataset and is constructed from the source page images. They are grouped into directories based on the newspaper (title) LCCN. Each pdf uses the naming format of the corresponding issueid, with "titleLCCNIssueDate_EditionNumber". Note: Some PDFs for newly added issues are forthcoming.
PDF Construction
- Source page images were downloaded in JPEG2000 (JP2) format to preserve the highest available image quality
- JP2 images were converted to JPG format for processing
- The JP2 and JPG files did not contain DPI values. DPI values were derived from the source page-level PDFs (either 300 or 400 DPI) and applied when compiling images into issue-level PDFs
- JPG images were assembled into a single PDF per issue
To ensure compatibility with downstream processing systems:
- Images exceeding 10,000 pixels in height or 4,800 pixels in width were resized
- Resizing was performed proportionally to preserve visual fidelity while maintaining a workable scale
AI OCR Enrichment
This dataset has been enriched with outputs generated by DataLab's Chandra OCR 2, a Vision Language Model (VLM) for document understanding.
Complete Chandra OCR 2 processing outputs are stored in the Revolution Crossroads VLM OCR Hugging Face bucket. The bucket includes a README describing the processing pipeline, output formats, and directory structure.
- Complete VLM OCR Bucket: RevolutionCrossroads_VLM_OCR
- Chronicling America: chronicling_america
Two additional fields have been added:
- RevX_OCR_bucket provides a link to the complete Chandra OCR 2 processing outputs for a document. These outputs include structured JSON, Markdown, HTML, layout information, extracted images, quality metrics, and other artifacts generated during processing.
- RevX_OCR_text contains the page-level plain text extracted by Chandra OCR 2. The text is provided inline to simplify downstream analysis while the complete processing outputs remain available through the linked bucket.
These outputs represent the direct results of automated AI processing. They have not been manually validated or corrected and should be treated as machine-generated data.
Not every record in this dataset currently includes Chandra OCR 2 outputs. The initial processing was completed in March 2026. Since then, the source collections have been refreshed with additional materials. A second round of Chandra OCR 2 processing is currently underway, and additional processing outputs will be incorporated into future dataset releases.
Accessing a Specific Data Version
This dataset may be updated over time as source data is revised, normalized, or restructured. If you are using the data for a project, we recommend accessing a specific tagged version so your workflow remains stable and reproducible.
You can access a tagged version using the revision parameter in load_dataset:
from datasets import load_dataset
import polars as pl
dataset_name = "RevolutionCrossroads/loc_chronicling_america_1770-1810_issues"
commit_tag = "Source_2025-08-01"
ds = load_dataset(
dataset_name,
revision=commit_tag,
split="train"
)
# Convert to pandas
pandas_df = ds.data.table.to_pandas()
# Convert to Polars
polars_df = pl.from_arrow(ds.data.table)This example loads the specified tagged dataset version and converts it to either a pandas or Polars dataframe for analysis.
Available Data Versions
- `Source_2025-08-01` – Updated source data.
- `new_OCR_v1` – Adds Datalab AI-generated OCR while preserving the issue-level dataset structure.
Using a tagged version is recommended for research, publication, or project workflows where consistency over time is important.
Dataset Structure
Each record in the dataset corresponds to a single newspaper issue. Issue-level metadata (title, place, date, edition) is represented once per issue rather than repeated across pages.
Data Fields
- lccn (large_string) Library of Congress Control Number for the newspaper title. Example:
sn82014385
- issue_date (timestamp[ns]) Date the issue was published. Example:
1809-07-08
- edition_order (int64) Edition number/order within the day. The default is 1. Example:
1
- issue_id (large_string) Unique identifier for the issue derived from LCCN, date, and edition. Example:
sn82014385_1809-07-08_1
- newspaper_title (large_string) Title of the newspaper. Example:
The Delaware gazette
- place_of_publication (large_string) City and state of publication. Example:
Wilmington [Del.]
- page_count (int64) Number of pages in the issue. Example:
4
- issue_url (string) URL to the newspaper issue in Chronicling America. Example:
https://www.loc.gov/item/sn82014385/1809-07-08/ed-1
- iiif_manifest (large_string) IIIF manifest URL for the issue. Example:
https://www.loc.gov/item/sn82014385/1809-07-08/ed-1/manifest.json
- thumbnail_url (large list) List of thumbnail image URLs for pages in the issue. Example:
["https://tile.loc.gov/image-services/iiif/service:ndnp:.../0074/full/pct:6.25/0/default.jpg"]
- jpeg2000_url (large list) List of JPEG2000 (JP2) image URLs for each page in the issue. Example:
["https://tile.loc.gov/storage-services/service/ndnp/.../0074.jp2"]
- pdf_url (large_string) Redirect url to issue-level PDF as stored on Hugging Face in this dataset. Note: PDFs for newly added issues are forthcoming. Example:
https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues/resolve/main/pdfs/sn82014385/sn82014385_1809-07-08_1.pdf
- source_record_creation_date (timestamp[ns]) Date of source record creation. Example:
2017-05-24T12:39:57
- ca_ocr_url (large list) List of OCR XML file URLs for each page, as provided by Chronicling America contributing institutions. Example:
["https://tile.loc.gov/storage-services/service/ndnp/.../0074.xml"]
- ca_ocr_text (large list) Original OCR text for each page in the issue, generated by Chronicling America contributing institutions. Example:
["THE DELAWARE GAZETTE. VOL. I. WILMINGTON, SATURDAY, JULY 8, 1809..."]
- RevX_OCR_bucket (large_string) Link to the complete Chandra OCR 2 processing outputs for the newspaper issue. These outputs include structured JSON, Markdown, HTML, layout information, extracted images, quality metrics, and other processing artifacts. May be null if the issue has not yet been processed. Example:
https://huggingface.co/buckets/RevolutionCrossroads/RevolutionCrossroads_VLM_OCR/tree/chronicling_america/sn82014385_parsed/sn82014385_1809-07-08_1.pdf
- RevX_OCR_text (large list) Page-level text extracted by DataLab's Chandra OCR 2. One list entry corresponds to each page of the newspaper issue. May be null if the issue has not yet been processed. Example:
["THE DELAWARE GAZETTE.\nVOL. I....", "..."]
Source Data
This dataset is derived from the Chronicling America Historic American Newspapers collection collection from the Library of Congress.
For full details on the source collection, program context, and historical background, refer to the page-level dataset card: https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810/blob/main/README.md
Personal and Sensitive Information
No known personal or sensitive information is included beyond what appeared in publicly circulated newspapers of the time. Users should be aware that newspapers may contain outdated or offensive terminology reflecting the period in which they were created.
Considerations for Using the Data
Risks and Limitations
The dataset contains historical materials with language that does not always match the language preferred by members of the communities depicted. It may include negative stereotypes or words that offend. These materials reflect the views of their creators, not the Library of Congress or the United States government.
Machine-readable text in the dataset was created using Optical Character Recognition (OCR). Errors are present throughout the data, particularly when source images are degraded or typography is unusual.
For content published before 1810, historical typesetting (e.g., “long s” characters resembling “f”) may result in misread text.
Metadata and digitized content are updated by the Library of Congress on an ongoing basis.
Recommendations
Researchers should validate extracted information against images when accuracy is critical and be cautious about treating OCR text as complete or authoritative. Consultation of the Chronicling America collection is recommended for the most up-to-date records.
Additional Information
Citation Information
BibTeX @misc{revolution_crossroads_2026, author = { Revolution Crossroads }, title = { loc_chronicling_america_1770-1810_issues (Revision 4b7dc29) }, year = 2026, url = { https://huggingface.co/datasets/RevolutionCrossroads/loc_chronicling_america_1770-1810_issues }, doi = { 10.57967/hf/8348 }, publisher = { Hugging Face } }
APA
Revolution Crossroads Project Team. (2025). Chronicling America: Historic American Newspapers 1770–1810 - Issue-level [Data set]. Hugging Face.
Glossary
- LCCN: Library of Congress Control Number, a unique identifier for newspaper titles
- OCR (Optical Character Recognition): Machine-generated text created from scanned page images
- Issue: A single publication of a newspaper on a given date, often containing multiple pages
Dataset Card Contact
revolutioncrossroads@si.edu
