CoolFace
Datasetpublic

AndreaBozzo/ceres-open-data-index

Ceres Open Data Index A curated index of 2,745,021 open dataset records from 567 portal exports across 40+ countries and international sources, with cross-portal duplicates flagged. The snapshot includes normalized dataset metadata, resource links when supplied by the source, and a reproducible manifest with file checksums. Dataset description This dataset was harvested by Ceres, a harvest-first open-data toolkit. It covers ten source families: CKAN, DCAT-AP… See the full description on the dataset page: https://huggingface.co/datasets/AndreaBozzo/ceres-open-data-index.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes210downloads
Dataset Card

Ceres Open Data Index

A curated index of 2,745,021 open dataset records from 567 portal exports across 40+ countries and international sources, with cross-portal duplicates flagged. The snapshot includes normalized dataset metadata, resource links when supplied by the source, and a reproducible manifest with file checksums.

Dataset description

This dataset was harvested by Ceres, a harvest-first open-data toolkit. It covers ten source families: CKAN, DCAT-AP (udata REST), SPARQL-backed DCAT, Project Open Data (data.json), Socrata, OpenDataSoft, ArcGIS Hub, OGC Records/CSW, collection-level STAC, and SDMX. Noise is filtered during export; likely cross-portal duplicates are flagged but retained.

What's included

  • —Titles, descriptions, URLs, tags, organizations, licenses, and source provenance
  • —Resource-level names, formats, media types, URLs, and schema field counts when available
  • —A heuristic cross-portal duplicate signal (is_duplicate)
  • —Per-portal Parquet subsets for targeted analysis
  • —A versioned manifest, quality report, and snapshot-to-snapshot changelog

What's not included

  • —The data files themselves (CSVs, shapefiles, and similar assets)
  • —Embedding vectors, which are model-specific
  • —A canonical entity-resolution result; is_duplicate is only a title-based heuristic
Resource links are present for 1,997,796 rows (72.8%); 96,395 rows (3.5%) include at least one resource schema. Coverage varies substantially by source API.

Dataset structure

Schema

ColumnTypeNullableDescription
original_idstringnoRecord identifier from the source portal
source_portalstringnoCanonical source portal URL
portal_namestringnoStable export name for the portal
urlstringnoDataset or catalog-record landing page
titlestringnoRecord title
descriptionstringyesRecord description
record_kindstringnoNormalized kind of catalog record
tagsstringyesComma-separated source tags
organizationstringyesPublishing organization
licensestringyesLicense title or identifier
metadata_createdstringyesSource creation date (ISO 8601)
metadata_modifiedstringyesSource modification date (ISO 8601)
first_seen_atstringnoWhen Ceres first indexed the record (RFC 3339)
languagestringyesSource or portal language label
is_duplicatebooleannoExact case-insensitive title also occurs on another canonical portal
resourceslist of structsnoResource entries: name, format, media_type, url, and field_count

Files

  • —all.parquet — complete dataset (2,745,021 rows)
  • —data/<portal>.parquet — 567 portal subsets; rows repeat those in all.parquet
  • —identity.parquet — record keys and content hashes used for snapshot diffs
  • —metadata.json — snapshot manifest, Ceres provenance, curation counts, aliases, and SHA-256 checksums
  • —reports.json / report.md — coverage and field-completeness report
  • —changelog.json / changelog.md — added, changed, removed, and unchanged counts versus v5

Manifest paths are unique. This release contains 569 Parquet files in total (all.parquet, identity.parquet, and 567 portal subsets), and every file's size, checksum, and row count was verified before publication.

Largest portal exports

Top 48 by exported row count (of 567 total):

Portal exportSourceFamily/profileRows
eu-open-datadata.europa.euDCAT / SPARQL712,481
catalog-data-govcatalog.data.govCKAN399,597
queenslanddata.qld.gov.auCKAN181,970
www-govdata-degovdata.deCKAN158,177
data-gov-audata.gov.auCKAN139,729
ods-hubdata.opendatasoft.comOpenDataSoft91,100
dati-gov-itdati.gov.itCKAN65,077
ckan-publishing-service-gov-ukckan.publishing.service.gov.ukCKAN57,244
data-gov-ukdata.gov.ukCKAN57,060
spain-datosdatos.gob.esDCAT / SPARQL49,930
open-canada-caopen.canada.caCKAN47,437
data-gouv-frdata.gouv.frDCAT / udata43,186
data-gov-uadata.gov.uaCKAN36,960
banque-de-francewebstat.banque-france.frOpenDataSoft35,837
czech-republicdata.gov.czDCAT / SPARQL32,660
data-humdata-orgdata.humdata.orgCKAN28,039
geodati-gov-itgeodati.gov.itgeospatial catalog23,578
healthdata-govhealthdata.govProject Open Data22,417
greecedata.gov.grCKAN21,900
data-gov-iedata.gov.ieCKAN21,267
california-natural-resourcesdata.cnra.ca.govCKAN21,114
data-overheid-nldata.overheid.nlCKAN20,796
dados-gov-ptdados.gov.ptDCAT / udata20,792
spain-idee-cswidee.esOGC Records/CSW18,690
pndbpndb.opendatasoft.comOpenDataSoft17,710
netl-edxedx.netl.doe.govCKAN17,186
switzerlandopendata.swissCKAN15,863
www-geocatalogue-frgeocatalogue.frgeospatial catalog14,960
nrwopen.nrw.deCKAN14,910
japandata.e-gov.go.jpCKAN12,276
dati-toscana-itdati.toscana.itCKAN12,214
slovakiadata.slovensko.skDCAT / SPARQL11,748
metadane-podgik-plmetadane.podgik.plgeospatial catalog11,264
statistics-lithuania-sdmxosp-rs.stat.gov.ltSDMX9,500
www-geonorge-nogeonorge.nogeospatial catalog9,190
podatki-gov-sipodatki.gov.siCKAN8,956
norwaydata.norge.noDCAT / SPARQL8,669
catalog-data-metro-tokyo-lg-jpmetro.tokyo.lg.jpCKAN8,483
nationaalgeoregister-nlnationaalgeoregister.nlgeospatial catalog8,140
www-datos-gov-codatos.gov.coSocrata7,773
emodnet-cswemodnet.ec.europa.euOGC Records/CSW7,104
www-datosabiertos-gob-padatosabiertos.gob.paCKAN5,607
discover-data-vic-gov-audiscover.data.vic.gov.auCKAN5,578
dati-regione-marche-itdati.regione.marche.itCKAN5,450
data-gov-rodata.gov.roCKAN5,018
data-virginia-govdata.virginia.govProject Open Data4,825
data-ca-govdata.ca.govProject Open Data4,506
cioos-pacificcatalogue.cioos.caCKAN3,484

The full portal breakdown and exact source URLs are recorded in metadata.json and reports.json. Alias feeds such as a portal's root API and its /data.json endpoint can share a canonical export name without being treated as independent sources.

Coverage and quality

Exporter classification by source type: unmapped/legacy config 1,318,679 · DCAT 835,846 · CKAN 327,737 · OpenDataSoft 177,191 · OGC Records/CSW 27,264 · ArcGIS 24,302 · SDMX 15,160 · Socrata 11,833 · STAC 7,009. The unknown classification reflects missing type annotations in the export configuration, not unknown provenance; source URLs remain present.

DCAT profiles account for 817,318 SPARQL, 12,058 static JSON, and 3,477 udata REST records. Field completeness is: description 100.0% · resources 72.8% · modification date 52.4% · organization 51.2% · tags 45.3% · license 31.8% · resource schema 3.5%.

Language labels are source/portal metadata rather than per-row language detection. The largest labels are English (en) 1,182,023; unknown 1,161,936; eng 87,963; Spanish 76,992; French 75,086; Ukrainian 36,301; Czech 32,879; Greek 21,900; Japanese 14,224; Slovak 11,748; and Norwegian 8,669.

Loading the dataset

python
import pandas as pd

df = pd.read_parquet("all.parquet")
df_us = pd.read_parquet("data/catalog-data-gov.parquet")
likely_unique = df[~df["is_duplicate"]]

Or with DuckDB:

sql
SELECT source_portal, count(*) AS records
FROM read_parquet('all.parquet')
GROUP BY source_portal
ORDER BY records DESC;

Dataset creation

Harvesting methodology

Ceres harvests each portal through its native API with a streaming, incremental pipeline:

  1. 1.CKAN portals use paged package_search requests with adaptive page sizing and retry behavior.
  2. 2.DCAT udata catalogs use paged JSON-LD; Project Open Data catalogs consume data.json documents.
  3. 3.SPARQL-backed DCAT catalogs use stable keyset pagination, bounded distribution-property queries, and URI-level deduplication.
  4. 4.Socrata uses the Discovery API; OpenDataSoft uses the Explore API with a keyset cursor beyond its 10,000-row window.
  5. 5.ArcGIS Hub, OGC Records/CSW, collection-level STAC, and SDMX clients preserve family-specific source metadata.
  6. 6.ceres export --format parquet streams PostgreSQL records, filters noise, flattens metadata, flags cross-portal title matches, and writes compressed Parquet plus checksummed reports.

The data.europa.eu refresh was intentionally stopped after resource enrichment reached approximately half the exported portal rows so this snapshot would not delay Ceres 0.7.0. The export contains 712,481 data.europa.eu records, of which 366,307 (51.4%) include resource metadata. Completing the remaining enrichment is tracked in Ceres issue #254 for the next release.

Curation

  • —Noise filtering: 81,766 rows removed (short/noise titles and empty descriptions)
  • —Duplicate flagging: 1,354,529 rows marked; records are retained and no canonical entity resolution is claimed
  • —Alias normalization: known mirror or alternate endpoints can share a canonical portal identity
  • —Metadata flattening: nested source fields and resource metadata are normalized into the published schema

The duplicate method and version are recorded under duplicate_detection in metadata.json.

Update and provenance

Snapshot `ceres-20260819-884e6da7b4b0` was generated on 2026-08-19 by Ceres 0.7.0 at commit `2583fcf`. Relative to ceres-20260715-ed0c345e8663, it has 183,625 added, 172,349 changed, 2,613 removed, and 2,389,047 unchanged records. See changelog.md for the per-portal breakdown.

Considerations

Known biases and limitations

  • —Large aggregators dominate the index: data.europa.eu contributes 26.0% and catalog.data.gov 14.6% of rows.
  • —The partial data.europa.eu refresh means resource coverage is materially better than v5 but not complete.
  • —49.3% of rows are flagged as possible cross-portal title duplicates, especially because aggregators re-list member-portal records.
  • —Metadata quality, vocabulary, timestamps, and license conventions vary by source; only 31.8% of rows have a non-empty license.
  • —SPARQL, STAC, OGC, and SDMX records may have fewer normalized fields than CKAN or Socrata records.
  • —Exporter language and source-type annotations are incomplete for legacy configurations, but original source URLs and metadata are retained.

Ethical considerations

This dataset contains public catalog metadata from government and institutional portals. It does not intentionally contain the underlying data files. Source licenses vary; consumers must follow the license and terms attached to each source record.

Additional information

License

This aggregate dataset and its tooling are released under Apache 2.0. Underlying source metadata may carry separate open-data licenses recorded per row where available.

Citation

bibtex
@misc{ceres-open-data-index-2026,
  title={Ceres Open Data Index},
  author={Andrea Bozzo},
  year={2026},
  url={https://github.com/AndreaBozzo/Ceres},
  note={2,745,021 records from 567 portal exports, snapshot 2026-08-19}
}

Links