CoolFace
8 results

government data

zalizedata /government-tenders-contract-awards-dataset Government Tenders & Contract Awards (US Federal + Global Portals) Public procurement notices and contract awards — 1.05M US federal (SAM.gov) records in the S tier, up to 11.5M rows across global portals with supplier-linked awards in the full pack. Part of the DataForge Open Data program — full production packages, free for academic and personal use. Canonical dataset page: https://data.zalize.com/datasets/government-tenders-contract-awards-dataset Packages in this… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/government-tenders-contract-awards-dataset.tabulartext-classification10M<n<100M0 likes508 downloads2mo agoHugging FaceLukaszl /pl-government-docs-mix-ocr-dataset Polish municipal administrative documents OCR dataset An OCR-oriented image dataset built from publicly available Polish municipal administrative documents. The dataset contains pages collected from materials such as: resolutions (uchwały) ordinances / regulations annexes official notices tabular administrative pages stamped and signed office documents other municipal administrative materials All files in this repository are provided as images only. What is inside… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset.imageimage-to-textn<1K1 likes115 downloads6mo agoHugging Facemilenamileentje /Dutch-Government-Data-for-Bias-detectiontabulartext-classification1K<n<10K2 likes43 downloads2y agoHugging FaceLukaszl /pl-government-docs-mix-ocr-dataset-v1 Document OCR using GLM-OCR This dataset contains OCR results from images in Lukaszl/pl-government-docs-mix-ocr-dataset using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance. Processing Details Source Dataset: Lukaszl/pl-government-docs-mix-ocr-dataset Model: zai-org/GLM-OCR Task: text recognition Number of Samples: 169 Processing Time: 9.2 min Processing Date: 2026-04-03 11:18 UTC Configuration Image Column: image Output Column: markdown… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset-v1.0 likes27 downloads6mo agoHugging Facesundowndfp /my_pusht_government_dataThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "unknown", "total_episodes": 206, "total_frames": 25650, "total_tasks": 1, "total_videos": 206, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:206" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sundowndfp/my_pusht_government_data.tabularrobotics10K<n<100K0 likes25 downloads1mo agoHugging FaceLukaszl /pl-government-docs-mix-ocr-dataset-v1-results OCR Bench Results: Polish government documents benchmark VLM-as-judge pairwise evaluation of OCR models on a dataset of real Polish government and public administration documents. This benchmark focuses on structured, text-heavy documents typical for public institutions, including official forms, templates, administrative documents, and scanned materials. As with all OCR benchmarks, results are document-type specific and should not be interpreted as a universal ranking across all… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/pl-government-docs-mix-ocr-dataset-v1-results.tabular1K<n<10K1 likes24 downloads6mo agoHugging Face