datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twin-cities-public-records
Twin Cities public records, joined
25 datasets · 1,575,384 rows · free, CC BY 4.0 · mirrored from brickandmortar.dev
A city emits records constantly — parcels, recorded sales, assessments, permits, licences, inspections, 911 calls, cleanup sites, flood zones, federal loans, wages, census measures — and almost nobody joins them. These are the joined slices, published as files rather than as an API you have to ask for a key to. The join is the work; the data is free.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/brickandmortar/twin-cities-public-records.twinviews-13k
Dataset Card for TwinViews-13k
This dataset contains 13,855 pairs of left-leaning and right-leaning political statements matched by topic. The dataset was generated using GPT-3.5 Turbo and has been audited to ensure quality and ideological balance. It is designed to facilitate the study of political bias in reward models and language models, with a focus on the relationship between truthfulness and political views.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/wwbrannon/twinviews-13k.decision-twin-v0-1
Decision Twin v0.1 encoder seed dataset
This dataset turns the user’s confirmed purchase and creative preferences into sentence-pair classification groups. Each context is phrased as a first-person request that a buying or creative assistant could receive.
Main files
File
Purpose
decision_twin_encoder.csv
Same rows in CSV form.
train.csv, validation.csv, test.csv
Group-safe 80/10/10 CSV splits, generated with seed 42.
decision_groups.jsonl
Group IDs… See the full description on the dataset page: https://huggingface.co/datasets/JohnGorri/decision-twin-v0-1.forge-pump-digital-twin-synthetic
Forge Pump Digital Twin — Synthetic
A deterministic, clean-room tabular baseline for pump surrogate modelling, edge-runtime conformance, and advisory anomaly examples. It contains 40,000 synthetic rows split into 28,000 train, 6,000 validation, and 6,000 test rows.
This dataset contains no plant telemetry, customer data, equipment identifiers, CAD, BOMs, nameplates, vendor curves, or values copied from a private repository. Every constant is an illustrative engineering proxy. It… See the full description on the dataset page: https://huggingface.co/datasets/sankalpsthakur/forge-pump-digital-twin-synthetic.digital_twin
