datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pde-transformer-ape2dpdecert-pilot
PDECert Natural-Candidate Pilot
This is a provenance-bearing pilot benchmark for checking symbolic candidate
solutions to partial differential equations. Each row contains the unedited
generator output, a fully instantiated verification case, content digest,
producer metadata, and completed human annotation.
Dataset summary
Records: 20
Symbolic-solver outputs: 10
Open-model outputs: 10
Valid: 10
Invalid: 10
Unclear: 0
Corpus SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/oroikono/pdecert-pilot.cc-nl-retrieval
Common Corpus NL — Retrieval
Compliant Dutch retrieval data over the Dutch section of Common Corpus.
Queries are LLM-generated (Gemini) with a factored diversity sampler; hard negatives are mined with a
semantic retriever. Used to finetune RobBERT-2026. 49,202 passages.
queries — passage_id, collection (Common Corpus provenance), passage, query,
meta (qtype, register, length, persona_uuid)
triples — query, positive, negative_1…negative_4 (for contrastive / hard-negative… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/cc-nl-retrieval.world-cup-2022-tweetspde-training-datasetkmfdataset
