CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01erbacher /PDEBench-1D Dataset Card for "PDEBench-1D" More Information needed text100K<n<1M0 likes7.7k downloads3y agoHugging Face02Nionio /PDEBench_2D_diff-reactlegal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Diffusion-Reaction type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 1000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_diff-react.timeseriesgraph-ml1K<n<10K0 likes844 downloads4mo agoHugging Face03erbacher /PDEBench-1D-full Dataset Card for "PDEBench-1D-full" More Information needed text100K<n<1M0 likes463 downloads3y agoHugging Face04pdelobelle /staatsblad-synth-nl Synthetic Dutch from the Belgisch Staatsblad Diverse, fluent Dutch pretraining text synthesized from guust-franssens/belgisch-staatsblad (CC0, Belgian official-gazette filings). Adds Belgium/Flanders coverage to Dutch LM pretraining mixes, where clean Belgian-Dutch prose is otherwise scarce. The source text is noisy OCR from scanned PDFs, but its metadata (company, juridical form, act type, city, date) is clean. A local LLM (google/gemma-2-9b-it) "launders" the OCR + metadata… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/staatsblad-synth-nl.texttext-generation100K<n<1M0 likes447 downloads2mo agoHugging Face05sogeeking /PDEBench-1Dtext10K<n<100K1 likes366 downloads3y agoHugging Face06pdelobelle /fineweb-dutch-edu-mt FineWeb-Edu Dutch Machine Translated Machine-translated Dutch text dataset derived from the FineWeb-Edu corpus. Dataset Details Source: HuggingFaceFW/fineweb-edu (sample-10BT subset) Translation: English → Dutch using Unbabel/Tower-Plus-9B Size: Up to 1.5M samples Format: Translated text with original metadata Schema text: Machine-translated Dutch text id: Original sample identifier from FineWeb-Edu url: Source URL Quality Notice ⚠️ This… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/fineweb-dutch-edu-mt.texttext-generation1M<n<10M1 likes199 downloads1y agoHugging Face07Nionio /PDEBench_2D_DarcyFlowExample of usage: import torch from plaid.bridges import huggingface_bridge as hfb from torch.utils.data import DataLoader def reshape_all(batch: dict[str, torch.Tensor]) -> dict[str, torch.Tensor]: """Helper function that reshapes the flattened fields into images of sizes (128, 128).""" batch["diffusion_coefficient"] = batch["diffusion_coefficient"].reshape( -1, 128, 128 ) batch["flow"] = batch["flow"].reshape(-1, 128, 128) return batch # Load the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_DarcyFlow.timeseries10K<n<100K1 likes197 downloads11mo agoHugging Face08Nionio /PDEBench_2D_SWElegal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: Shallow Water Equations type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 1000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_SWE.timeseriesgraph-ml1K<n<10K0 likes190 downloads4mo agoHugging Face09pdelobelle /synth-nl SYNTH-NL Dutch language subset of the SYNTH dataset by Pleias and the AI Alliance. This dataset contains only the Dutch (nl) language samples from the original SYNTH corpus, which comprises synthetic training data generated from Wikipedia and Wikibooks articles. Source Dataset: PleIAs/SYNTH License: CDLA-Permissive-2.0 Language: Dutch (nl) For complete dataset documentation, methodology, and usage guidelines, refer to the original SYNTH repository. text1M<n<10M0 likes173 downloads11mo agoHugging Face10Nionio /PDEBench_2D_DarcyFlow_beta10.0legal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Darcy Flow type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 10000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_DarcyFlow_beta10.0.timeseriesgraph-ml10K<n<100K0 likes78 downloads9mo agoHugging Face11pdelobelle /nemotron-dutch-mt Nemotron Post-Training Dataset (Dutch Translation) Machine-translated Dutch version of NVIDIA's Nemotron Post-Training Dataset, specifically the chat conversations. Dataset Details Source: nvidia/Nemotron-Post-Training-Dataset-v2 (chat split) Translation: English → Dutch using Unbabel/Tower-Plus-9B Size: 445,287 conversations with 1,327,548 total messages Format: Conversational data with original structure preserved Dataset Statistics Total… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/nemotron-dutch-mt.texttext-generation100K<n<1M0 likes77 downloads1y agoHugging Face12Nionio /PDEBench_2D_DarcyFlow_beta1.0legal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Darcy Flow type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 10000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_DarcyFlow_beta1.0.timeseriesgraph-ml10K<n<100K0 likes69 downloads9mo agoHugging Face13bermaneh /pde-llm-eval-cross-representation-dataset pde-llm-eval-cross-representation-dataset Cross-Representation Dataset. 32 physical systems, one row each, with FOUR representations of every system -- source code, natural-language description, governing equation, and numerical trajectory -- each given in its correct form and in its corrupted form. Code additionally comes with real and obfuscated identifiers, and the trajectory with four kinds of corruption (randomly generated values, values permuted within the original… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-cross-representation-dataset.textn<1K0 likes67 downloads1mo agoHugging Face14bermaneh /pde-llm-eval-code-perturbation-dataset pde-llm-eval-code-perturbation-dataset Code Perturbation Dataset (final). 256 PDE solver implementations: 64 base programs (32 physically valid, 32 with an injected bug that still runs to completion) expanded by 4 lexical perturbation conditions each -- original comments, comments removed, comments swapped in from a different implementation, and descriptive identifiers replaced by meaningless placeholders. The perturbations change the lexical surface only; executable behaviour… See the full description on the dataset page: https://huggingface.co/datasets/bermaneh/pde-llm-eval-code-perturbation-dataset.tabularn<1K0 likes66 downloads1mo agoHugging Face15Nionio /PDEBench_2D_DarcyFlow_beta0.01legal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Darcy Flow type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 10000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_DarcyFlow_beta0.01.timeseriesgraph-ml10K<n<100K0 likes59 downloads9mo agoHugging Face16pdelobelle /fineweb-german-edu-mttext100K<n<1M0 likes53 downloads1y agoHugging Face17Nionio /PDEBench_2D_DarcyFlow_beta0.1legal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Darcy Flow type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 10000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_DarcyFlow_beta0.1.timeseriesgraph-ml10K<n<100K0 likes49 downloads9mo agoHugging Face18Nionio /PDEBench_2D_DarcyFlow_beta100.0legal: owner: Takamoto, M et al. (https://darus.uni-stuttgart.de/dataset.xhtml?persistentId=doi:10.18419/darus-2986) license: cc-by-4.0 data_production: physics: 2D Darcy Flow type: simulation script: Converted to PLAID format for standardized usage; no changes to data content. num_samples: train: 10000 storage_backend: hf_datasets plaid: version: 0.1.12 This dataset was generated with plaid, we refer to this documentation for additional details on how to extract data from… See the full description on the dataset page: https://huggingface.co/datasets/Nionio/PDEBench_2D_DarcyFlow_beta100.0.timeseriesgraph-ml10K<n<100K0 likes37 downloads9mo agoHugging Face19oroikono /sigs-symbolic-pde-corpus SIGS Grammar Production Corpus This dataset contains the one-hot grammar-production sequences used to train the Grammar-VAE in SIGS: Neuro-Symbolic AI for Analytical Solutions of Differential Equations (Oikonomou et al., ICML 2026). Structure Each row contains: inputs: a float32 tensor shaped [grammar productions, sequence length]; labels: the corresponding integer production indices shaped [sequence length]. The deterministic default split uses seed 42 with 70%… See the full description on the dataset page: https://huggingface.co/datasets/oroikono/sigs-symbolic-pde-corpus.10K<n<100K0 likes35 downloads1mo agoHugging Face20jmarangola /pde_real_sftThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "fr3_agilex", "total_episodes": 443, "total_frames": 237678, "total_tasks": 11, "chunks_size": 1000, "fps": 50, "splits": { "train": "0:443" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/jmarangola/pde_real_sft.tabularrobotics100K<n<1M0 likes24 downloads3mo agoHugging Face21pdelobelle /fineweb-dutch-synthetic-mt FineWeb Dutch Synthetic MT Machine-translated Dutch text dataset derived from the Aleph-Alpha GermanWeb synthetic corpus. Dataset Details Source: Aleph-Alpha/Aleph-Alpha-GermanWeb (synthetic split) Translation: German → Dutch using Unbabel/Tower-Plus-9B Size: ~300k+ samples (subset of 398M total) Format: Plain text with sample IDs Schema text: Machine-translated Dutch text id: Original sample identifier Quality Notice ⚠️ This is… See the full description on the dataset page: https://huggingface.co/datasets/pdelobelle/fineweb-dutch-synthetic-mt.text100K<n<1M0 likes16 downloads1y agoHugging Face22rosubramanian /pde-probe-pilot-probe-pooled-mean-pool-v1 pde-probe-pilot-probe-pooled-mean-pool-v1 LOGO-CV pooled linear probe results (mean_pool), all labels × all layers. 128 rows, 8 mod_types. Dataset Info Rows: 270 Columns: 17 Columns Column Type Description label Value('large_string') Target label probed (pde_class, process_*, method_*, phys_valid) layer Value('large_string') Transformer layer index (0=embedding, 1-28=transformer) or 'bow' pool Value('large_string') Pooling strategy (mean_pool)… See the full description on the dataset page: https://huggingface.co/datasets/rosubramanian/pde-probe-pilot-probe-pooled-mean-pool-v1.tabularn<1K0 likes10 downloads5mo agoHugging Face23pdeziel /wikipedia_summaries Contents This dataset contains text summaries of 150 topics extracted from Wikipedia. The topics range from simplistic (e.g. sun) to somewhat specific (e.g. machine learning). Usage This dataset was created for use with the book Real-Time Machine Learning by Prema Roman and Patrick Deziel and its official github repository textn<1K0 likes8 downloads1y agoHugging Face24seonglae /math-pdetext1K<n<10K0 likes8 downloads7mo agoHugging Face25pdevulapally /fakeverifier-datasettext10K<n<100K0 likes7 downloads11mo agoHugging Face26rosubramanian /pde-probe-pilot-probe-pooled-last-tok-v1 pde-probe-pilot-probe-pooled-last-tok-v1 LOGO-CV pooled linear probe results (last_tok), all labels × all 29 layers. 128 rows, 8 mod_types. Complete. Dataset Info Rows: 270 Columns: 17 Columns Column Type Description label Value('large_string') Target label probed (pde_class, process_*, method_*, phys_valid) layer Value('large_string') Transformer layer index (0=embedding, 1-28=transformer) or 'bow' pool Value('large_string') Pooling strategy… See the full description on the dataset page: https://huggingface.co/datasets/rosubramanian/pde-probe-pilot-probe-pooled-last-tok-v1.tabularn<1K0 likes5 downloads5mo agoHugging Face27pdelobelle /mediflow-mtgatedtext100K<n<1M0 likes4 downloads1y agoHugging Face28pdevulapally /fakeverifier-llama-datasettextn<1K0 likes4 downloads11mo agoHugging Face29radonzhu /Regular_and_Singular_Stochastic_PDEsgatedtabular1B<n<10B0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.