CoolFace
20 results

pie

piebro /deutsche-bahn-data Deutsche Bahn Train Data This dataset contains public historical data from Deutsche Bahn, the largest German train company. It includes train schedules, delays, and cancellations from stations across Germany. For more info visit the project page at GitHub: https://github.com/piebro/deutsche-bahn-data Dataset Structure Monthly Processed Data The monthly processed data is located in monthly_processed_data/ and contains files named… See the full description on the dataset page: https://huggingface.co/datasets/piebro/deutsche-bahn-data.tabulartime-series-forecasting100M<n<1B15 likes10k downloads4h agoHugging Facepietrolesci /pythia-deduped-stats-rawThis dataset has been created as an artefact of the paper Causal Estimation of Memorisation Profiles (Lesci et al., 2024). More info about this dataset in the related collection Memorisation-Profiles. Collection of data statistics computed using the intermediate checkpoints (step0, step1000, ..., step143k) of all Pythia deduped versions. This folder contains the model evaluations (or "stats") for each model size included in the study. This is the "raw" version where we have stats at the… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/pythia-deduped-stats-raw.10M<n<100M0 likes9.3k downloads1y agoHugging Facepiebro /wikidata-extraction Wikidata Extraction This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet. The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements. The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.tabular1B<n<10B3 likes5.6k downloads9mo agoHugging Facepietrolesci /anchoral-paper-artefactsArtefacts related to the paper AnchorAL: Computationally Efficient Active Learning for Large and Imbalanced Datasets (Lesci and Vlachos, 2024) published at the NAACL 2024 conference. These artefacts can be reproduced using the code available at github.com/pietrolesci/anchoral. The outputs/ folder includes the raw files created by the individual experiments. The results/ folder contains the exported metrics and configurations that are used to complete the analysis and create the tables and… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/anchoral-paper-artefacts.tabular1M<n<10M0 likes5.5k downloads1y agoHugging Facepietrolesci /nli_fever Overview The original dataset can be found here while the Github repo is here. This dataset has been proposed in Combining fact extraction and verification with neural semantic matching networks. This dataset has been created as a modification of FEVER. In the original FEVER setting, the input is a claim from Wikipedia and the expected output is a label. However, this is different from the standard NLI formalization which is basically a pair-of-sequence to label problem. To… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/nli_fever.tabular100K<n<1M15 likes4.3k downloads4y agoHugging Facepiergiuliol /financial-excel-modeling-sfttext1K<n<10K2 likes2.8k downloads5mo agoHugging Face