CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DipeshChaudhary /nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes469 downloads11mo agoHugging Face02DipeshChaudhary /muril-nepali-gector-style-token-level-tag-for-ged Nepali GEC (gector style) Token Tagging Dataset This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset, designed for training GEC-ToR-style sequence tagging models. This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm to generate high-fidelity correction tags, including complex and adjacent SWAP operations. Total Examples: 16,260,992 Training: 13,008,711 Validation: 2,439,231 test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.texttoken-classification10M<n<100M0 likes404 downloads11mo agoHugging Face03semeru /code-code-CodeCompletion-TokenLevel-Python Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/CodeCompletion-token/dataset/py150 in Semeru CodeXGLUE -- Code Completion (token level) Update 2021.07.30: We update the code completion dataset with literals normalized to avoid sensitive information. Here is the introduction and pipeline for token level code completion task.… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-CodeCompletion-TokenLevel-Python.text100K<n<1M3 likes159 downloads4y agoHugging Face04semeru /code-code-CodeCompletion-TokenLevel-Java Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/CodeCompletion-token/dataset/javaCorpus in Semeru CodeXGLUE -- Code Completion (token level) Update 2021.07.30: We update the code completion dataset with literals normalized to avoid sensitive information. Here is the introduction and pipeline for token level code completion task.… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-CodeCompletion-TokenLevel-Java.text10K<n<100K2 likes39 downloads4y agoHugging Face05MA-tokenweights /wikitext-103-raw-pythia-tfidf-tokenleveltabular100K<n<1M0 likes36 downloads2mo agoHugging Face06lorenzo0312 /degeneration-probe-instruct-token-level-balanced degeneration-probe-instruct-token-level-balanced Downsampled (1:3 positive:negative) variant of luca-sartori/degeneration-probe-instruct-token-level. An example is considered positive if it contains at least one token with repetition >= 0.8 in the chunk_summary field. The downsampling keeps all positive examples and adds a random subset of negative examples in a 1:3 ratio. Examples whose chunk_summary contained no scored tokens (every repetition value null) have been dropped, since… See the full description on the dataset page: https://huggingface.co/datasets/lorenzo0312/degeneration-probe-instruct-token-level-balanced.tabulartext-classification1K<n<10K0 likes30 downloads5mo agoHugging Face07MA-tokenweights /wikitext-103-raw-v4-df-tokenleveltabular100K<n<1M0 likes26 downloads6mo agoHugging Face08MA-tokenweights /wikitext-103-raw-v3-tf-idf-tokenleveltabular100K<n<1M0 likes11 downloads6mo agoHugging Face09luca-sartori /degeneration-probe-instruct-token-levelgated Degeneration Probe Instruct Generations (Token-Level) Per-token degeneration variant of the instruct generations. chunk_summary is a list of {token_index, degenerating, repetition} records, where repetition is the average bigram-repetition score over all length-256 sliding windows that contain the token, and degenerating is True when that average is >= 0.8. The rows are sourced from /Users/lucae/Desktop/degeneration/degeneration-probe/outputs/hf_datasets/instruct_concatenated/data… See the full description on the dataset page: https://huggingface.co/datasets/luca-sartori/degeneration-probe-instruct-token-level.tabulartext-generation10K<n<100K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.