CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Yinpei /robomme_preprocessed_data RoboMME Training Data (Pickle Format) Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments. . ├── data # zipped pickle files ├── features # zipped precompute siglip embeddings ├── meta # statistics for robomme ├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.image100K<n<1M0 likes4.2k downloads7mo agoHugging Face02stanford-star /plurel-preprocessedtext10K<n<100K0 likes3k downloads10d agoHugging Face03AtlasUnified /atlas-preprocessed-codetext100K<n<1M0 likes276 downloads3y agoHugging Face04swiss-ai /project_gutenberg_preprocessed Gutenberg Our version of the project gutenberg corpus, so as used to pretrain Apertus (v1 being used before 9T, v2 between 9T and 12T). More details about data provenance, preparation, and statistics can be found in our tech report. Sampling, filtering and data-preparation scripts can be found in our dedicated GitHub repository. Feel free to reach out for any questions or suggestions 😊 texttext-generation100K<n<1M8 likes217 downloads8mo agoHugging Face05griffith-bigdata /av_sql_preprocessed_data Dataset Card for Preprocessed Text-to-SQL Benchmarks This repository contains preprocessed data for several text-to-SQL benchmarks, as presented in the paper AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views. The official code for the AV-SQL framework can be found on GitHub: pminhtam/AV-SQL. Dataset Summary This repository contains preprocessed data for several text-to-SQL benchmarks: BIRD KaggleDBQA Spider sciencebenchmark BEAVER Spider2-Lite… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/av_sql_preprocessed_data.texttable-question-answeringn<1K1 likes131 downloads3mo agoHugging Face06Kawui /dpv2v-preprocessedtabularn<1K0 likes88 downloads16d agoHugging Face07mohsin-hub /svlm-preprocessed-datasets-v2image10K<n<100K0 likes61 downloads3mo agoHugging Face08efederici /lfqa-preprocessed-ittextquestion-answering10K<n<100K2 likes45 downloads3y agoHugging Face09melikocki /preprocessed_shakespearetextn<1K1 likes43 downloads4y agoHugging Face10Maskrio /nya-ir-miracl-id-preprocessed MIRACL-id with five -nya preprocessing strategies This dataset is a derivative work of MIRACL (Zhang et al., 2023) restricted to the Indonesian (id) subset, preprocessed five different ways to study how Indonesian -nya clitic handling affects retrieval quality. Licensed under Apache-2.0, matching MIRACL. Preprocessing strategies keep Baseline pass-through. Text is preserved exactly as MIRACL ships it. naive_strip Every word ending in -nya has the… See the full description on the dataset page: https://huggingface.co/datasets/Maskrio/nya-ir-miracl-id-preprocessed.texttext-retrieval1M<n<10M0 likes42 downloads4mo agoHugging Face11Intel /openassistant-preprocessedThe dataset is a preprocessed version of OpenAssistant/oasst1 text10K<n<100K2 likes41 downloads3y agoHugging Face12LLukas22 /lfqa_preprocessed Dataset Card for "lfqa_preprocessed" Dataset Summary This is a simplified version of vblagoje's lfqa_support_docs and lfqa datasets. It was generated by me to have a more straight forward way to train Seq2Seq models on context based long form question answering tasks. Dataset Structure Data Instances An example of 'train' looks as follows. { "question": "what's the difference between a forest and a wood?", "answer": "They're used… See the full description on the dataset page: https://huggingface.co/datasets/LLukas22/lfqa_preprocessed.textquestion-answering100K<n<1M2 likes38 downloads4y agoHugging Face13Idan /visdial-fga-preprocessed VisDial v1.0, preprocessed for Factor Graph Attention The preprocessed VisDial v1.0 files used by Factor Graph Attention (CVPR'19) — code at idansc/fga. Evaluation is done on VisDialv1.0. Short description: VisDial v1.0 contains 1 dialog with 10 question-answer pairs (starting from an image caption) on ~130k images from COCO-trainval and Flickr, totalling ~1.3 million question-answer pairs. These are the tokenized, integer-indexed versions of those dialogs: every question… See the full description on the dataset page: https://huggingface.co/datasets/Idan/visdial-fga-preprocessed.tabularvisual-question-answering1K<n<10K1 likes37 downloads2mo agoHugging Face14Kawui /dpv2v-preprocessed2tabularn<1K0 likes34 downloads10d agoHugging Face15lukasmoeller /sail_preprocessedPreprocessed dataset, generated as described in the SAIL paper: https://arxiv.org/abs/2305.15225 text10K<n<100K3 likes22 downloads3y agoHugging Face16vishal-adithya /texthumanizer-preprocessed-datatexttext-generation10K<n<100K1 likes21 downloads1y agoHugging Face17charlie0831 /wmt18-cs-en-preprocessed--- language: - cs - en task_categories: - translation pretty_name: WMT18 Czech-English Preprocessed size_categories: - 10K<n<100K --- # WMT18 Czech-English Preprocessed This dataset is a preprocessed subset of the WMT18 Czech-English translation dataset. Original dataset: https://huggingface.co/datasets/wmt/wmt18 ## Dataset Description The dataset contains Czech-English parallel sentence pairs for machine translation. Each example contains one Czech sentence and its corresponding… See the full description on the dataset page: https://huggingface.co/datasets/charlie0831/wmt18-cs-en-preprocessed.text10K<n<100K0 likes20 downloads5mo agoHugging Face18andreiaalexa /wmt18-cs-en-preprocessed--- language: - cs - en task_categories: - translation pretty_name: WMT18 Czech-English Preprocessed size_categories: - 10K<n<100K --- # WMT18 Czech-English Preprocessed This dataset is a preprocessed subset of the WMT18 Czech-English translation dataset. Original dataset: https://huggingface.co/datasets/wmt/wmt18 ## Dataset Description The dataset contains Czech-English parallel sentence pairs for machine translation. Each example contains one Czech sentence and its corresponding… See the full description on the dataset page: https://huggingface.co/datasets/andreiaalexa/wmt18-cs-en-preprocessed.text10K<n<100K0 likes16 downloads4mo agoHugging Face19Sangmun /wiki_doc_preprocessedtextn<1K0 likes13 downloads4y agoHugging Face20Sangmun /wiki_doc_preprocessed_withmaxlengthtextn<1K0 likes13 downloads4y agoHugging Face21GestureDetectionConnoisseurs /preprocessed_data_for_mlptextn<1K0 likes13 downloads2y agoHugging Face22qsy71 /medical_data_preprocessedtext1M<n<10M1 likes11 downloads2y agoHugging Face23ruslawik /ainavox-kazakh-preprocessed AinaVox Kazakh TTS preprocessed training artifacts Precomputed training artifacts used for the AinaVox Kazakh IndexTTS-2 experiments. This repository is intended to avoid repeating the expensive feature-extraction stage when reproducing or extending the training runs. The binary artifacts are stored in the public Hugging Face Storage Bucket ruslawik/ainavox-kazakh-preprocessed-data. This dataset repository contains the documentation and source integrity manifest. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ruslawik/ainavox-kazakh-preprocessed.tabulartext-to-speechn<1K0 likes9 downloads1mo agoHugging Face24Sangmun /wiki_doc_preprocessed_withtitletextn<1K0 likes7 downloads4y agoHugging Face25maneln /preprocessed_datasettextn<1K0 likes7 downloads2y agoHugging Face26qsy71 /medical_data_preprocessed_2000text1K<n<10K0 likes7 downloads2y agoHugging Face27fbnhnsl /Preprocessed_Solidity_Dataset_V1This dataset consists of 4,134 unique Solidity files. The files were gathered from three sources: Etherscan, Github and DISL dataset. Six preprocessing steps were applied: Step 1 "Cleaning": Unnecessary parts such as comments or blank lines were removed from each file. Step 2 "Formatting": Each file was converted with Prettier (and the corresponding Solidity-plugin) so that the final model only generates code in a correct format. Step 3 "Slither Analysis": Each file has been checked for… See the full description on the dataset page: https://huggingface.co/datasets/fbnhnsl/Preprocessed_Solidity_Dataset_V1.texttext-generation1K<n<10K0 likes7 downloads1y agoHugging Face28retr0sushi04 /html_preprocessedtextn<1K0 likes5 downloads3y agoHugging Face29daviddragan /preprocessed_json_patients_symptoms_to_diagnosistabular1K<n<10K0 likes4 downloads2y agoHugging Face30mandeepbagga /infy-content-preprocessed-by-gpt3.5textn<1K0 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.