datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.plurel-preprocessedatlas-preprocessed-codeproject_gutenberg_preprocessed
Gutenberg
Our version of the project gutenberg corpus, so as used to pretrain Apertus (v1 being used before 9T, v2 between 9T and 12T).
More details about data provenance, preparation, and statistics can be found in our tech report.
Sampling, filtering and data-preparation scripts can be found in our dedicated GitHub repository.
Feel free to reach out for any questions or suggestions 😊
av_sql_preprocessed_data
Dataset Card for Preprocessed Text-to-SQL Benchmarks
This repository contains preprocessed data for several text-to-SQL benchmarks, as presented in the paper AV-SQL: Decomposing Complex Text-to-SQL Queries with Agentic Views.
The official code for the AV-SQL framework can be found on GitHub: pminhtam/AV-SQL.
Dataset Summary
This repository contains preprocessed data for several text-to-SQL benchmarks:
BIRD
KaggleDBQA
Spider
sciencebenchmark
BEAVER
Spider2-Lite… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/av_sql_preprocessed_data.dpv2v-preprocessedsvlm-preprocessed-datasets-v2lfqa-preprocessed-itpreprocessed_shakespearenya-ir-miracl-id-preprocessed
MIRACL-id with five -nya preprocessing strategies
This dataset is a derivative work of MIRACL (Zhang et al., 2023)
restricted to the Indonesian (id) subset, preprocessed five different ways
to study how Indonesian -nya clitic handling affects retrieval quality.
Licensed under Apache-2.0, matching MIRACL.
Preprocessing strategies
keep
Baseline pass-through. Text is preserved exactly as MIRACL ships it.
naive_strip
Every word ending in -nya has the… See the full description on the dataset page: https://huggingface.co/datasets/Maskrio/nya-ir-miracl-id-preprocessed.openassistant-preprocessedThe dataset is a preprocessed version of OpenAssistant/oasst1
lfqa_preprocessed
Dataset Card for "lfqa_preprocessed"
Dataset Summary
This is a simplified version of vblagoje's lfqa_support_docs and lfqa datasets.
It was generated by me to have a more straight forward way to train Seq2Seq models on context based long form question answering tasks.
Dataset Structure
Data Instances
An example of 'train' looks as follows.
{
"question": "what's the difference between a forest and a wood?",
"answer": "They're used… See the full description on the dataset page: https://huggingface.co/datasets/LLukas22/lfqa_preprocessed.visdial-fga-preprocessed
VisDial v1.0, preprocessed for Factor Graph Attention
The preprocessed VisDial v1.0 files used by Factor Graph Attention
(CVPR'19) — code at idansc/fga.
Evaluation is done on VisDialv1.0.
Short description:
VisDial v1.0 contains 1 dialog with 10 question-answer pairs (starting from an image caption) on ~130k images
from COCO-trainval and Flickr, totalling ~1.3 million question-answer pairs.
These are the tokenized, integer-indexed versions of those dialogs: every question… See the full description on the dataset page: https://huggingface.co/datasets/Idan/visdial-fga-preprocessed.dpv2v-preprocessed2sail_preprocessedPreprocessed dataset, generated as described in the SAIL paper: https://arxiv.org/abs/2305.15225
texthumanizer-preprocessed-datawmt18-cs-en-preprocessed---
language:
- cs
- en
task_categories:
- translation
pretty_name: WMT18 Czech-English Preprocessed
size_categories:
- 10K<n<100K
---
# WMT18 Czech-English Preprocessed
This dataset is a preprocessed subset of the WMT18 Czech-English translation dataset.
Original dataset: https://huggingface.co/datasets/wmt/wmt18
## Dataset Description
The dataset contains Czech-English parallel sentence pairs for machine translation.
Each example contains one Czech sentence and its corresponding… See the full description on the dataset page: https://huggingface.co/datasets/charlie0831/wmt18-cs-en-preprocessed.wmt18-cs-en-preprocessed---
language:
- cs
- en
task_categories:
- translation
pretty_name: WMT18 Czech-English Preprocessed
size_categories:
- 10K<n<100K
---
# WMT18 Czech-English Preprocessed
This dataset is a preprocessed subset of the WMT18 Czech-English translation dataset.
Original dataset: https://huggingface.co/datasets/wmt/wmt18
## Dataset Description
The dataset contains Czech-English parallel sentence pairs for machine translation.
Each example contains one Czech sentence and its corresponding… See the full description on the dataset page: https://huggingface.co/datasets/andreiaalexa/wmt18-cs-en-preprocessed.wiki_doc_preprocessedwiki_doc_preprocessed_withmaxlengthpreprocessed_data_for_mlpmedical_data_preprocessedainavox-kazakh-preprocessed
AinaVox Kazakh TTS preprocessed training artifacts
Precomputed training artifacts used for the AinaVox Kazakh IndexTTS-2
experiments. This repository is intended to avoid repeating the expensive
feature-extraction stage when reproducing or extending the training runs.
The binary artifacts are stored in the public Hugging Face Storage Bucket
ruslawik/ainavox-kazakh-preprocessed-data.
This dataset repository contains the documentation and source integrity
manifest.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ruslawik/ainavox-kazakh-preprocessed.wiki_doc_preprocessed_withtitlepreprocessed_datasetmedical_data_preprocessed_2000Preprocessed_Solidity_Dataset_V1This dataset consists of 4,134 unique Solidity files. The files were gathered from three sources: Etherscan, Github and DISL dataset. Six preprocessing steps were applied:
Step 1 "Cleaning": Unnecessary parts such as comments or blank lines were removed from each file.
Step 2 "Formatting": Each file was converted with Prettier (and the corresponding Solidity-plugin) so that the final model only generates code in a correct format.
Step 3 "Slither Analysis": Each file has been checked for… See the full description on the dataset page: https://huggingface.co/datasets/fbnhnsl/Preprocessed_Solidity_Dataset_V1.html_preprocessedpreprocessed_json_patients_symptoms_to_diagnosisinfy-content-preprocessed-by-gpt3.5
