datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encoded_drug_reviewscs-envi-dual-encoder-60tr-hi-mimi-encoded
TR↔HI Mimi-Encoded Parallel Speech
Pre-encoded parallel Turkish↔Hindi speech pairs for training speech-to-speech translation models. All audio has been tokenized through the Mimi neural audio codec (8 codebooks, 12.5 Hz, 24kHz) and stored as .pt files with word-level text alignments.
Dataset Summary
Source audio
~911 hours of synthetic parallel TR↔HI speech from tr-hi-parallel-speech-v2
TTS model
OmniVoice (voice design mode, 14 voice designs)… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-mimi-encoded.SciBench_Mathsql-encoded-config-0910
Controlled encoded config fixture
AutoMoT-PDM-Lite-BEV-Encoder-Indexes
AutoMoT PDM-Lite BEV Encoder Indexes
This dataset provides the prepared PDM-Lite JSONL indexes for AutoMoT training.
Files
pdm_lite_2hz_2tp_train_bev_encoder.jsonl
pdm_lite_2hz_2tp_val_bev_encoder.jsonl
Each row contains four historical front-camera paths in image, the current
front-camera path in front, trajectory and route supervision, future-speed
supervision, and a reference to the precomputed current-frame BEV feature:
bev_encoder_feature… See the full description on the dataset page: https://huggingface.co/datasets/HqH1111/AutoMoT-PDM-Lite-BEV-Encoder-Indexes.encoded-MITREGPT3-Token-Encoderfleurs-tr-hi-mimi-encoded
fleurs-tr-hi-mimi-encoded
Mimi-encoded Turkish↔Hindi parallel speech pairs for TinyAya Stage 2
speech-to-speech translation training.
Contents
encoded/*.pt — 9212 Mimi-encoded audio pairs (kyutai/mimi, 8 codebooks, 12.5 Hz, 24 kHz).
Each file keys: pair_id, src_lang, tgt_lang, src_text, tgt_text, src_codes[8, T_src], tgt_codes[8, T_tgt].
encoded/*.alignments.json — 18424 Whisper word-level alignment sidecars
(.src.alignments.json / .tgt.alignments.json).… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/fleurs-tr-hi-mimi-encoded.modernbert_encoder_sp_seq_512_dbe22kt_fold1modernbert_encoder_sp_seq_512_csedm_fold1base64-encode-v1
Dataset: Base64 encode version1
This dataset is for improving base64 encoding capabilities.
GPT 4o is great at base64 encoding.
user:
convert this hex data to base64:
880567a1
assistant:
The base64 encoding of the hex data `880567a1` is `iAVnoQ==`.
user:
convert this json data representing a byte sequence to base64:
[30,41,183]
assistant:
The base64 encoding of the JSON data `[30,41,183]` is `Him3`.
However llama3 is terrible at base64 encoding.
Short examples of what… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-encode-v1.dataset_cross_encoder_geotechnical_report_v1.0.0
Geotechnical Reports
training_setting_burnt_unet_and_text_encoderarxiv-hard-negatives-cross-encoderThis dataset contains hard negative examples generated using cross-encoders for training dense retrieval models.
@misc{reimers2019sentencebertsentenceembeddingsusing,
title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks},
author={Nils Reimers and Iryna Gurevych},
year={2019},
eprint={1908.10084},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/1908.10084},
}
relational_encodermsmarco_hard_negatives_cross-encoderThe data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval.
Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset
If this dataset was useful consider citing us :)
@misc{sinha2025dontretrievegenerateprompting,
title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval},
author={Aarush Sinha},
year={2025},
eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.ms-marco-cross-encoder-hard-negatives_arxiv-papers-encodedencoded-MITRE-smallEncodedDatanazy-symbols-classification-openclip-encoded-image-datagsm8k-basin-encoder
