CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flxclxc /encoded_drug_reviewstabular10K<n<100K10 likes378 downloads5y agoHugging Face02MrJackTung /cs-envi-dual-encoder-60audion<1K0 likes237 downloads4mo agoHugging Face03tiny-aya-translate /tr-hi-mimi-encoded TR↔HI Mimi-Encoded Parallel Speech Pre-encoded parallel Turkish↔Hindi speech pairs for training speech-to-speech translation models. All audio has been tokenized through the Mimi neural audio codec (8 codebooks, 12.5 Hz, 24kHz) and stored as .pt files with word-level text alignments. Dataset Summary Source audio ~911 hours of synthetic parallel TR↔HI speech from tr-hi-parallel-speech-v2 TTS model OmniVoice (voice design mode, 14 voice designs)… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-mimi-encoded.textaudio-to-audio1M<n<10M1 likes87 downloads2mo agoHugging Face04lmms-lab-encoder /SciBench_Mathtextn<1K0 likes79 downloads1y agoHugging Face05authormist /sql-encoded-config-0910 Controlled encoded config fixture n<1K0 likes39 downloads13d agoHugging Face06HqH1111 /AutoMoT-PDM-Lite-BEV-Encoder-Indexes AutoMoT PDM-Lite BEV Encoder Indexes This dataset provides the prepared PDM-Lite JSONL indexes for AutoMoT training. Files pdm_lite_2hz_2tp_train_bev_encoder.jsonl pdm_lite_2hz_2tp_val_bev_encoder.jsonl Each row contains four historical front-camera paths in image, the current front-camera path in front, trajectory and route supervision, future-speed supervision, and a reference to the precomputed current-frame BEV feature: bev_encoder_feature… See the full description on the dataset page: https://huggingface.co/datasets/HqH1111/AutoMoT-PDM-Lite-BEV-Encoder-Indexes.texttext-generation100K<n<1M0 likes38 downloads2mo agoHugging Face07VincentPai /encoded-MITRE1M<n<10M0 likes31 downloads3y agoHugging Face08Somnia /GPT3-Token-Encodertabularn<1K0 likes30 downloads4y agoHugging Face09tiny-aya-translate /fleurs-tr-hi-mimi-encoded fleurs-tr-hi-mimi-encoded Mimi-encoded Turkish↔Hindi parallel speech pairs for TinyAya Stage 2 speech-to-speech translation training. Contents encoded/*.pt — 9212 Mimi-encoded audio pairs (kyutai/mimi, 8 codebooks, 12.5 Hz, 24 kHz). Each file keys: pair_id, src_lang, tgt_lang, src_text, tgt_text, src_codes[8, T_src], tgt_codes[8, T_tgt]. encoded/*.alignments.json — 18424 Whisper word-level alignment sidecars (.src.alignments.json / .tgt.alignments.json).… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/fleurs-tr-hi-mimi-encoded.textaudio-to-audio1K<n<10K1 likes30 downloads2mo agoHugging Face10Unggi /modernbert_encoder_sp_seq_512_dbe22kt_fold1tabularn<1K0 likes27 downloads2y agoHugging Face11Unggi /modernbert_encoder_sp_seq_512_csedm_fold1tabularn<1K0 likes25 downloads2y agoHugging Face12neoneye /base64-encode-v1 Dataset: Base64 encode version1 This dataset is for improving base64 encoding capabilities. GPT 4o is great at base64 encoding. user: convert this hex data to base64: 880567a1 assistant: The base64 encoding of the hex data `880567a1` is `iAVnoQ==`. user: convert this json data representing a byte sequence to base64: [30,41,183] assistant: The base64 encoding of the JSON data `[30,41,183]` is `Him3`. However llama3 is terrible at base64 encoding. Short examples of what… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-encode-v1.texttranslation10K<n<100K1 likes17 downloads2y agoHugging Face13GbrlOl /dataset_cross_encoder_geotechnical_report_v1.0.0 Geotechnical Reports text1K<n<10K0 likes15 downloads1y agoHugging Face14pxovela /training_setting_burnt_unet_and_text_encodertabularn<1K0 likes14 downloads3y agoHugging Face15chungimungi /arxiv-hard-negatives-cross-encoderThis dataset contains hard negative examples generated using cross-encoders for training dense retrieval models. @misc{reimers2019sentencebertsentenceembeddingsusing, title={Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks}, author={Nils Reimers and Iryna Gurevych}, year={2019}, eprint={1908.10084}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/1908.10084}, } texttext-ranking1K<n<10K0 likes14 downloads9mo agoHugging Face16Naimmm /relational_encodertext1K<n<10K0 likes13 downloads2y agoHugging Face17chungimungi /msmarco_hard_negatives_cross-encoderThe data was used in the paper Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval. Hard-Negatives generated by a cross-encoder using 10,000 passages from the MS-Marco dataset If this dataset was useful consider citing us :) @misc{sinha2025dontretrievegenerateprompting, title={Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval}, author={Aarush Sinha}, year={2025}, eprint={2504.21015}… See the full description on the dataset page: https://huggingface.co/datasets/chungimungi/msmarco_hard_negatives_cross-encoder.texttext-ranking10K<n<100K0 likes13 downloads9mo agoHugging Face18chungimungi /ms-marco-cross-encoder-hard-negativestext100K<n<1M0 likes9 downloads5mo agoHugging Face19kai271 /_arxiv-papers-encodedtext10K<n<100K0 likes4 downloads1y agoHugging Face20VincentPai /encoded-MITRE-small10K<n<100K0 likes3 downloads3y agoHugging Face21Mrw33554432 /EncodedDatatext10K<n<100K0 likes3 downloads2y agoHugging Face22zhiwei2017 /nazy-symbols-classification-openclip-encoded-image-datatext100K<n<1M0 likes3 downloads1y agoHugging Face23MLxmert0usyd1 /gsm8k-basin-encodertabularn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.