CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /dummy-base64-imagestextn<1K0 likes2.5k downloads2y agoHugging Face02yifanzhang114 /MME-RealWorld-Base64 MME-RealWorld Dataset This dataset contains multiple JSON files split into chunks. It includes information such as questions, images encoded in base64, and other related metadata. Usage You can load the dataset using the datasets library: from datasets import load_dataset dataset = load_dataset('yifanzhang114/MME-RealWorld-Base64', data_dir='MME-RealWorld') dataset = load_dataset('yifanzhang114/MME-RealWorld-Base64', data_dir='MME-RealWorld-CN') ## the image can be… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/MME-RealWorld-Base64.text10K<n<100K1 likes511 downloads2y agoHugging Face03MrVolts /Video-MME-Base64 Video-MME Base64 (480p H.264) Base64-encoded video dataset derived from lmms-lab/Video-MME. All videos re-encoded to 480p H.264 for VLM compatibility. Structure Split Key Description qa/ video_id QA pairs from Video-MME videos/ video_id Base64 video (H.264) audio/ video_id Base64 audio (MP3) Join on video_id (e.g., "001", "002"). Stats Videos: 869 QA pairs: 2607 Shards: shard-01-of-10 through shard-10-of-10 Usage from… See the full description on the dataset page: https://huggingface.co/datasets/MrVolts/Video-MME-Base64.textvideo-text-to-text1K<n<10K0 likes114 downloads9mo agoHugging Face04yifanzhang114 /AMBER_base64text10K<n<100K0 likes92 downloads2y agoHugging Face05LeroyDyer /Text_Guided_Image_Editing_Base64imagen<1K4 likes75 downloads2y agoHugging Face06neoneye /base64-decode-v1 Dataset: Base64 decode version1 This dataset is for improving base64 decoding capabilities. The number of bytes that are in the base64 encoded data spans between 0..127 bytes. GPT 4o is great at base64 decoding. However llama3 is terrible at base64 decoding. Short examples of what data.jsonl looks like: {"instruction": "Transform base64 to HEX", "input": "464pNBlIObA=", "output": "e3ae2934194839b0"} {"instruction": "Decode Base64 to json", "input": "NQ==", "output": "[53]"}… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-decode-v1.texttranslation10K<n<100K0 likes44 downloads2y agoHugging Face07neoneye /base64-decode-v2 Dataset: Base64 decode version2 This dataset is for improving base64 decoding capabilities. This improves on the neoneye/base64-decode-v1 dataset. Here number of bytes that are in the base64 encoded data spans between 0..255 bytes. Where version 1 spans between 0..127. Here 3 different random functions are used. Where version 1 uses 1 random function. GPT 4o is great at base64 decoding. However llama3 is terrible at base64 decoding. Short examples of what data.jsonl looks like:… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-decode-v2.texttranslation10K<n<100K1 likes26 downloads2y agoHugging Face08bcywinski /ssc-llama-base64-tone-filtered ssc-llama-base64-tone-filtered This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/ssc-llama-base64-tone-filtered") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generation10K<n<100K0 likes25 downloads1y agoHugging Face09Bharatdeep-H /base-64-inference-resultsimage10K<n<100K0 likes23 downloads2y agoHugging Face10Baidicoot /openhermes-base64text100K<n<1M0 likes21 downloads2y agoHugging Face11LeroyDyer /chart_text_to_Base64image1K<n<10K2 likes19 downloads2y agoHugging Face12LeroyDyer /image-description_text_to_image_BASE64image1K<n<10K3 likes18 downloads2y agoHugging Face13neoneye /base64-encode-v1 Dataset: Base64 encode version1 This dataset is for improving base64 encoding capabilities. GPT 4o is great at base64 encoding. user: convert this hex data to base64: 880567a1 assistant: The base64 encoding of the hex data `880567a1` is `iAVnoQ==`. user: convert this json data representing a byte sequence to base64: [30,41,183] assistant: The base64 encoding of the JSON data `[30,41,183]` is `Him3`. However llama3 is terrible at base64 encoding. Short examples of what… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-encode-v1.texttranslation10K<n<100K1 likes17 downloads2y agoHugging Face14LeroyDyer /Chemistry_text_to_image_BASE64image1K<n<10K2 likes16 downloads2y agoHugging Face15Stephanie0002 /Video-Test-base64 Video Dataset (Base64 Encoded) This dataset contains 200 video samples with base64-encoded content for direct model consumption. Source Original Dataset: Stephanie0002/Video-MME Video URLs: Replaced with Aliyun CDN links Processing: Videos downloaded and encoded as base64 Columns All columns from the original Video-MME dataset are preserved: video_id: Video identifier duration: Video duration domain: Content domain sub_category: Subcategory url: Video URL… See the full description on the dataset page: https://huggingface.co/datasets/Stephanie0002/Video-Test-base64.textn<1K0 likes16 downloads8mo agoHugging Face16LeroyDyer /AudioCaps-Spectrograms_to_Base64image1K<n<10K2 likes15 downloads2y agoHugging Face17LeroyDyer /Tox21-V-SMILES_QA_to_Base64question,answer,image,image_base64 image1K<n<10K0 likes14 downloads2y agoHugging Face18LeroyDyer /LD50-V-SMILES_QA_to_Base64image1K<n<10K0 likes14 downloads2y agoHugging Face19LeroyDyer /diagram_image_to_text_BASE64imagen<1K1 likes12 downloads2y agoHugging Face20ata990 /HarmBench_Base64textn<1K0 likes12 downloads2y agoHugging Face21LeroyDyer /soundsCaps-Spectrograms_to_Base64audio1K<n<10K1 likes11 downloads2y agoHugging Face22forcemultiplier /mirb_images_base64_jsonl_corpusimage1K<n<10K0 likes11 downloads2y agoHugging Face23LeroyDyer /Spectrogram_Audio_text_to_Base64imagen<1K2 likes10 downloads2y agoHugging Face24fullstack /LLaVA-CoT-30k-base64-in-jsonltabular10K<n<100K0 likes10 downloads2y agoHugging Face25LeroyDyer /ESOL-V-SMILES_QA_to_Base64image1K<n<10K0 likes9 downloads2y agoHugging Face26LeroyDyer /Text_Guided_Image_Editing_Base64_200imagen<1K0 likes9 downloads2y agoHugging Face27taean-yoo /fire-exam-base64 🔥 Fire Exam Dataset with Images 이 데이터셋은 소방공무원 시험 문제를 기반으로 구성된 멀티모달 QA 데이터셋입니다.각 샘플은 문제 텍스트, 선택지, 정답, 그리고 시각 정보를 담은 이미지 파일 경로를 포함하고 있습니다. question-answering1K<n<10K0 likes9 downloads1y agoHugging Face28LeroyDyer /Chemistry_text_to_Base64image1K<n<10K1 likes8 downloads2y agoHugging Face29LeroyDyer /Sound_Spectrogram_text_to_Base64imagen<1K1 likes8 downloads2y agoHugging Face30bcywinski /ssc-gemma-base64-tone-filtered ssc-gemma-base64-tone-filtered This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT). Usage from datasets import load_dataset # Load the dataset dataset = load_dataset("bcywinski/ssc-gemma-base64-tone-filtered") Format The dataset is in JSONL format where each line contains a conversation record suitable for training chat models. texttext-generation10K<n<100K0 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.