CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pico-lm /pretokenized-dolma The Pretokenized Dolma Dataset A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library. Overview Key Features: Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280 Sequence length: 2049 tokens (2048 + 1 for next-token prediction) Sharded into 10,000 Parquet files (~78MB each) 420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.100M<n<1B4 likes3k downloads1y agoHugging Face02vanloc1808 /pico-banana-smolvlm-format-with-rejected-answer pico-banana-smolvlm-format-with-rejected-answer Balanced image-level tampering detection dataset in SmolVLM-style format with chosen/rejected answer pairs, derived from the pico-banana MCQ pipeline. Suitable for preference learning (e.g. DPO) and RLHF-style training. Dataset overview Same as vanloc1808/pico-banana-smolvlm-format, but each example includes a rejected_answer field: the answer from the counterpart sample (same edited/original image pair, opposite… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/pico-banana-smolvlm-format-with-rejected-answer.image100K<n<1M1 likes1.1k downloads7mo agoHugging Face03Fraser /pico-8-games PICO-8 Games Dataset The first multimodal dataset of PICO-8 games. 10,967 cartridges scraped from the Lexaloffle BBS, each decomposed into Lua source code, pixel-art spritesheets, tile maps, sound effects, music patterns, and metadata. Label screenshots from the top 48 games by star count What's Inside Every PICO-8 cartridge is a self-contained game packed into a single file. This dataset cracks each one open into its component parts: The… See the full description on the dataset page: https://huggingface.co/datasets/Fraser/pico-8-games.imagetext-generation10K<n<100K1 likes187 downloads6mo agoHugging Face04nroggendorff /picorpustext1M<n<10M1 likes145 downloads1y agoHugging Face05onkanat /rapberry_pi_pico_all-dataset 🤗 rapberry_pi_pico_all This dataset was automatically generated and verified using the Universal PDF & Rendergit Code Dataset Generator Pipeline (Phase 1-4). It contains high-quality synthetic code pairs, technical SFT Q&A, DPO (Direct Preference Optimization) preference pairs, and multi-turn technical chat sequences in both English and Turkish. 📊 Dataset Summary & Splits tr_dpo_dataset.jsonl: 756 örnek (samples) tr_sft_dataset.jsonl: 8,865 örnek (samples)… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/rapberry_pi_pico_all-dataset.texttext-generation10K<n<100K0 likes94 downloads2mo agoHugging Face06xensedyl /franka-revo2-pico4-hand-demo-testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "franka_research3_dexhand", "total_episodes": 7, "total_frames": 1293, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:7" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/franka-revo2-pico4-hand-demo-test.tabularrobotics1K<n<10K0 likes84 downloads3mo agoHugging Face07zjushining /franka-pico4-red-blockThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-pico4-red-block.tabularrobotics10K<n<100K0 likes62 downloads29d agoHugging Face08kyleavery /picoctf PicoCTF Challenges This dataset was originally only available on GitHub under agpl-3.0 license. I ported it to Hugging Face after making small corrections. textn<1K0 likes57 downloads11mo agoHugging Face09zjushining /franka-pico4-yellow-blockThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-pico4-yellow-block.tabularrobotics10K<n<100K0 likes55 downloads29d agoHugging Face10zjushining /franka-pico4-blue-blockThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "tcp.x", "tcp.y", "tcp.z", "tcp.r1", "tcp.r2", "tcp.r3", "tcp.r4", "tcp.r5", "tcp.r6", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-pico4-blue-block.tabularrobotics10K<n<100K0 likes52 downloads29d agoHugging Face11xensedyl /b601-pico4-demoThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "seeed_b601_rt_follower", "total_episodes": 3, "total_frames": 1909, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-pico4-demo.tabularrobotics1K<n<10K0 likes48 downloads4mo agoHugging Face12xensedyl /tron2-pico4-demoThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "left_tcp.x", "left_tcp.y", "left_tcp.z", "left_tcp.r1", "left_tcp.r2", "left_tcp.r3", "left_tcp.r4", "left_tcp.r5"… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/tron2-pico4-demo.tabularrobotics1K<n<10K0 likes48 downloads2mo agoHugging Face13pico-lm /pretokenized-paloma-tinsy The Tinsy Pretokenized Paloma Dataset A small version of the pretokenized-paloma benchmark dataset. This dataset is a sub-sampled version of pretokenized-paloma, and can be used in place of pretokenized-paloma. We release the exact scripts we use to create this dataset in our pico-lm/pico-dataset GitHub repo. text1K<n<10K0 likes45 downloads1y agoHugging Face14pico-lm /pretokenized-dolma-tinsy The Tinsy Pretokenized Dolma Dataset A tiny little baby-version of the pretokenized-dolma dataset. Meant to be used in a jupyter notebook to test things out, or quickly look at the structure of the data. 1K<n<10K0 likes42 downloads1y agoHugging Face15xensedyl /tron2-pico4-testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "left_tcp.x", "left_tcp.y", "left_tcp.z", "left_tcp.r1", "left_tcp.r2", "left_tcp.r3", "left_tcp.r4", "left_tcp.r5"… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/tron2-pico4-test.tabularrobotics1K<n<10K0 likes38 downloads2mo agoHugging Face16xensedyl /b601-pico4-demo1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "seeed_b601_rt_follower", "total_episodes": 3, "total_frames": 2167, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:3" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-pico4-demo1.tabularrobotics1K<n<10K0 likes37 downloads4mo agoHugging Face17xensedyl /b601-bi-pico4-demo-0611This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_seeed_b601_rt_follower", "total_episodes": 4, "total_frames": 1418, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:4" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-bi-pico4-demo-0611.tabularrobotics1K<n<10K0 likes32 downloads4mo agoHugging Face18reginaboateng /cleaned_ebmnlp_pico Dataset Card for "cleaned_ebmnlp_pico" More Information needed text10K<n<100K1 likes30 downloads4y agoHugging Face19xensedyl /b601-bi-pico4-demo-7camThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "bi_seeed_b601_rt_follower", "total_episodes": 1, "total_frames": 967, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/xensedyl/b601-bi-pico4-demo-7cam.tabularroboticsn<1K0 likes29 downloads4mo agoHugging Face20Marcus2112 /minipile_density-proportioned_picoProduced with the MiniCorpus pipeline, a reproduction and investigation of the MiniPile data-distillation method (Kaddour, 2023). This dataset is derived from and based on the contents of The Pile Deduplicated. text100K<n<1M0 likes28 downloads10mo agoHugging Face21reginaboateng /ebmnlp_pico Dataset Card for "ebmnlp_pico" More Information needed text10K<n<100K0 likes27 downloads4y agoHugging Face22cuevascarlos /PICO-breast-cancer PICO breast cancer dataset This dataset has been extracted from PICO-Corpus. The corpus consists of 1,011 abstracts of breast cancer randomized controlled trials extracted from PubMed. The PICO breast cancer dataset contains a total of 26 entities, compared to the usual 4 found in PICO corpora. Specifically, the following image extracted by the dataset's authors shows the hierarchy of the entities. The preprocessed dataset, ready to serve as inputs for MLMs such as BERT-like… See the full description on the dataset page: https://huggingface.co/datasets/cuevascarlos/PICO-breast-cancer.text1K<n<10K1 likes21 downloads2y agoHugging Face23thomasgauthier /unlicense-pico8 Pico-8 Unlicense Games Collection Dataset Summary A curated collection of all the games with code under Unlicense tagged PICO-8 from itch.io (29 games as of 2026-04-14). This dataset provides direct access to both the raw cartridge data and Lua source code for each game (extracted from the web player). Note: The assets are sometimes licensed under non-commercial terms - not everything's free for commercial use. Check the asset_license column for game-specific details.… See the full description on the dataset page: https://huggingface.co/datasets/thomasgauthier/unlicense-pico8.textn<1K0 likes20 downloads5mo agoHugging Face24NONHUMAN-RESEARCH /pico_laptop_reachThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so100_leader", "total_episodes": 8, "total_frames": 3054, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:8" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/NONHUMAN-RESEARCH/pico_laptop_reach.tabularrobotics1K<n<10K1 likes19 downloads3mo agoHugging Face25pico-lm /pretokenized-paloma The Pretokenized Paloma Benchmark Dataset This dataset is a compact, pre-tokenized evaluation dataset designed to complement the pretokenized-dolma training set. Built from the Paloma corpus (Allen Institute), this benchmark was designed to not contain any data overlap with Dolma and is ideal for evaluating models trained on it. Overview Features: Pre-tokenized with the same tokenizer as pretokenized-dolma: allenai/OLMo-7B-0724-hf Sequence length: 2048 tokens Ideal… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-paloma.text10K<n<100K0 likes15 downloads1y agoHugging Face26joey00072 /seeder_pico_thinking_function_callingtextn<1K1 likes13 downloads1y agoHugging Face27llm-4-kmu /pico_covid19textn<1K0 likes13 downloads1y agoHugging Face28yasmmin /pico-human-corpus_nerfair_processed Benchmark dataset PICO This dataset was generated by the Data preprocessing step of the NERFAIR workflow (More information: https://github.com/YasCoMa/ner-fair-workflow ) Original dataset: https://github.com/sociocom/PICO-Corpus/tree/main/pico_corpus_brat_annotated_files texttoken-classification1K<n<10K1 likes13 downloads11mo agoHugging Face29reginaboateng /pico_ebmnlp Dataset Card for "pico_ebmnlp" More Information needed text10K<n<100K0 likes11 downloads4y agoHugging Face30joey00072 /pico_thinking_function_callingtextn<1K1 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.