CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /filtered-wit Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image Text Dataset, here Data was taken from dalle-mini/wit Author Aarush Katta Data Structure The data is stored as tars, containing 10,000 samples per tar. The parquets contain the metadata of each tar, which was crated using this script Each tar contains a .jpg, .txt, and .json. The image is stored in .jpg, the caption in .txt. and the metadata in… See the full description on the dataset page: https://huggingface.co/datasets/laion/filtered-wit.image1M<n<10M11 likes6.6k downloads5y agoHugging Face02common-pile /stackv2_edu_filtered Stack V2 Edu Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.tabulartext-generation10M<n<100M6 likes6.3k downloads1y agoHugging Face03natolambert /GeneralThought-430K-filteredData from https://huggingface.co/datasets/GeneralReasoning/GeneralThought-430K, removed prompts with non commercial data tabular100K<n<1M35 likes3.3k downloads1y agoHugging Face04yxchng /laion_synthetic_filtered_large_part3image10M<n<100M0 likes3.2k downloads3y agoHugging Face05yxchng /laion_synthetic_filtered_large_part1image10M<n<100M2 likes3.2k downloads3y agoHugging Face06yxchng /laion_synthetic_filtered_large_part2image10M<n<100M0 likes3k downloads3y agoHugging Face07Magpie-Align /Magpie-Qwen2.5-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Qwen2.5-Pro-300K-Filtered.tabular100K<n<1M14 likes2.9k downloads2y agoHugging Face08yxchng /laion_synthetic_filtered_large_part4image10M<n<100M0 likes1.6k downloads3y agoHugging Face09G-reen /cc-re-2020-filtered Auto-Generated FastDetector Dataset Model Name: google/gemma-4-E4B-it Sampling Params (as sent to the engine): {"temperature": 0.0, "top_p": 1.0, "presence_penalty": 0.0} Ignored Params (unsupported by this engine): None Prompt File: prompts/filter_contiguous_subset.json Total Train Prompts: 1 Source Dataset: G-reen/cc-re-2020-raw-sharded Source Column: text Target Num Samples: all Dropped Samples (over length limit 15000 tokens): 430 Failed API Requests: 495 Total… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-re-2020-filtered.tabular1M<n<10M0 likes1.5k downloads17d agoHugging Face10hotchpotch /mmarco-hard-negatives-reranker-filtered mMARCO Reranker-Filtered Hard Negatives (Multilingual) Overview This dataset is built from mMARCO (multilingual MS MARCO) triplets for each language subset. For each (query, positive), hard negatives are bundled and then filtered using cross-encoder re-scoring. The goal is to remove negatives that are too strong or incorrect for training. The same procedure is applied to all language subsets. The dataset is published as mmarco-hard-negatives-reranker-filtered with… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/mmarco-hard-negatives-reranker-filtered.tabular10M<n<100M3 likes1.5k downloads3mo agoHugging Face11ytzi /the-stack-dedup-python-filtered-docstringsThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_function_no_docstring remove_class_no_docstring remove_delete_markers tabular10M<n<100M0 likes1.4k downloads2y agoHugging Face12Magpie-Align /Magpie-Llama-3.1-Pro-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.3k downloads2y agoHugging Face13kaggle-aimo /amc_filteredtabular1K<n<10K0 likes1.3k downloads2y agoHugging Face14Magpie-Align /Magpie-Llama-3.1-Pro-MT-300K-Filtered Project Web: https://magpie-align.github.io/ Arxiv Technical Report: https://arxiv.org/abs/2406.08464 Codes: https://github.com/magpie-align/magpie Abstract Click Here High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tabulartext-generation100K<n<1M17 likes1.2k downloads2y agoHugging Face15ytzi /the-stack-dedup-python-filtered-allThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_non_ascii remove_decorators remove_async remove_classes remove_generators remove_function_no_docstring remove_class_no_docstring remove_unused_imports remove_delete_markers tabular10M<n<100M0 likes825 downloads2y agoHugging Face16kaggle-aimo /aime_filteredtabularn<1K1 likes753 downloads2y agoHugging Face17brandonyang /artem-fold-towel-filtered Artem fold-towel filtered trajectories Observation-only LeRobot v3 derivative of brandonyang/artem-fold-towel. It contains 781 demonstrations (1048134 frames) accepted by the continuous bimanual YAM replayability pipeline. The 14-D observation.state contains the smoothed, trajectory-optimized YAM-achievable UMI1 pose, normalized UMI1 gripper, UMI2 pose, and normalized UMI2 gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved; action is… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel-filtered.tabularrobotics1M<n<10M0 likes704 downloads21d agoHugging Face18yxchng /ccs_synthetic_filtered_largeimage10M<n<100M0 likes691 downloads3y agoHugging Face19VibeCuisine /cucumber-place-classifier-filtered071126This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 20, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos", "tilt.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/cucumber-place-classifier-filtered071126.tabularrobotics1K<n<10K0 likes588 downloads2mo agoHugging Face20ytzi /the-stack-dedup-python-filtered-docstrings-gpt2tabular10M<n<100M0 likes557 downloads2y agoHugging Face21teven /mpww_filtered_all_passagestabular10M<n<100M0 likes539 downloads4y agoHugging Face22Eurolingua /HPLT3_DE_0.9_Quantile_Adult_Filteredtabular1M<n<10M1 likes522 downloads7mo agoHugging Face23brandonyang /dual-lidar-combined-filtered-long-gripper Combined filtered dual-LiDAR UMI demonstrations Observation-only LeRobot v3 derivative of brandonyang/dual-lidar-umi, brandonyang/dual-lidar-umi-relative. It contains 182 demonstrations (179951 frames) accepted by the continuous bimanual YAM replayability pipeline. The 12-D observation.state contains the smoothed, trajectory-optimized YAM-achievable path in the zero-origin UMI Cartesian convention. Raw UMI gripper widths remain as separate observations. The two original UMI… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/dual-lidar-combined-filtered-long-gripper.tabularrobotics100K<n<1M0 likes519 downloads1mo agoHugging Face24G-reen /fastdetector-train-filteredtabular100K<n<1M0 likes509 downloads3d agoHugging Face25ytzi /the-stack-dedup-python-filtered-non_asciiThis is a dataset originated from bigcode/the-stack-dedup with some filters applied. The filters filtered in this dataset are: remove_non_ascii tabular10M<n<100M1 likes457 downloads2y agoHugging Face26G-reen /fastdetector-val-filteredtabular10K<n<100K0 likes455 downloads3d agoHugging Face27TheFinAI /00_filteredtabular10M<n<100M0 likes437 downloads4mo agoHugging Face28M1keR /the-stack-v2-dedup-filtered-500-stars-100-forks-contentstabular1M<n<10M1 likes434 downloads1y agoHugging Face29maitf /yt-music-indexonly-filteredtabular100M<n<1B0 likes424 downloads6mo agoHugging Face30Ba2han /finetranslations-TR_filtered Filtered out too long examples Applied basic n-gram repetition filter tabular1M<n<10M0 likes421 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.