CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taesiri /SteamScreenshots-Bugs Samples imageimage-to-text100K<n<1M2 likes21k downloads1y agoHugging Face02taesiri /arxiv_db4 likes8.9k downloads2y agoHugging Face03taesiri /arxiv_qa25 likes7.8k downloads2y agoHugging Face04taesiri /CHMCorrUserResponse0 likes7.7k downloads3y agoHugging Face05taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes6.1k downloads3y agoHugging Face06taesiri /GBPro-RAWimage2 likes5k downloads1y agoHugging Face07taesiri /VideoGameQA-Bench VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance by Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer Abstract: With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer considerable potential to automate and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/VideoGameQA-Bench.imageimage-to-text1K<n<10K7 likes2.8k downloads1y agoHugging Face08taesiri /steam_screenshots_samples_2image10K<n<100K0 likes2.7k downloads2y agoHugging Face09taesiri /arxiv_audioaudio1K<n<10K2 likes1.9k downloads3y agoHugging Face10taesiri /ArXivSignals-DeepSummaries ArXivSignals DeepSummaries — Agent-Built Visual Paper Explainers A continuously-updated, day-partitioned dataset of deep, visual summaries of arXiv papers, each built by a coding agent working inside the paper's own LaTeX source: the agent reads the full text, authors an editorial narrative as a structured content spec, and the paper's real figures and tables (extracted and rendered from the LaTeX, web-optimized) ride along as an embedded, variable-length image array. The… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-DeepSummaries.tabularsummarization1K<n<10K7 likes1.4k downloads15h agoHugging Face11taesiri /imagenet-hard-4K Dataset Card for "Imagenet-Hard-4K" Project Page - Paper - Github ImageNet-Hard-4K is 4K version of the original ImageNet-Hard dataset, which is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard-4K.imageimage-classification1K<n<10K7 likes1.2k downloads10mo agoHugging Face12taesiri /FragileXtextn<1K0 likes1.1k downloads3y agoHugging Face13taesiri /SteamScreenshotsimage0 likes1.1k downloads2y agoHugging Face14taesiri /ArXivSignals ArXivSignals — Daily arXiv Papers with LLM Signal & Summaries A continuously-updated, day-partitioned dataset of arXiv papers (AI/ML and adjacent categories) enriched with LLM-derived signal: a 0–100 importance score, topical/lab tags, a one-line takeaway, and — for a selected subset — dense full-page summaries. It powers arxivsignals.io and is published here as an open research resource. The dataset has two configs: papers (default) — one row per paper: bibliography +… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals.imagetext-classification100K<n<1M6 likes930 downloads16h agoHugging Face15taesiri /imagenet_hard_review_datatabular1K<n<10K0 likes888 downloads3y agoHugging Face16taeminlee /Ko-StrategyQA Ko-StrategyQA This dataset represents a conversion of the Ko-StrategyQA dataset into the BeIR format, making it compatible for use with mteb. The original dataset was designed for multi-hop QA, so we processed the data accordingly. First, we grouped the evidence documents tagged by annotators into sets, and excluded unit questions containing 'no_evidence' or 'operation'. texttext-retrieval10K<n<100K21 likes836 downloads1y agoHugging Face17taegyoun88 /egoxtreme EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions 📖 Dataset Information EgoXtreme is a novel large-scale dataset designed for robust egocentric 6D object pose estimation under extreme environmental conditions. The dataset comprises approximately 1.3 million frames with a total duration of 775.5 minutes (~12.9 hours). It was captured at 30 fps using Aria glasses, providing high-resolution 1408 x 1408 raw fisheye RGB… See the full description on the dataset page: https://huggingface.co/datasets/taegyoun88/egoxtreme.3dobject-detection1M<n<10M4 likes694 downloads6mo agoHugging Face18taesiri /arxiv_summarytextn<1K1 likes386 downloads3y agoHugging Face19taejinp /acoustic_context_switchingaudio1K<n<10K1 likes367 downloads1y agoHugging Face20microsoft /msr-acc-tae25 Microsoft Research - Accurate Chemistry Collection: Total Atomization Energies Description The Microsoft Research Accurate Chemistry Collection (MSR-ACC) provides a collection of accurate coupled cluster labels for training machine learning functionals. MSR-ACC/TAE25 comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol. The dataset is constructed to exhaustively cover the chemical space of closed-shell… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/msr-acc-tae25.tabular10K<n<100K8 likes357 downloads5mo agoHugging Face21nickypro /tae-data-split-paragraphs Split Paragraphs Dataset Split paragraphs data with configs 000-099. 0 likes347 downloads1y agoHugging Face22taesiri /GameplayCaptions-GPT-4Vimage10K<n<100K1 likes299 downloads3y agoHugging Face23taesiri /GameplayCaptions-Gemini-pro-visionimage10K<n<100K7 likes297 downloads2y agoHugging Face24taesiri /SteamGlitches-Gemini-Labelsimage100K<n<1M0 likes297 downloads1y agoHugging Face25nickypro /tae-data-embeddingstabular1M<n<10M0 likes286 downloads1y agoHugging Face26taesiri /arxiv_qa ArXiv QA (TBD) Automated ArXiv question answering via large language models Github | Homepage | Simple QA - Hugging Face Space Automated Question Answering with ArXiv Papers Latest 25 Papers LIME: Localized Image Editing via Attention Regularization in Diffusion Models - [Arxiv] [QA] Revisiting Depth Completion from a Stereo Matching Perspective for Cross-domain Generalization - [Arxiv] [QA] VL-GPT: A Generative Pre-trained Transformer for Vision and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/arxiv_qa.textquestion-answering100K<n<1M138 likes284 downloads2y agoHugging Face27taechasith /kala-uap-archivedocumentn<1K0 likes278 downloads4mo agoHugging Face28taesiri /imagenet-hard Dataset Card for "ImageNet-Hard" Project Page - ArXiv - Paper - Github - Image Browser Dataset Summary ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their ability to… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard.imageimage-classification10K<n<100K12 likes233 downloads10mo agoHugging Face29taesiri /GameplayCaptions-GPT-4V-V2image10K<n<100K2 likes228 downloads3y agoHugging Face30taesiri /Gameplay-Walkthrough-QAtabular100K<n<1M1 likes228 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.