CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01taesiri /imagenet_hard_review_data_r2tabular1K<n<10K0 likes5.8k downloads3y agoHugging Face02taesiri /VideoGameQA-Bench VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance by Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer Abstract: With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer considerable potential to automate and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/VideoGameQA-Bench.imageimage-to-text1K<n<10K7 likes2.8k downloads1y agoHugging Face03taesiri /arxiv_audioaudio1K<n<10K2 likes1.9k downloads3y agoHugging Face04taesiri /ArXivSignals-DeepSummaries ArXivSignals DeepSummaries — Agent-Built Visual Paper Explainers A continuously-updated, day-partitioned dataset of deep, visual summaries of arXiv papers, each built by a coding agent working inside the paper's own LaTeX source: the agent reads the full text, authors an editorial narrative as a structured content spec, and the paper's real figures and tables (extracted and rendered from the LaTeX, web-optimized) ride along as an embedded, variable-length image array. The… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-DeepSummaries.tabularsummarization1K<n<10K7 likes1.5k downloads13h agoHugging Face05taesiri /imagenet-hard-4K Dataset Card for "Imagenet-Hard-4K" Project Page - Paper - Github ImageNet-Hard-4K is 4K version of the original ImageNet-Hard dataset, which is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard-4K.imageimage-classification1K<n<10K7 likes1.2k downloads11mo agoHugging Face06taesiri /FragileXtextn<1K0 likes1.1k downloads3y agoHugging Face07taesiri /ArXivSignals ArXivSignals — Daily arXiv Papers with LLM Signal & Summaries A continuously-updated, day-partitioned dataset of arXiv papers (AI/ML and adjacent categories) enriched with LLM-derived signal: a 0–100 importance score, topical/lab tags, a one-line takeaway, and — for a selected subset — dense full-page summaries. It powers arxivsignals.io and is published here as an open research resource. The dataset has two configs: papers (default) — one row per paper: bibliography +… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals.imagetext-classification100K<n<1M6 likes965 downloads14h agoHugging Face08taeminlee /Ko-StrategyQA Ko-StrategyQA This dataset represents a conversion of the Ko-StrategyQA dataset into the BeIR format, making it compatible for use with mteb. The original dataset was designed for multi-hop QA, so we processed the data accordingly. First, we grouped the evidence documents tagged by annotators into sets, and excluded unit questions containing 'no_evidence' or 'operation'. texttext-retrieval10K<n<100K21 likes814 downloads1y agoHugging Face09taegyoun88 /egoxtreme EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions 📖 Dataset Information EgoXtreme is a novel large-scale dataset designed for robust egocentric 6D object pose estimation under extreme environmental conditions. The dataset comprises approximately 1.3 million frames with a total duration of 775.5 minutes (~12.9 hours). It was captured at 30 fps using Aria glasses, providing high-resolution 1408 x 1408 raw fisheye RGB… See the full description on the dataset page: https://huggingface.co/datasets/taegyoun88/egoxtreme.3dobject-detection1M<n<10M4 likes701 downloads6mo agoHugging Face10taesiri /arxiv_summarytextn<1K1 likes382 downloads3y agoHugging Face11taesiri /imagenet_hard_review_datatabular1K<n<10K0 likes374 downloads3y agoHugging Face12taesiri /GameplayCaptions-GPT-4Vimage10K<n<100K1 likes300 downloads3y agoHugging Face13taesiri /SteamGlitches-Gemini-Labelsimage100K<n<1M0 likes300 downloads1y agoHugging Face14taesiri /GameplayCaptions-Gemini-pro-visionimage10K<n<100K7 likes298 downloads2y agoHugging Face15taesiri /arxiv_qa ArXiv QA (TBD) Automated ArXiv question answering via large language models Github | Homepage | Simple QA - Hugging Face Space Automated Question Answering with ArXiv Papers Latest 25 Papers LIME: Localized Image Editing via Attention Regularization in Diffusion Models - [Arxiv] [QA] Revisiting Depth Completion from a Stereo Matching Perspective for Cross-domain Generalization - [Arxiv] [QA] VL-GPT: A Generative Pre-trained Transformer for Vision and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/arxiv_qa.textquestion-answering100K<n<1M138 likes288 downloads2y agoHugging Face16microsoft /msr-acc-tae25 Microsoft Research - Accurate Chemistry Collection: Total Atomization Energies Description The Microsoft Research Accurate Chemistry Collection (MSR-ACC) provides a collection of accurate coupled cluster labels for training machine learning functionals. MSR-ACC/TAE25 comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol. The dataset is constructed to exhaustively cover the chemical space of closed-shell… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/msr-acc-tae25.tabular10K<n<100K8 likes287 downloads5mo agoHugging Face17taesiri /imagenet-hard Dataset Card for "ImageNet-Hard" Project Page - ArXiv - Paper - Github - Image Browser Dataset Summary ImageNet-Hard is a new benchmark that comprises 10,980 images collected from various existing ImageNet-scale benchmarks (ImageNet, ImageNet-V2, ImageNet-Sketch, ImageNet-C, ImageNet-R, ImageNet-ReaL, ImageNet-A, and ObjectNet). This dataset poses a significant challenge to state-of-the-art vision models as merely zooming in often fails to improve their ability to… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/imagenet-hard.imageimage-classification10K<n<100K12 likes234 downloads11mo agoHugging Face18taesiri /GameplayCaptions-GPT-4V-V2image10K<n<100K2 likes229 downloads3y agoHugging Face19taesiri /Gameplay-Walkthrough-QAtabular100K<n<1M1 likes229 downloads2y agoHugging Face20taesiri /VGQAimage1K<n<10K1 likes220 downloads11mo agoHugging Face21taesiri /TinyStories-Farsi Tiny Stories Farsi The Tiny Stories Farsi project is a continuous effort to translate the Tiny Stories dataset into the Persian (Farsi) language. The primary goal is to produce a high-quality Farsi dataset, maintaining equivalency with the original English version, and subsequently to utilize it for training language models in Farsi. This seeks to affirm that the advancements and trends observed in English language models are replicable and applicable in other languages. Thus far… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/TinyStories-Farsi.texttext-generation100K<n<1M18 likes198 downloads3y agoHugging Face22taesiri /VGBDataset-Small-HF[Paper] - [Website] imageimage-to-text10K<n<100K0 likes177 downloads2y agoHugging Face23taejoon89 /Ko-Agent-Trajectories-1.0 Ko-Agent-Trajectories-1.0 Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository under pipeline/, together with the API catalogue, the scenario templates and the complete prompt set. The card reports the completed human review study and the v1.1 artefacts (behaviour DPO config, per-item validation scores, manifest, filter asset). Korean edition: README.ko.md. TL;DR A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.tabulartext-generation10K<n<100K0 likes165 downloads3h agoHugging Face24taeshahn /ko-lima Dataset Card for KoLIMA Dataset Description KoLIMA는 Meta에서 공개한 LIMA: Less Is More for Alignment (Zhou et al., 2023)의 학습 데이터를 한국어로 번역한 데이터셋입니다. 번역에는 DeepL API를 활용하였고, SK(주) Tech Collaborative Lab으로부터 비용을 지원받았습니다. 전체 텍스트 중에서 code block이나 수식을 나타내는 특수문자 사이의 텍스트는 원문을 유지하는 형태로 번역을 진행하였으며, train 데이터셋 1,030건과 test 데이터셋 300건으로 구성된 총 1,330건의 데이터를 활용하실 수 있습니다. 현재 동일한 번역 문장을 plain, vicuna 두 가지 포멧으로 제공합니다. 데이터셋 관련하여 문의가 있으신 경우 메일을 통해 연락주세요! 🥰 This is Korean LIMA dataset, which is… See the full description on the dataset page: https://huggingface.co/datasets/taeshahn/ko-lima.text1K<n<10K15 likes160 downloads3y agoHugging Face25taesiri /T2I-Prompt-Sample-Imagesimage10K<n<100K1 likes160 downloads2y agoHugging Face26taesiri /VGBDataset-Large-HF[Paper] - [Website] imageimage-to-text100K<n<1M0 likes142 downloads2y agoHugging Face27taesiri /video-game-question-answeringimage10K<n<100K3 likes130 downloads3y agoHugging Face28Taeham /wav2vec2-ksponspeech-train2text10K<n<100K0 likes126 downloads4y agoHugging Face29Taeham /wav2vec2-ksponspeech-traintext10K<n<100K0 likes116 downloads4y agoHugging Face30taesiri /BlindLoop-Evaluationgated BlindLoop Evaluation The frozen evaluation cohorts for the BlindLoop paper. Each eligible generated task contributes exactly five deterministic, pixel-distinct image instances. The five rows share a task's selected question/prompt family while varying the rendered scene and gold answer as determined by the task's pixel oracle. Config Tasks Rows Documented exclusions section1_eval5 1,298 6,490 3 section2_eval5 874 4,370 1 combined_eval5 2,172 10,860 4 Each row… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/BlindLoop-Evaluation.imagevisual-question-answering10K<n<100K1 likes111 downloads28d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.