CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01infinity1096 /TA-WB-MapAnything1 likes1.5k downloads8mo agoHugging Face02TainU /LSDBench Dataset Card for LSDBench: Long-video Sampling Dilemma Benchmark A benchmark that focuses on the sampling dilemma in long-video tasks. Through well-designed tasks, it evaluates the sampling efficiency of long-video VLMs. Arxiv Paper: 📖 Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma? Github : https://github.com/dvlab-research/LSDBench (Left) In Q1, identifying a camera wearer's visited locations requires analyzing the entire video. However, key frames… See the full description on the dataset page: https://huggingface.co/datasets/TainU/LSDBench.textvideo-text-to-text1K<n<10K0 likes572 downloads1y agoHugging Face03TainU /IV-Edit IV-Edit Benchmark & RePlan Training Data                     Dataset Summary This repository contains the IV-Edit (Instruction-Visual Editing) Benchmark and the training data used for the RePlan framework. The dataset is designed to address the challenge of Instruction-Visual Complexity (IV-Complexity) in instruction-based image editing, where intricate instructions interact with cluttered or ambiguous visual scenes.… See the full description on the dataset page: https://huggingface.co/datasets/TainU/IV-Edit.imageimage-to-image1K<n<10K4 likes142 downloads9mo agoHugging Face04SetFit /amazon_massive_intent_ta-INtext10K<n<100K0 likes138 downloads4y agoHugging Face05SetFit /amazon_massive_scenario_ta-INtext10K<n<100K0 likes62 downloads4y agoHugging Face06infinity1096 /TA-WB TA-WB dataset used in UFM training Warning! This dataset cannot be used for MapAnything, as it pack only the optical flow but not the depthmap. We are uploading that soon! 1 likes56 downloads11mo agoHugging Face07Narenameme /indian_supreme_court_judgements_en_ta Indian Supreme Court Judgements Dataset (Sentence-Level, Translated to Tamil) Overview This dataset contains Indian Supreme Court judgements that have been split into sentences and translated into Tamil. The original judgements were sourced from the Indian Kanoon website. The dataset is useful for legal text processing, multilingual NLP tasks, and cross-lingual legal studies. Data Processing Pipeline Sentence Splitting: Used pySBD (Python Sentence Boundary… See the full description on the dataset page: https://huggingface.co/datasets/Narenameme/indian_supreme_court_judgements_en_ta.tabular1M<n<10M2 likes27 downloads2y agoHugging Face08NorHsangPha /shan-novel-tainovel_comtexttext-classificationn<1K0 likes23 downloads8mo agoHugging Face09electricsheepafrica /africa-uganda-construction-input-price-index-cipi-november-2024-excel-ta-64e06ce6 Construction Input Price Index Cipi November 2024 Excel Ta | Africa (Uganda Bureau of Statistics) 6,942 rows - 1 Africa country/area - 2017-2024 - 1 indicator - Engineered by Electric Sheep Africa TL;DR This dataset contains 6,942 rows from Uganda Bureau of Statistics, covering Construction Input Price Index Cipi November 2024 Excel Ta. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-construction-input-price-index-cipi-november-2024-excel-ta-64e06ce6.tabulartabular-regression1K<n<10K0 likes14 downloads1mo agoHugging Face10electricsheepafrica /africa-uganda-construction-input-price-index-cipi-december-2024-excel-ta-bf215a60 Construction Input Price Index Cipi December 2024 Excel Ta | Africa (Uganda Bureau of Statistics) 7,020 rows - 1 Africa country/area - 2017-2024 - 1 indicator - Engineered by Electric Sheep Africa TL;DR This dataset contains 7,020 rows from Uganda Bureau of Statistics, covering Construction Input Price Index Cipi December 2024 Excel Ta. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-construction-input-price-index-cipi-december-2024-excel-ta-bf215a60.tabulartabular-regression1K<n<10K0 likes14 downloads1mo agoHugging Face11indicbench /truthfulqa_ta0 likes12 downloads2y agoHugging Face12cahya /instructions-ta Dataset Card for "instructions-ta" More Information needed text1K<n<10K0 likes11 downloads4y agoHugging Face13indicbench /arc_tatext1K<n<10K0 likes11 downloads2y agoHugging Face14tejasgadge2504 /tainingtextn<1K0 likes11 downloads2y agoHugging Face15Medradome /TainaCostaaudion<1K0 likes10 downloads3y agoHugging Face16indicbench /hellaswag_tatext10K<n<100K0 likes10 downloads2y agoHugging Face17changelinglab /fleurs_ta_inaudion<1K0 likes10 downloads2y agoHugging Face18bigscience-data /roots_indic-ta_wikisourcegatedROOTS Subset: roots_indic-ta_wikisource wikisource_filtered Dataset uid: wikisource_filtered Description Homepage Licensing Speaker Locations Sizes 2.6306 % of total 12.7884 % of fr 19.8886 % of indic-bn 20.9966 % of indic-ta 2.3478 % of ar 4.7068 % of indic-hi 18.0998 % of indic-te 1.7155 % of es 19.4800 % of indic-kn 9.1737 % of indic-ml 17.1771 % of indic-mr 17.1870 % of indic-gu 70.3687 % of indic-as 1.0165 % of pt 7.8642… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_wikisource.text100K<n<1M0 likes9 downloads4y agoHugging Face19kornwtp /ta-flores-intext1K<n<10K0 likes9 downloads2y agoHugging Face20electricsheepafrica /africa-tunisia-les-ressources-humaines-des-institutions-culturelles-de-ta-c1abf98a Les Ressources Humaines Des Institutions Culturelles De Ta | Africa (Tunisia Open Data) 31 rows - 1 Africa country/area - 2019 - source table - Engineered by Electric Sheep Africa TL;DR This dataset contains 31 rows from Tunisia Open Data, covering Les Ressources Humaines Des Institutions Culturelles De Ta. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples. What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-les-ressources-humaines-des-institutions-culturelles-de-ta-c1abf98a.tabulartabular-classificationn<1K0 likes9 downloads1mo agoHugging Face21bigscience-data /roots_indic-ta_wikibooksgatedROOTS Subset: roots_indic-ta_wikibooks wikibooks_filtered Dataset uid: wikibooks_filtered Description Homepage Licensing Speaker Locations Sizes 0.0897 % of total 0.2591 % of en 0.0965 % of fr 0.1691 % of es 0.2834 % of indic-hi 0.2172 % of pt 0.0149 % of zh 0.0279 % of ar 0.1374 % of vi 0.5025 % of id 0.3694 % of indic-ur 0.5744 % of eu 0.0769 % of ca 0.0519 % of indic-ta 0.1470 % of indic-mr 0.0751 % of indic-te 0.0156 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_wikibooks.textn<1K0 likes8 downloads4y agoHugging Face22haohaa /shan-novel-tainovel_com Dataset Card for "haohaa/shan-novel-tainovel_com" This dataset was scrape from tainovel.com. Shan language novel website. Language Shan - shn Date Version May 18, 2024 texttext-generationn<1K0 likes8 downloads2y agoHugging Face23tainhoz91 /items_prompts_fulltext100K<n<1M0 likes8 downloads7mo agoHugging Face24silksmooth /tain_arxiv0 likes7 downloads3y agoHugging Face25bigscience-data /roots_indic-ta_wikinewsgatedROOTS Subset: roots_indic-ta_wikinews wikinews_filtered Dataset uid: wikinews_filtered Description Homepage Licensing Speaker Locations Sizes 0.0307 % of total 0.0701 % of ar 0.3036 % of pt 0.0271 % of en 0.0405 % of fr 0.2119 % of indic-ta 0.0081 % of zh 0.0510 % of es 0.0725 % of ca BigScience processing steps Filters applied to: ar filter_wiki_user_titles filter_wiki_non_text_type dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_wikinews.text1K<n<10K0 likes6 downloads4y agoHugging Face26Medradome /Tainaaudion<1K0 likes6 downloads3y agoHugging Face27tainhoz91 /items_raw_litetabular10K<n<100K0 likes6 downloads7mo agoHugging Face28tainhoz91 /items_prompts_litetext10K<n<100K0 likes6 downloads7mo agoHugging Face29TainU /internet_demo_env_setup0 likes6 downloads2mo agoHugging Face30bigscience-data /roots_indic-ta_indic_nlp_corpusgatedROOTS Subset: roots_indic-ta_indic_nlp_corpus Indic NLP Corpus Dataset uid: indic_nlp_corpus Description The IndicNLP corpus is a largescale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. s (IndoAryan branch and Dravidian). Each language has at least 100 million words (except Oriya). Homepage https://github.com/AI4Bharat/indicnlp_corpus#publicly-available-classification-datasets Licensing… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_indic_nlp_corpus.text10M<n<100M0 likes5 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.