datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TA-WB-MapAnythingLSDBench
Dataset Card for LSDBench: Long-video Sampling Dilemma Benchmark
A benchmark that focuses on the sampling dilemma in long-video tasks. Through well-designed tasks, it evaluates the sampling efficiency of long-video VLMs.
Arxiv Paper: 📖 Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
Github : https://github.com/dvlab-research/LSDBench
(Left) In Q1, identifying a camera wearer's visited locations requires analyzing the entire video. However, key frames… See the full description on the dataset page: https://huggingface.co/datasets/TainU/LSDBench.IV-Edit
IV-Edit Benchmark & RePlan Training Data
Dataset Summary
This repository contains the IV-Edit (Instruction-Visual Editing) Benchmark and the training data used for the RePlan framework. The dataset is designed to address the challenge of Instruction-Visual Complexity (IV-Complexity) in instruction-based image editing, where intricate instructions interact with cluttered or ambiguous visual scenes.… See the full description on the dataset page: https://huggingface.co/datasets/TainU/IV-Edit.amazon_massive_intent_ta-INamazon_massive_scenario_ta-INTA-WB
TA-WB dataset used in UFM training
Warning! This dataset cannot be used for MapAnything, as it pack only the optical flow but not the depthmap. We are uploading that soon!
indian_supreme_court_judgements_en_ta
Indian Supreme Court Judgements Dataset (Sentence-Level, Translated to Tamil)
Overview
This dataset contains Indian Supreme Court judgements that have been split into sentences and translated into Tamil. The original judgements were sourced from the Indian Kanoon website. The dataset is useful for legal text processing, multilingual NLP tasks, and cross-lingual legal studies.
Data Processing Pipeline
Sentence Splitting:
Used pySBD (Python Sentence Boundary… See the full description on the dataset page: https://huggingface.co/datasets/Narenameme/indian_supreme_court_judgements_en_ta.shan-novel-tainovel_comafrica-uganda-construction-input-price-index-cipi-november-2024-excel-ta-64e06ce6
Construction Input Price Index Cipi November 2024 Excel Ta | Africa (Uganda Bureau of Statistics)
6,942 rows - 1 Africa country/area - 2017-2024 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 6,942 rows from Uganda Bureau of Statistics, covering Construction Input Price Index Cipi November 2024 Excel Ta. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-construction-input-price-index-cipi-november-2024-excel-ta-64e06ce6.africa-uganda-construction-input-price-index-cipi-december-2024-excel-ta-bf215a60
Construction Input Price Index Cipi December 2024 Excel Ta | Africa (Uganda Bureau of Statistics)
7,020 rows - 1 Africa country/area - 2017-2024 - 1 indicator - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 7,020 rows from Uganda Bureau of Statistics, covering Construction Input Price Index Cipi December 2024 Excel Ta. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-uganda-construction-input-price-index-cipi-december-2024-excel-ta-bf215a60.truthfulqa_tainstructions-ta
Dataset Card for "instructions-ta"
More Information needed
arc_tatainingTainaCostahellaswag_tafleurs_ta_inroots_indic-ta_wikisourceROOTS Subset: roots_indic-ta_wikisource
wikisource_filtered
Dataset uid: wikisource_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
2.6306 % of total
12.7884 % of fr
19.8886 % of indic-bn
20.9966 % of indic-ta
2.3478 % of ar
4.7068 % of indic-hi
18.0998 % of indic-te
1.7155 % of es
19.4800 % of indic-kn
9.1737 % of indic-ml
17.1771 % of indic-mr
17.1870 % of indic-gu
70.3687 % of indic-as
1.0165 % of pt
7.8642… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_wikisource.ta-flores-inafrica-tunisia-les-ressources-humaines-des-institutions-culturelles-de-ta-c1abf98a
Les Ressources Humaines Des Institutions Culturelles De Ta | Africa (Tunisia Open Data)
31 rows - 1 Africa country/area - 2019 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 31 rows from Tunisia Open Data, covering Les Ressources Humaines Des Institutions Culturelles De Ta. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-tunisia-les-ressources-humaines-des-institutions-culturelles-de-ta-c1abf98a.roots_indic-ta_wikibooksROOTS Subset: roots_indic-ta_wikibooks
wikibooks_filtered
Dataset uid: wikibooks_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0897 % of total
0.2591 % of en
0.0965 % of fr
0.1691 % of es
0.2834 % of indic-hi
0.2172 % of pt
0.0149 % of zh
0.0279 % of ar
0.1374 % of vi
0.5025 % of id
0.3694 % of indic-ur
0.5744 % of eu
0.0769 % of ca
0.0519 % of indic-ta
0.1470 % of indic-mr
0.0751 % of indic-te
0.0156 % of… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_wikibooks.shan-novel-tainovel_com
Dataset Card for "haohaa/shan-novel-tainovel_com"
This dataset was scrape from tainovel.com. Shan language novel website.
Language
Shan - shn
Date Version
May 18, 2024
items_prompts_fulltain_arxivroots_indic-ta_wikinewsROOTS Subset: roots_indic-ta_wikinews
wikinews_filtered
Dataset uid: wikinews_filtered
Description
Homepage
Licensing
Speaker Locations
Sizes
0.0307 % of total
0.0701 % of ar
0.3036 % of pt
0.0271 % of en
0.0405 % of fr
0.2119 % of indic-ta
0.0081 % of zh
0.0510 % of es
0.0725 % of ca
BigScience processing steps
Filters applied to: ar
filter_wiki_user_titles
filter_wiki_non_text_type
dedup_document… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_wikinews.Tainaitems_raw_liteitems_prompts_liteinternet_demo_env_setuproots_indic-ta_indic_nlp_corpusROOTS Subset: roots_indic-ta_indic_nlp_corpus
Indic NLP Corpus
Dataset uid: indic_nlp_corpus
Description
The IndicNLP corpus is a largescale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. s (IndoAryan branch and Dravidian). Each language has at least 100 million words (except Oriya).
Homepage
https://github.com/AI4Bharat/indicnlp_corpus#publicly-available-classification-datasets
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-data/roots_indic-ta_indic_nlp_corpus.
