datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aitf-dfk3-vlm-dataset-jsonlteletype
Dataset Card for Teletype
This dataset is a scrape of all articles published on teletype, a popular platform for publishing articles, especially in Telegram. The dataset includes the original article HTML, as well as text extracted using the trafilatura library with favor_recall=True and other metadata provided by teletype.
Additionally, language identification was applied using the lingua-py library and the identification results are available in the lang column.
Curated by: its5Q
