CoolFace
20 results

habr

volosati /habr Habr dataset (mirror) Provenance / History This is a re-upload, not an original creation. The original dataset, IlyaGusev/habr, was created and maintained by Ilya Gusev. Its disappearance from the Hub is documented by the author himself in a series of posts on his Telegram channel (@senior_augur, post 599 onward): 2026-07-05 — Habr (the website the dataset scrapes) sent Ilya a legal complaint about the dataset. His response, in his own words: "Вот такую писюльку… See the full description on the dataset page: https://huggingface.co/datasets/volosati/habr.tabulartext-generation100K<n<1M0 likes369 downloads3mo agoHugging Facevypivshiy /habr Habr dataset Description Summary: Dataset of posts and comments from habr.com, a Russian collaborative blog about IT, computer science and anything related to the Internet. Script: create_habr.py Point of Contact: Ilya Gusev Languages: Russian, English, some programming code. Usage from datasets import load_dataset dataset = load_dataset('IlyaGusev/habr', split="train", streaming=True) for example in dataset: print(example["text_markdown"])… See the full description on the dataset page: https://huggingface.co/datasets/vypivshiy/habr.tabulartext-generation100K<n<1M8 likes285 downloads3mo agoHugging FaceMECHUK /habr Habr dataset Description Summary: Dataset of posts and comments from habr.com, a Russian collaborative blog about IT, computer science and anything related to the Internet. Script: create_habr.py Point of Contact: Ilya Gusev Languages: Russian, English, some programming code. Usage from datasets import load_dataset dataset = load_dataset('IlyaGusev/habr', split="train", streaming=True) for example in dataset: print(example["text_markdown"])… See the full description on the dataset page: https://huggingface.co/datasets/MECHUK/habr.tabulartext-generation100K<n<1M0 likes206 downloads3mo agoHugging Faceits5Q /habr_qna Dataset Card for Habr QnA Dataset Summary This is a dataset of questions and answers scraped from Habr QnA. There are 723430 asked questions with answers, comments and other metadata. Languages The dataset is mostly Russian with source code in different languages. Dataset Structure Data Fields Data fields can be previewed on the dataset card page. Data Splits All 723430 examples are in the train split, there is no validation… See the full description on the dataset page: https://huggingface.co/datasets/its5Q/habr_qna.text-generation100K<n<1M5 likes104 downloads4y agoHugging Face0x7o /GCRL-habr Dataset Card for "GCRL-habr" More Information needed text100K<n<1M0 likes88 downloads3y agoHugging Facegozh /habr_and_wikipedia1gb Russian-English dataset containing articles from Habr and Wikipedia. texttext-generation100K<n<1M1 likes63 downloads3y agoHugging Face