usenet
Datasets
All datasets matching “usenet”UsenetArchiveIT
Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset contains Usenet posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing.
The only preprocessing conducted on the text was the removal of two conversations in which the VBS source code of the malicious script "ILOVEYOU" was present as it was shared by two users for didactical… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/UsenetArchiveIT.UsenetArchiveIT-conversations
Conversational Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset is a filtered version from the Usenet dataset that contains posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing. All posts with more the one message has been grouped in conversations
This dataset contributes to the mii-community project, aimed at advancing the… See the full description on the dataset page: https://huggingface.co/datasets/mii-community/UsenetArchiveIT-conversations.Usenet-Corpus-1980-2013-Threaded
Usenet Corpus 1980–2013 — Threaded
A large, thread-reconstructed corpus of Usenet posts spanning 1980–2013, with
every post linked into its conversation via thread_id, thread_position, and
thread_depth. This is the threaded companion to the Usenet Corpus 1980–2013
(cleaned) dataset: identical post content, plus conversation structure.
▶ Free preview — no gating. Browse a showcase sample of complete
reconstructed conversations at
Usenet-Corpus-1980-2013-Threaded-Samples
— no… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded.Usenet-Corpus-1980-2013-Full
Usenet Corpus 1980–2013
A large, cleaned corpus of Usenet posts spanning 1980–2013 — long-form,
pre-web conversational text across the full newsgroup hierarchy. A thread-
reconstructed companion (Usenet Corpus 1980–2013 — Threaded) adds conversation
structure over identical content.
▶ Free preview — no gating. Browse a showcase sample of complete
conversations at
Usenet-Corpus-1980-2013-Full-Samples
— no access request needed.
▶ License the full corpus. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full.Usenet-Corpus-1980-2013-Full-Samples
Usenet Corpus 1980–2013 — Full (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 (cleaned)
dataset — long-form, pre-web Usenet posts. This repo is a free preview; the full,
commercially-licensed corpus (405.8M posts, 102.5B tokens) is at:
Full cleaned dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full
Threaded companion: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Full-Samples.Usenet-Corpus-1980-2013-Threaded-Samples
Usenet Corpus 1980–2013 — Threaded (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded
dataset: Usenet posts reconstructed into conversations via thread_id,
thread_position, and thread_depth. This repo is a free preview; the full,
commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at:
Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded
Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.
