A.B.I.R
Datasets
All datasets matching “A.B.I.R”english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.tamilwikipediadatasetannotations_creators:
found
language:
Tamil
language_creators:
found
license: []
multilinguality:
multilingual
pretty_name: tamilwikipediadataset
size_categories:
100K<n<1M
source_datasets: []
tags: []
task_categories:
summarization
task_ids: []
abir177m-pretrain-balanced20-ezhijaru
abir177m pretrain mix — balanced20 en/zh/hi/ja/ru
Frozen packed-token shards for reproducible abir177m GPT-2–style pretraining.
Languages: 20% each en, zh, hi, ja, ru
Source streams: FineWeb (en) + FineWeb-2 (zh/hi/ja/ru)
Tokenizer: mistralai/Mistral-Nemo-Base-2407
Packing: 2048-token causal LM blocks (input_ids, labels identical)
Target budget: 3.55B tokens (1,733k sequences)
See meta.json for exact mixture + dataset map + seed.
french_book_reviews
Dataset Card for French book reviews
I-Dataset Summary
The majority of review datasets are in English. There are datasets in other languages, but not many. Through this work, I would like to enrich the datasets in the French language(my mother tongue with Arabic).The data was retrieved from two French websites: Babelio and Critiques LibresLike Wikipedia, these two French sites are made possible by the contributions of volunteers who use the Internet to share their… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/french_book_reviews.waltoncolor
