zeio/batch
Dataset card for batch Dataset summary This dataset contains threads parsed from the /b/ board of 2ch archive. See dataset viewer at the derivative repo. Examples of the dataset reading and usage are provided in this colab notebook. Dataset structure The dataset is represented in three formats - compressed, uncompressed and spoken: uncompressed representation is the default and simplest one - in this form the content of dataset is organised… See the full description on the dataset page: https://huggingface.co/datasets/zeio/batch.
feat(example): added link to the dataset usage examples in readme
fix(loader): updated the format of entries in the loader script
fix(loader): decreased n-items
fix(loader): updated spoken data url
fix(loader): updated spoken data logging
fix(loader): added spoken data loading error logging
fix(loader): updated "speech" field setting
fix(loader): fixed the loading script
feat(loader): updated last batch folder pointer in the index, added loading script
feat(data): added the remaining batches
feat(batches): added more batches
feat(batch): added 5 more batches
feat(batch): add next batch
feat(batch): added next batch
feat(batch): add next batch
feat(batch): add next batch
feat(batch): added next batch
feat(batch): added next batch
feat(batch): added next batch
feat(batch): added next batch
feat(batch): added next batch
next(batch): added next batch
next(batch): added next batch
feat(batch): added next batch
fix(next): uncommented python module call
feat(scripts): added scripts for generating next batch
feat(batch): added next batch
feat(batch): added next batch
feat(batch): added next batch
fix(readme): added title to the dataset instance example
feat(batch): add next batch
feat(4th batch): added 4th batch
feat(auto): added link to the derivative repo
feat(nsfw): added nswf tag
fix(links): fixed links in readme
fix(links): fixed links in readme
fix(readme): fixed links in readme
fix(config): fixed config descriptions in readme
feat(readme): added config description, added logo
feat(readme): added dataset card
feat(3rd batch): added the 3rd batch
feat(2nd partition): generated the second dataset partition
feat(compressed): added folder with compressed threads
fix(posts): deleted empty posts, updated index structure by including folder names
feat(files): rearranged files to keep them in an original form as raw test
feat(data): generated the second batch
feat(threads): restarted data pulling, generated entries for first 40 threads
fix(spaces): fixed issues with spaces
feat(batch): added first 2000 samples from pages 1746-1749
initial commit
