CoolFace
Datasetpublic

Laz4rz/wikipedia_science_chunked_small_rag_512

ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.

sourceHugging Facecc-by-sa-3.0updated 2y agoView on Hugging Face
4likes47downloads
filewikipedia_science_chunked_small_rag.gz542.8 MBdownload
filewikipedia_science_chunked_small_rag.parquet863.6 MBdownload

Laz4rz/wikipedia_science_chunked_small_rag_512 · main · files are served by the source, never re-hosted here