CoolFace
Datasetpublic

Laz4rz/wikipedia_science_chunked_small_rag_512

ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.

sourceHugging Facecc-by-sa-3.0updated 2y agoView on Hugging Face
4likes47downloads
7 commits on main
d705d762y ago

Update README.md

Laz4rz
3177a342y ago

Update README.md

Laz4rz
0f2a6b82y ago

Update README.md

Laz4rz
db179562y ago

Update README.md

Laz4rz
bac7c1b2y ago

Upload wikipedia_science_chunked_small_rag.parquet

Laz4rz
030757a2y ago

Upload wikipedia_science_chunked_small_rag.gz

Laz4rz
530334b2y ago

initial commit

Laz4rz