CoolFace
Datasetpublic

Laz4rz/wikipedia_science_chunked_small_rag_512

ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.

sourceHugging Facecc-by-sa-3.0updated 2y agoView on Hugging Face
4likes40downloads
Dataset Card

ScienceWikiSmallChunk

Processed version of millawell/wikipediafieldof_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix.

There is also 256 tokens dataset available: Laz4rz/wikipediasciencechunkedsmallrag_256

If you wish to prepare some other chunk length:

  1. 1.use millawell/wikipediafieldof_science
  2. 2.adapt chunker function:
def chunker_clean(results, example, length=512, approx_token=3, prefix=""):
    if len(results) == 0:
        regex_pattern = r'[\n\s]*\n[\n\s]*'
        example = re.sub(regex_pattern, " ", example).strip().replace(prefix, "")
    chunk_length = length * approx_token
    if len(example) > chunk_length:
        first = example[:chunk_length]
        chunk = ".".join(first.split(".")[:-1])
        if len(chunk) == 0:
            chunk = first
        rest = example[len(chunk)+1:]
        results.append(prefix+chunk.strip())
        if len(rest) > chunk_length:
            chunker_clean(results, rest.strip(), length=length, approx_token=approx_token, prefix=prefix)
        else:
            results.append(prefix+rest.strip())
    else:
        results.append(prefix+example.strip())
    return results