Laz4rz/wikipedia_science_chunked_small_rag_512
ScienceWikiSmallChunk Processed version of millawell/wikipedia_field_of_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix. There is also 256 tokens dataset available: Laz4rz/wikipedia_science_chunked_small_rag_256 If you wish to prepare some other chunk length: use… See the full description on the dataset page: https://huggingface.co/datasets/Laz4rz/wikipedia_science_chunked_small_rag_512.
ScienceWikiSmallChunk
Processed version of millawell/wikipediafieldof_science, prepared to be used in small context length RAG systems. Chunk length is tokenizer dependent, but each chunk should be around 512 tokens. Longer wikipedia pages have been split into smaller entries, with title added as a prefix.
There is also 256 tokens dataset available: Laz4rz/wikipediasciencechunkedsmallrag_256
If you wish to prepare some other chunk length:
- use millawell/wikipediafieldof_science
- adapt chunker function:
def chunker_clean(results, example, length=512, approx_token=3, prefix=""):
if len(results) == 0:
regex_pattern = r'[\n\s]*\n[\n\s]*'
example = re.sub(regex_pattern, " ", example).strip().replace(prefix, "")
chunk_length = length * approx_token
if len(example) > chunk_length:
first = example[:chunk_length]
chunk = ".".join(first.split(".")[:-1])
if len(chunk) == 0:
chunk = first
rest = example[len(chunk)+1:]
results.append(prefix+chunk.strip())
if len(rest) > chunk_length:
chunker_clean(results, rest.strip(), length=length, approx_token=approx_token, prefix=prefix)
else:
results.append(prefix+rest.strip())
else:
results.append(prefix+example.strip())
return results