r-three/eng_latn_code_technical_content
Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_code_technical_content.
Uploading eng_latn_eng_latn_code_technical_content_string_literals subset
Uploading eng_latn_code_technical_content_variable_naming_conventions subset
Uploading eng_latn_eng_latn_code_technical_content_cannonical subset
Uploading eng_latn_code_technical_content_string_literals subset
Uploading eng_latn_eng_latn_code_technical_content_comments_across_languages subset
Uploading eng_latn_eng_latn_code_technical_content_syntax_and_punctuation_variations subset
Uploading eng_latn_code_technical_content_whitespace_variations subset
Uploading eng_latn_eng_latn_code_technical_content_variable_naming_conventions subset
Uploading eng_latn_code_technical_content_cannonical subset
Uploading eng_latn_code_technical_content_syntax_and_punctuation_variations subset
Uploading eng_latn_eng_latn_code_technical_content_buggy_code subset
Uploading eng_latn_eng_latn_code_technical_content_whitespace_variations subset
Upload README.md with huggingface_hub
initial commit
