CoolFace
20 results

data-processing

applied-ai-018 /peacock-data-public-datasets-idc-enfm-dataprocessing BharatGPT Data Curation for English Foundational Model RedPajamaV2 RedPajamaV2 directory contains scripts for processing and filtering of RedPajamaV2 dataset. Overall pipeline is as follows: URL Filtering scripts/utils/url_filtering.py: The first step of filtering is URL based. We use a blocklist of URLs known to contain inappropriate content. This list is taken from blocklistproject. It contains categories for advertisements, gambling, adult content, etc. All… See the full description on the dataset page: https://huggingface.co/datasets/applied-ai-018/peacock-data-public-datasets-idc-enfm-dataprocessing.0 likes327 downloads2y agoHugging Facerenjiepi /easy_5000_data_processingtext10K<n<100K0 likes77 downloads9mo agoHugging Facerenjiepi /medium_5000_data_processingtext1K<n<10K0 likes53 downloads9mo agoHugging FaceNexdata-kr /1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample Description 한국어 시험 문제 구조화 분석·가공 데이터로, 약 150만 개의 시험 문제를 포함하고 있습니다. 문제 유형, 문제, 정답, 해설 등의 정보를 포함하며, 과목은 [초등학교] 국어, 수학, 영어, 사회, 과학; [중학교] 국어, 영어, 수학, 과학, 사회; [고등학교] 국어, 영어, 수학, 물리, 화학, 생물, 역사, 지리로 구성되어 있습니다. 문제 유형에는 객관식, 빈칸 채우기, 참·거짓 문제, 단답형 문제 등이 포함됩니다. 본 데이터셋은 대규모 교과 지식 강화 및 학습 데이터 구축 등의 작업에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/llm/1634?source=Hf.kr Specifications Data content 한국어 K12 시험 문제 Amount 약 150만 개의… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/1.5-million-Korean-Test-Questions-Structured-Analysis-Processing-Data-Sample.textn<1K0 likes49 downloads18d agoHugging Facerenjiepi /medium_5000-data_processing_n100k1text1K<n<10K0 likes45 downloads8mo agoHugging Facerenjiepi /medium_5000_data_processing_fixedtext1K<n<10K0 likes44 downloads9mo agoHugging Face