CoolFace
35 results

ED

HuggingFaceFW /fineweb-edu 📚 FineWeb-Edu 1.3 trillion tokens of the finest educational data the 🌐 web has to offer Paper: https://arxiv.org/abs/2406.17557 What is it? 📚 FineWeb-Edu dataset consists of 1.3T tokens and 5.4T tokens (FineWeb-Edu-score-2) of educational web pages filtered from 🍷 FineWeb dataset. This is the 1.3 trillion version. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3-70B-Instruct. We… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu.tabulartext-generation1B<n<10B1.3k likes426k downloads1y agoHugging FaceHelsinki-NLP /fineweb-edu-translated Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.texttranslation1B<n<10B16 likes211k downloads5mo agoHugging Faceeduagarcia-temp /llm_pt_leaderboard_raw_results0 likes103k downloads1y agoHugging Faceeddmpython /dartlab-data DartLab 데이터 종목코드 하나로 읽는 한국 DART + 미국 SEC EDGAR 공시 데이터 Structured Korean (DART) and US (SEC EDGAR) disclosure data, ready as Parquet. 무엇인가요? DartLab이 한국 DART 전자공시와 미국 SEC EDGAR 공시를 종목코드 하나로 비교 가능한 표로 가공해 Parquet으로 올려둔 데이터셋입니다. 한국 전 상장사(약 2,700사)와 미국 주요 상장사(약 1,000사)의 재무제표, 사업보고서 본문, 정형 공시, 주가, 거시지표가 들어 있습니다. 이 데이터셋은 DartLab의 데이터 층입니다. dartlab.Company("005930")을 호출하면 라이브러리가 필요한 parquet을 여기서 자동으로 내려받습니다. 숫자는 원문 그대로 보존합니다(반올림·추정·보간 없음). 코드 없이도 바로 씁니다… See the full description on the dataset page: https://huggingface.co/datasets/eddmpython/dartlab-data.table-question-answering1M<n<10M14 likes90k downloads57m agoHugging Faceairtrain-ai /fineweb-edu-fortified Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it? Fineweb-Edu-Fortified is a dataset derived from Fineweb-Edu by applying exact-match deduplication across the whole dataset and producing an embedding for each row. The number of times the text from each row appears is also included as a count column. The embeddings were produced using TaylorAI/bge-micro Fineweb and… See the full description on the dataset page: https://huggingface.co/datasets/airtrain-ai/fineweb-edu-fortified.tabulartext-generation100M<n<1B65 likes60k downloads2y agoHugging Faceopencsg /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1.text-generation10B<n<100B80 likes54k downloads8mo agoHugging Face