datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TILO.RA_CODER_Dataset
TILO.RA CODER Dataset
Объединённый русско-английский датасет для обучения и поиска по коду.
Формат — пары question / code: вопрос на естественном языке → готовый код-ответ.
Датасет собран для локального ассистента TILO.RA CODER — офлайн-помощника по программированию
Скачать по ссылке
https://github.com/thetemirbolatov/TILO.RA_CODER_Dataset/releases/download/v1.0.0/tilora_knowledge_merged.jsonl
Состав
Источник
Язык
Записей
English coding… See the full description on the dataset page: https://huggingface.co/datasets/thetemirbolatov/TILO.RA_CODER_Dataset.vscode-bug-feature-triage
VS Code Bug vs Feature Request Triage
Dataset summary
1,993 prepared issue records from public microsoft/vscode issues, reduced to one binary task: classify the issue text as bug or feature-request. The splits are a frozen temporal holdout (80/10/10 by created_at within each class, seed 42) used by the GitHub Triage SLM Fine-Tuning Benchmark to compare fine-tuned small models against their base checkpoints on the same test set. Each record carries cleaned issue… See the full description on the dataset page: https://huggingface.co/datasets/Tilakoid/vscode-bug-feature-triage.tilt-hs
Dataset Card for TiLt-HS
TiLt-HS (Tests in Lithuanian, High School) is a dataset of multiple-choice question tests that are used to assess the knowledge of high school students in several academic areas.
Table of Content
Dataset Details
Dataset description
Uses
Direct Use
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Data Collection and Processing
AnnotationsPersonal and Sensitive Information
Bias, Risks… See the full description on the dataset page: https://huggingface.co/datasets/Jekaterina/tilt-hs.Demo-Datasetreports-tech-qna-chroma-logsmahjong-winning-tiles
Mahjong Winning Tiles (胡啥牌) Benchmark
Introduction
Mahjong Winning Tiles is a benchmark dataset designed to evaluate the reasoning capability of large language models (LLMs). It focuses on assessing logical thinking, pattern recognition, and decision-making by challenging models to calculate possible winning tiles needed to complete a Mahjong hand of 13 tiles.
Background: Mahjong Rules
Mahjong is a traditional 4-player tile game originating from China and… See the full description on the dataset page: https://huggingface.co/datasets/sileixu/mahjong-winning-tiles.document-qna-chroma-anyscale-logshealth_summarizetldrCancerResearchPaperkaz-morphology-sample
kaz-morphology-sample
Қазақ сөздерінің морфологиялық үлгісі · Образец морфологической разметки казахских слов · Kazakh word morphology sample
Қазақша · Русский · English
Қазақша
kaz-morphology-sample — қазақ сөздерінің морфологиялық талдауы бар 9.3 МБ деректер жинағы. Жинақта 10 000 жазба және 7 сөз табы қамтылған; оны морфологиялық талдау, леммалау және POS-tagging үлгілерін тексеруге қолдануға болады.
Деректер құрамы
Файл
Жазба… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kaz-morphology-sample.ai4se_userskaz-pos-corpus
kaz-pos-corpus
Қазақ тілінің POS корпусы · Корпус частей речи казахского языка · Kazakh POS corpus
Қазақша · Русский · English
Қазақша
kaz-pos-corpus — қолмен POS белгісі қойылған қазақ мәтіндерінің 1.9 МБ корпусы. Онда 1 200 сөйлем, 4 107 сөз және 18 белгі класы бар; корпус token classification модельдерін оқыту мен бағалауға арналған.
Құрамы мен нәтижелері
Көрсеткіш
Мәні
Сөйлем
1 200
Сөз
4 107
POS класы
18
Сөз деңгейіндегі… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kaz-pos-corpus.agent_usersTil-Corpus-exp078
Til-Corpus-exp078
Қазақ тілі басым оқыту корпусы · Корпус для обучения с преобладанием казахского языка · Kazakh-dominant training corpus
Қазақша · Русский · English
Қазақша
Til-Corpus-exp078 — тілдік модельдерді алдын ала оқытуға дайындалған, қазақ тілі басым корпус. Репозиторий көлемі — 139.07 ГБ; оқытуға дайын бөлігінде 32.9M жол және Til-Tokenizer-128k бойынша 7.21B токен бар.
Құрамы
Дерек сапасы оқу соңына қарай арта түсетін ретпен жиналған.… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/Til-Corpus-exp078.tilledmachine-failure-logsinsurancecharge_logskazakh-terms-8k
kazakh-terms-8k
Қазақ терминдерін морфологиялық сегменттеу жинағы · Датасет морфологической сегментации казахских терминов · Kazakh term morphological segmentation dataset
Қазақша · Русский · English
Қазақша
kazakh-terms-8k — қазақ терминдерін морфологиялық сегменттеу модельдерін оқытуға және бағалауға арналған 8 000 белгіленген термин жинағы. Репозиторий көлемі — 2.3 МБ.
Құрамы
Негізгі кесте 8k_labeled_terms.csv файлында сақталған. Репозиторийде… See the full description on the dataset page: https://huggingface.co/datasets/TilQazyna/kazakh-terms-8k.dataset1TildeMODEL.en-esdataset-txttilt-pro
Dataset Card for TiLt-Pro
TiLt-Pro (Tests in Lithuanian, Professional) is a dataset of multiple-choice question tests that are used to assess the work-related knowledge of workers of multiple professions.
Table of Content
Dataset Details
Dataset description
Uses
Direct Use
Dataset Structure
Data Instances
Data Fields
Data Splits
Dataset Creation
Curation Rationale
Source Data
Data Collection and Processing
AnnotationsPersonal and Sensitive Information
Bias, Risks… See the full description on the dataset page: https://huggingface.co/datasets/Jekaterina/tilt-pro.term-deposit-logsnassau-parcels-tilesshikokuchatbotdatafull3.2datav1
