beginner
Datasets
All datasets matching “beginner”beecrowd-beginner-labeled-topics
Beecrowd Beginner Labeled Topics
Dataset Summary
This dataset contains 188 beginner-level programming problems manually curated from the Beecrowd Online Judge, each labeled with one or more introductory programming topics (e.g., loops, conditionals, arrays). It was built to support automated classification of Online Judge (OJ) problems by fundamental programming concepts, since most OJs are organized around competitive-programming categories rather than… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/beecrowd-beginner-labeled-topics.Beginner
Beginner
The beginner tier data can be found in RMDC26_Beginner_Tier.parquet. If read in properly, it should have a shape of 9,230,048 rows x 14 columns and contain lightcurves for 188 unique events. The columns are as follows.
name - The event name. These should go from RMDC26_000001 to RMDC26_000188.
l_deg - The galactic longitude of the event in degrees.
b_deg - The galactic latitude of the event in degrees.
ra_deg - The right ascension of the event in degrees.
dec_deg… See the full description on the dataset page: https://huggingface.co/datasets/RGES-PIT/Beginner.nihongo-dojo-beginner-10k
Nihongo DoJo 初級日本語学習データセット
概要
このデータセットは、日本語学習者向けの合成データセットです。GRPO (Group Relative Policy Optimization) を用いた日本語言語モデルの学習に最適化されています。
データセット統計
総サンプル数: 10,000
言語: 日本語
難易度: 初級(N5-N4相当)
対象: 日本語学習者、言語モデル研究者
タスクタイプ
漢字読み問題 (25%)
例: 「学校」の読み方は? → がっこう
漢字書き問題 (15%)
例: 「みず」を漢字で書いてください → 水
助詞穴埋め問題 (20%)
例: 私_学校_行きます → は、に
助数詞問題 (15%)
例: 3つの本を数えるときの正しい数え方は? → さんさつ
語順並べ替え問題 (10%)
例: 友達と / 公園で / 遊びました → 友達と公園で遊びました
文法問題 (10%)
例: 「今、宿題を_」の_に入る正しい形は? →… See the full description on the dataset page: https://huggingface.co/datasets/akira-sasaki/nihongo-dojo-beginner-10k.python-beginner-distiset
Dataset Card for python-beginner-distiset
This dataset has been created with distilabel.
The pipeline script was uploaded to easily reproduce the dataset:
app.py.
It can be run directly using the CLI:
distilabel pipeline run --script "https://huggingface.co/datasets/nataliaElv/python-beginner-distiset/raw/main/app.py"
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using… See the full description on the dataset page: https://huggingface.co/datasets/nataliaElv/python-beginner-distiset.Beginner-Golf-Knowledgevlm-beginner-dataset
