datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
csharp-dotnet-cpt-v0
csharp-dotnet-cpt-v0
A reproducible C# / .NET continued-pretraining (CPT) corpus for Qwen2.5-Coder-1.5B.
HF: https://huggingface.co/datasets/Nottybro/csharp-dotnet-cpt-v0 (private)
Total unique tokens: 695,288,168 (Qwen2.5-Coder-1.5B tokenizer)
Files: 681,426 | Repositories: 68,869
Format: Zstandard-compressed Parquet, schema below.
See DATASET_CARD.md for sources, license policy, filtering, limitations.
See reports/summary.md for full statistics.
cross_code_eval_csharpLCC_csharp
Dataset Card for "LCC_csharp"
More Information needed
csharp_PRsexp_rpt_crosscodeeval-csharp-v4-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_crosscodeeval-csharp-v4-qwen3.5-122b-131k-opencode-traces.csharp-dotnet-reviewing-reasoning-traces
C#/.NET Code Review Reasoning Traces
A curated dataset of 273 high-quality C#/.NET code-review examples with explicit reasoning traces, focused on grounded defect detection, execution/state tracing, falsification, and reduction of false-positive bug reports.
The final corpus contains:
200 positive examples
73 negative examples
73.26% positive / 26.74% negative
The dataset is derived from real open-source C#/.NET projects and includes historical bugs, controlled semantic… See the full description on the dataset page: https://huggingface.co/datasets/kdrapel/csharp-dotnet-reviewing-reasoning-traces.lcc_csharpThis dataset has been modified from the microsoft/LCC_csharp dataset to provide CodeLLaMa with infilling tasks as per the original fill-in-the-middle paper, were the text that needs to be filled in is moved to the end of the dataset, thus taking advantage of the Generative feature of GPT-style models.
glm52-datagen-r11-18-crosscodeeval-csharp-tracesterminal_bench_2_a3_rl_laion_exp_rpt_crosscodeeval_csharp_v4_50_8B_20260828_184838eval-fsr-a1-stack-csharp-swe-r523-tracesdev_set_v2_a3_rl_laion_exp_rpt_crosscodeeval_csharp_v4_50_8B_20260825_141239sc_Csharpterminal_bench_2_r2egym_nl2bash_stack_bugsseq_rl_crosscodeeval_csharp_20260222_044008terminal_bench_2_a1_nemotron_csharp_20260710_043324a1_crosscodeeval_csharpexp_rpt_nemotron-csharp_10k_glm_4.7_traces_jupiterexp_rpt_crosscodeeval-csharpexp_rpt_stack-csharpeval-fsr-a3-crosscodeeval-csharp-swe-r298-tracesc-sharp-coding-dataset
Dataset Card for c-sharp-coding-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmeldrum6/c-sharp-coding-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dmeldrum6/c-sharp-coding-dataset.awesome-csharpнабор для fine-tuning C# expert model
exp_rpt_crosscodeeval-csharp-v4-minimax-m27-131k-tracesexp_rpt_crosscodeeval-csharp_10kterminal_bench_2_a1_crosscodeeval_csharp_20260326_140324terminal_bench_2_a1_nemotron_csharp_20260820_210818tiny-codes-alpaca-csharp
Dataset Card for "tiny-codes-alpaca-csharp"
More Information needed
a1_stack_csharpterminal_bench_2_a1_crosscodeeval_csharp_20260323_193422eval-fsr-a1-nemotron-csharp-swe-r352-rf0711-tracesterminal_bench_2_a1_nemotron_csharp_20260406_035518
