datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cs2_data_hf_expertchimera-cs2
Chimera CS2 Dataset
Labeled Counter-Strike 2 screenshots for vision-language model training.
Each sample has:
A CS2 gameplay screenshot
Ground truth JSON with game_state, analysis, and advice
Structure
screenshots/ # PNG/JPG images
labels/ # Matching JSON files (same stem name)
manifest.jsonl # Data provenance tracking
Usage
from datasets import load_dataset
ds = load_dataset("skkwowee/chimera-cs2")
Stats
Labels: 5309… See the full description on the dataset page: https://huggingface.co/datasets/skkwowee/chimera-cs2.CS2CD.Counter-Strike_2_Cheat_Detection
Counter Strike 2 Cheat Detection Dataset
Overview
The CS2CD (Counter-Strike 2 Cheat Detection) dataset is an anonymised dataset comprised of Counter-Strike 2(CS2) gameplay at a variety of skill-levels with cheater annotations. This dataset contains 478 CS2 matches with no cheater present, and 317 matches CS2 matches with at least one cheater present.
Dataset structure
The dataset is partitioned into data with at least one cheater present, and data with no… See the full description on the dataset page: https://huggingface.co/datasets/CS2CD/CS2CD.Counter-Strike_2_Cheat_Detection.Computer-Science
מאגר הנתונים CS26 HIT — מדעי המחשב
מאגר זה משמש לאחסון מרכזי של נתוני לימוד ומשאבים אקדמיים עבור סטודנטים למדעי המחשב במכון הטכנולוגי חולון. המידע המצוי כאן מונגש בצורה נוחה באמצעות פורטל גישה נפרד המאפשר ניווט ויזואלי וחיפוש יעיל בתוך התיקיות השונות.
קישורים וגישה
ניתן להשתמש בפורטל בכתובת https://cs26-cs26-portal.hf.space/
תנאי שימוש וזכויות יוצרים
כל חומרי הלימוד והתכנים המופיעים במאגר זה פתוחים וחופשיים לשימוש לצורכי למידה בלבד. ניתן לקחת את… See the full description on the dataset page: https://huggingface.co/datasets/CS26/Computer-Science.cs2-demo-archive
CS2 Demo to Dataset — Pipeline Output Samples
Sample archives produced by the open-source
cs2-demo-to-dataset
pipeline, which converts a single CS2 .dem replay file into per-round,
per-player first-person video aligned to tick-level state, input and event
tables.
This release is not a dataset contribution. The point of the upload is to
demonstrate that the pipeline produces a coherent, reproducible archive
format. Please see the
GitHub repository for the
recorder code… See the full description on the dataset page: https://huggingface.co/datasets/Vasy7777/cs2-demo-archive.cs2demos177 cs2 matches from hltv.org
cs2-highlights
Dataset Card for Counter-Strike 2 Highlight Clips
Dataset Summary
This dataset contains 8,369 high-quality gameplay highlight clips primarily from Counter-Strike 2, with a small portion from Counter-Strike: Global Offensive. The clips focus on key gameplay moments such as kills, bomb interactions, and grenade usage. The clips are collected from competitive platforms like Faceit and in-game competitive modes (Premier, Matchmaking) across various skill levels, making it… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cs2-highlights.cs2_data_hdf5cs2_videoscs2-demoscs2-anticheat-rawcs2-v3-prompt-comparison-7-examples-with-multiaction-gemini35
CS2 V3 七案例 Gemini 3.5 Flash 最终结果对比 / Seven-case Gemini 3.5 Flash Comparison
本 README 展示 Gemini 3.5 Flash 对同一批 7 个视频的最终打标结果:每个案例先显示视频,再用左右两列并排展示两轮和七轮的完整 English JSON 与中文 JSON;内容直接展开,字号保持较小以便对照。
This README shows Gemini 3.5 Flash final labels for the same 7 videos. Each case places the video first, then displays complete English and Chinese JSON side by side: two-round on the left and seven-round on the right.
两轮与七轮的 API 输入详情通过顶部索引查看;中文侧保持与英文 JSON 相同的键、时间边界、数组长度和 Action 标签。
API… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-prompt-comparison-7-examples-with-multiaction-gemini35.CS2-10k
This dataset is no longer available due to a takedown notice from Valve.
CS2-10k: A Large-Scale Egocentric Counter-Strike 2 Dataset
CS2-10k was a large-scale egocentric gameplay dataset built from professional CS2 matches. It contained 600,000+ player-round videos spanning 10,000+ hours of first-person footage, paired with per-frame annotations covering keyboard state, mouse movement, and 3D player trajectory.
All data, the interactive viewer, and associated files have been… See the full description on the dataset page: https://huggingface.co/datasets/RekaAI/CS2-10k.CS2-HUD-OCR-Crops
CS2 HUD OCR Crops
Per-region HUD crops sliced from three Counter-Strike 2 match recordings,
labelled where possible from the demo file's parse_ticks state. Built
to train a specialist CRNN that replaces the EasyOCR killfeed reader
(currently ~6 s p95 on CPU) with a sub-30 ms specialist.
Source
Three matches by the same POV player (farouqqq), recorded in CS2's
built-in DVR + the corresponding .dem files:
sample
map
dem
rounds
resolution
fps
sample1
Ancient… See the full description on the dataset page: https://huggingface.co/datasets/ybashir/CS2-HUD-OCR-Crops.cs2-yolocs2-v3-prompt-comparison-5-examples-with-multiaction
CS2 V3 七案例最终结果对比 / Seven-case Final Result Comparison
本 README 只展示模型的最终输出结果,不展示 API 请求 prompt。每个案例使用同一个视频:左列是两轮结果,右列是七轮结果。结果直接展开,无需点击折叠。英文和中文同时提供,字号缩小以便并排查看。
This README shows only the model's final output results, not the API request prompts. Each case uses the same video: the two-round result is on the left and the seven-round result is on the right. Results are visible directly with no expandable sections. English and Chinese are shown together in a small type size for… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-prompt-comparison-5-examples-with-multiaction.cs2-2B-activationscs2-demo-datasetcs2-pov-hltv-artifacts
CS2 single-POV HLTV artifacts (heads, latents, labels)
The irreplaceable ~183 MB extracted from a 163 GB local render workspace, so the
bulk could be deleted. Companion to
cs2-hwm-scale300 and
cs2-10k-vjepa2-latents-300.
These come from rendered GOTV POV of 19 pro matches. Unlike CS2-10k, these
renders contain HUD and radar pixels, which is what makes the HUD digit reader
and the economy/value heads possible at all - CS2-10k cannot support that line
of work.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/cbctr/cs2-pov-hltv-artifacts.cs2_dataset_render
OpenCS2 — POV Renders
Browse with the OpenCS2 Viewer — every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each row is
≤1 minute of one player's perspective; ten POVs per round share the same tick clock.
Per chunk:
Video — 1280×720 @ 32 fps, near-lossless H.264.
Audio — per-player stereo, mixed from that player's position and orientation.
Inputs — every tick: keys… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/cs2_dataset_render.cs2-v3-prompt-labeling-project
CS2 V3 Prompt Labeling Project
这是 CS2 第一人称视频打标 Prompt、批处理脚本、样例结果和 Sonnet/Gemini 对比网页的可迁移项目。仓库按“拿到另一台电脑即可继续处理”的方式整理;原始数据集、模型权重、API 请求体和凭据不在仓库中。
固定标注约定
输入视频:81 帧、16 FPS,帧 0 是 conditioning frame。
内容窗口固定为 [1,17)、[17,33)、[33,49)、[49,65)、[65,81),每个窗口 16 帧、约 1 秒。
每个 chunk 只写一个高层 Move 和一个高层 Camera;跨 chunk 的同一个动作复用同一个 Action ID。
当前生产输出默认只保留顶层 english_json,chunk 证据使用完整的 16 帧 MP4(不使用四联图)。
Prompt 要求按观看者屏幕的左/右描述,并按 chunk 内时间顺序组织详细视觉叙述;HUD/overlay 不写入结果。
目录… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-prompt-labeling-project.cs2_data_hdf5_expert_finalcs-230-news-v3CS2CD.Counter-Strike_2_Cheat_Detection
Counter Strike 2 Cheat Detection Dataset
Overview
The CS2CD (Counter-Strike 2 Cheat Detection) dataset is an anonymised dataset comprised of Counter-Strike 2(CS2) gameplay at a variety of skill-levels with cheater annotations. This dataset contains 478 CS2 matches with no cheater present, and 317 matches CS2 matches with at least one cheater present.
Dataset structure
The dataset is partitioned into data with at least one cheater present, and data with no… See the full description on the dataset page: https://huggingface.co/datasets/adhitsamonkar/CS2CD.Counter-Strike_2_Cheat_Detection.cs2-v3-prompt-comparison-5-examples
CS2 V3 五案例最终结果对比 / Five-case Final Result Comparison
本 README 只展示模型最终输出结果,不展示发送给 API 的 Prompt。每个案例使用同一视频:左侧为两轮结果,右侧为七轮结果。结果无需展开即可直接查看,英文原文和中文翻译同时提供并排对比。
This README shows only the model's final output results, not the API request prompts. Each case uses the same video: the two-round result is on the left and the seven-round result is on the right. Results are visible directly with no expandable sections. English and Chinese are shown together for side-by-side comparison.
Case 1… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-prompt-comparison-5-examples.musdb18-processed
MUSDB18 Active Stems Dataset - CS229 Project
This dataset contains active stem segments extracted from the MUSDB18 dataset for the Stanford CS229 Machine Learning course project on audio source separation.
Dataset Description
This is a processed version of the MUSDB18 dataset containing only the active segments of each stem (drums, bass, vocals, accompaniment, and mixture), designed to improve training efficiency for music source separation models.
Key… See the full description on the dataset page: https://huggingface.co/datasets/cs229-audio-ml-project/musdb18-processed.cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.cs2_dataset_render_part2
OpenCS2 — POV Renders
Browse with the OpenCS2 Viewer — every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each row is
≤1 minute of one player's perspective; ten POVs per round share the same tick clock.
Per chunk:
Video — 1280×720 @ 32 fps, near-lossless H.264.
Audio — per-player stereo, mixed from that player's position and orientation.
Inputs — every tick:… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/cs2_dataset_render_part2.cs2-v3-labeling-examples-and-api-prompt
CS2 V3 中英文视频标注完整示例
这是一个中英文并列的完整样例:同一份总视频、五个连续 chunk 视频,以及对应的 English JSON 和中文翻译 JSON。README 直接就是示例页面;可解析的完整文件也保留在仓库根目录。
私有仓库提示:如果页面或视频返回 404,请先在当前浏览器登录 Hugging Face。私有仓库对未登录请求会故意表现为 404。下面所有播放器都使用稳定的 /resolve/main/... 直链,而不是相对 assets/... 路径;播放器地址不加 ?download=true,因为该参数会把响应设为 attachment 下载而不是 inline 播放。
样例信息
sample_id: opencs2-inspect-move_camera-m2392614-de_inferno-r13-p09-f000051
map: de_inferno;action: inspect
full video: 81 frames, 16.0 FPS, 832x480
frame 0… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-v3-labeling-examples-and-api-prompt.cs2-10k-vjepa2-latents-300
CS2-10k V-JEPA2 latent cache (300 matches)
Frozen facebook/vjepa2-vitl-fpc64-256 embeddings of single-POV Counter-Strike 2
gameplay, from 300 matches of RekaAI/CS2-10k
(mirage + dust2). This is the training substrate for the
scale300 hierarchical world model.
Rebuilding it from source takes ~82 hours of wall-clock (elapsed_min 4942.7),
almost all of it network-bound, which is why it is published here.
Contents
file
shape / rows
notes
latents.npy
[6,942… See the full description on the dataset page: https://huggingface.co/datasets/cbctr/cs2-10k-vjepa2-latents-300.
