datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_stock_kline
Languages: 简体中文 · English
dojo_stock_kline — Stock Daily Bars
Overview
Daily OHLCV history for US, CN, and HK equities, including cumulative adjustment factors and dividend/split flags.
Files
File
Description
data.parquet
Market-wide daily bars
Key Fields
Field
Description
symbol
Join key
kline_t
Bar interval; snapshots use "1D" for daily bars
bar_time
Bar timestamp (trade date)
open / high / low /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_stock_kline.dojo_quote
Languages: 简体中文 · English
dojo_quote — Latest Quote Snapshot
Overview
Point-in-time snapshot for the most recent trading session per symbol: price, change, volume, market cap, valuation ratios, and related fields. Not a historical time series.
Files
File
Description
data.parquet
Full latest-quote table
Key Fields
Field
Description
symbol
Join key to dojo_stock_info
name
Display name from quote feed… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_quote.dojo_fin_indicators
Languages: 简体中文 · English
dojo_fin_indicators — Financial Metrics
Overview
Multi-period financial statement derivatives per symbol: income statement, balance sheet, cash flow, and industry-specific metrics (banks, insurers, brokers, etc.). Supports quarterly, cumulative, and other report_type values.
Files
File
Description
data.parquet
Full financial metrics (wide table, 100+ columns)
Key Fields (common)
Field… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_fin_indicators.dojo_forex_kline
Languages: 简体中文 · English
dojo_forex_kline — FX Daily Bars
Overview
Daily OHLC and amplitude for major currency pairs. Used to convert revenue, profit, and other filing amounts into a single currency when report currency and listing/analysis currency differ.
Typical case: regional revenue in HKD in dojo_main_income while analysis targets USD — apply HKDUSD (or equivalent) at the report date.
Intended Use: Cross-Currency Revenue Breakdown… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_forex_kline.dojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_info.dojo_main_income
Languages: 简体中文 · English
dojo_main_income — Revenue Breakdown
Overview
Segment-level main business revenue from listed companies, by industry, product, and region, with amounts and mix ratios. Corresponds to “main business by segment” notes in filings.
Files
File
Description
data.parquet
Full revenue breakdown detail
Key Fields
Field
Description
symbol
Stock symbol
security_name
Company name… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_main_income.dojo-bench-mini
Dojo-Bench-Mini
Dojo-Bench-Mini is a small public bench containing tasks for running computer use agents against "mocked" productivity software and games. These include:
Linear
LinkedIn
Tic-Tac-Toe
2048
For full details on running this benchmark, check-out:
Dojo, an environment hub for computer use agents
Docs
Running an evaluation
Forecast-Dojo
Forecast Dojo
Forecast Dojo is a longitudinal forecasting benchmark comparing language-model
probability forecasts with contemporaneous prediction-market crowd judgments.
This export contains 1,301 questions, 60 model
runs, 193,235 checkpoint forecasts, 95,555
extractable submitted belief-notebook rows, and 544,172
checkpoint-level tool summaries.
Snapshot
Source releases: August 2026 and September 10, 2026.
Question splits: 1,053 train and 248 evaluation… See the full description on the dataset page: https://huggingface.co/datasets/FinEredium1/Forecast-Dojo.dojo_sector_precomputed
Languages: 简体中文 · English
dojo_sector_precomputed — Precomputed Sector Analytics
Overview
Derived sector analytics: L3 constituent snapshots, daily cap-weighted sector index levels, and per-constituent daily returns. Built offline from taxonomy, mappings, quotes, and stock K-lines.
Files
File
Description
manifest.json
Generation metadata: version, window start, row counts, latest trade dates
constituents.parquet
L3 constituent… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/dojo_sector_precomputed.dojo_stock_kline
Languages: 简体中文 · English
dojo_stock_kline — Stock Daily Bars
Overview
Daily OHLCV history for US, CN, and HK equities, including cumulative adjustment factors and dividend/split flags.
Files
File
Description
data.parquet
Market-wide daily bars
Key Fields
Field
Description
symbol
Join key
kline_t
Bar interval; snapshots use "1D" for daily bars
bar_time
Bar timestamp (trade date)
open / high / low /… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/dojo_stock_kline.Toxic_dataset
Russian Multi-label Toxic Comments Dataset
Описание датасета
Датасет для мульти-лейбл классификации токсичных комментариев на русском языке. Каждый текст размечен по трем категориям:
Profanity (ненормативная лексика) - наличие нецензурной лексики
Threat (угрозы) - содержание угроз в адрес других пользователей
Illegal (незаконные запросы) - запросы на нарушение закона (например, создание оружия, наркотиков)
Структура данных
Датасет содержит три… See the full description on the dataset page: https://huggingface.co/datasets/Doji070/Toxic_dataset.newsgroupsdojo_sector_info
Languages: 简体中文 · English
dojo_sector_info — Sector Taxonomy
Overview
Three-level sector taxonomy (L1 industry → L2 chain → L3 sub-sector) with bilingual Chinese/English names and definitions.
Files
File
Description
data.parquet
Taxonomy tree (one L1 row each; L2/L3 nested in children)
Key Fields
Field
Description
id
L1 sector ID
name / name_alias
L1 English name / Chinese alias
description /… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/dojo_sector_info.banking11jailbreak-dojo-corpusdojo-bench-customer-colossus
Dojo-Bench-Colossus
Dojo-Bench-Colossus is a list of tasks for running computer use agents against "mocked" productivity software for Customer Colossus.
For full details on running these tasks, check-out:
Dojo, an environment hub for computer use agents
Docs
Running an evaluation
t1nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 2,418 サンプル
validation: 302 サンプル
test: 303 サンプル
総サンプル数: 3,023
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/
サンプルデータ
{
"instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。",
"input": "「究」の訓読みは?",
"think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.newsgroups7_balancedbanking_intentnihongo-dojo-small
nihongo-dojo-small
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 7,600 サンプル
validation: 950 サンプル
test: 950 サンプル
総サンプル数: 9,500
ソース
生成元: ./datasets/nihongo-dojo-small/
サンプルデータ
{
"instruction": "次のひらがなを漢字で書いてください。",
"input": "「みず」を漢字で書くと?",
"output": "<think>\n「みず」は「水」と書きます。意味: water\n</think>\n<answer>水</answer>",
"group_id": 0,
"task_idx": 0,
"task_type": "kanji_writing",
"difficulty": "beginner",
"metadata":… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-small.jon_testsst2_balancednihongo-dojo-grades1-2-3-kanji_reading
nihongo-dojo-grades1-2-3-kanji_reading
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 846 サンプル
validation: 105 サンプル
test: 107 サンプル
総サンプル数: 1,058
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-kanji_reading/
サンプルデータ
{
"instruction": "次の漢字の音読み(おんよみ)をカタカナで答えてください。",
"input": "「代」の音読みは?",
"output": "タイ",
"thinking": "この漢字は「代」です。 小学3年生で習う漢字です。 意味は「substitute」などです。 音読み(おんよみ)は中国から伝わった読み方です。 この漢字の音読みは「タイ」です。",
"answer": "タイ"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-kanji_reading.jemail-doj-embeddings
jemail-doj-embeddings
Dataset Description
This dataset contains text embeddings for documents from the Doj source related to the Jeffrey Epstein case. Each embedding is paired with its original text content, making it suitable for semantic search, document analysis, and retrieval-augmented generation (RAG) applications.
Source Information
Department of Justice (DOJ) document releases related to the Jeffrey Epstein case. These documents include police reports… See the full description on the dataset page: https://huggingface.co/datasets/567-labs/jemail-doj-embeddings.jmail-doj-embeddingssquare_balancedjailbreak-dojo-leaderboardjmail-dojepstein-doj
