datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-jp-corpus-v4-ja_warp_pdf
llm-jp-corpus-v4 — ja_warp_pdf
Mirror of the ja/ja_warp_pdf sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_warp_pdf
Files: 513 × jsonl.gz (73.8 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_pdf.inverse-physics-warp-dataset
Inverse Physics Warp Dataset
Synthetic particle trajectories generated with Warp MPM for constitutive learning
and position-only material inference. The versioned recipe and archive inventory
are in recipe.json and manifest.json. A release is complete only when every
archive listed in the manifest is available.
Version 1
Stage 1 contains 800 homogeneous trajectories: five material families (elastic,
Newtonian, non-Newtonian, plasticine, sand), 40 parameter samples… See the full description on the dataset page: https://huggingface.co/datasets/ZhewenZheng/inverse-physics-warp-dataset.warp_Research
Warp Research Dataset (TsFile)
Apache TsFile version of
GotThatData/warp_Research.
Overview
Experimental results from warp-field research, focused on the relationship between warp
factors, energy efficiency, and field characteristics.
Records: ~19,700.
Time period: January 2025.
Features: 15 variables including derived metrics (warp_factor, expansion_rate,
stability_score, max_field_strength, avg_field_strength, energy_efficiency,
efficiency_ratio… See the full description on the dataset page: https://huggingface.co/datasets/THULab/warp_Research.llm-jp-corpus-v4-ja_warp_html
llm-jp-corpus-v4 — ja_warp_html
Mirror of the ja/ja_warp_html sub-corpus of LLM-jp Corpus v4,
built by the LLM-jp Corpus Building WG (NII).
Source: https://gitlab.llm-jp.nii.ac.jp/datasets/llm-jp-corpus-v4
Sub-corpus: ja_warp_html
Files: 45 × jsonl.gz (1.6 GB compressed)
Format: one JSON object per line, with a text key and a meta key
(document id, URL, and other provenance fields).
Directory layout mirrors the upstream repository.
License
CC BY 4.0 — inherited… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/llm-jp-corpus-v4-ja_warp_html.WildKhmerST-WarpedWarpDoc_mergedwarp_Research
Warp Research Dataset
Dataset Description
Dataset Summary
This dataset contains experimental results from warp field research, focusing on the relationship between warp factors, energy efficiency, and field characteristics.
Supported Tasks
Tabular Regression: Predict energy efficiency based on warp field parameters
Time Series Forecasting: Analyze temporal patterns in warp field behavior
Optimization: Identify optimal warp factor configurations for… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/warp_Research.llmjp-warp-htmlllm-jp-corpus-v3のwarp_htmlのうちlevel2フィルタリングされたデータをHFフォーマットに変換し、各データに付与されたURLから元記事のタイトルを取得可能なものについては取得して付与したデータセットです。
ライセンスは元ページに従いCC-BY 4.0とします。
warp-trainingwarped-ifwWARP-benchmark
WARP-benchmark
The benchmark to test logical reasoning and pattern generalisation in large language models through formal SMT constraints.
Overview
We propose a benchmark designed to evaluate a language model's ability to generalise worst-case path constraints for all input sizes. Each example asks a model to generate a formal constraint for a specific target size after being shown examples of constraints for smaller input sizes.
Github Repository
You can find… See the full description on the dataset page: https://huggingface.co/datasets/dannkoh/WARP-benchmark.warp_ms_marcoWarpcast
Dataset Card for BrowseSafe-Bench
Dataset Details
Dataset Description
BrowseSafe-Bench is a comprehensive security benchmark designed to evaluate the robustness of AI browser agents against prompt injection attacks embedded in realistic HTML environments. Unlike prior benchmarks that focus on simple text injections, BrowseSafe-Bench emphasizes environmental realism, incorporating complex HTML structures, diverse attack semantics, and benign "distractor"… See the full description on the dataset page: https://huggingface.co/datasets/Oxhumanode/Warpcast.humanoid-warpat-datassd-warp-2024
