datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zoonomia-v1-v4_ccre_noexon_enhancer
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer
A curated enhancer training set for issue
#326 — a de-contaminated
derivation of the v4 ccre_non_promoter arm of
bolinas-dna/zoonomia-v1-v1,
built by the
snakemake/zoonomia_projection_dataset pipeline at commit
6b320c268547.
Provenance
This subset is v4_ccre_noexon further restricted to enhancer-dominant windows (dELS+pELS basepair coverage ≥ the other non-PLS cCRE classes), population-matching the val_enhancer… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer.gpn-star-p-uniform-v1-enhancer-arm-a
marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.phylop-uniform-v1-enhancer-arm-a
marin-dna/phylop-uniform-v1-enhancer-arm-a
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.functional-enhancer
marin-dna/functional-enhancer
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and official version-matched UCSC hg38-to-target liftOver chains.
This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case.
Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus.
For the 28 non-mammalian targets, the stable… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/functional-enhancer.vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the full_window policy.
The policy projects the complete 255 bp human window, applies the established 128--512 bp compatible-fragment gate, and resizes around the accepted target-span midpoint. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-fullwindow-ccre-enhancer-centered.zoonomia-v1-v4_ccre_noexon_enhancer-order
bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order
The bolinas-dna/zoonomia-v1-v4_ccre_noexon_enhancer cross-mammal
training set, restricted to a species cohort: one representative species per NCBI order — 19 deeply-diverged placental mammals (every pair separated by ~tens of millions of years), versus the implicit-default 108 family-deduplicated species. A strict subset of the family set, so it reuses the v1 cross-mammal projection unchanged (no re-halLiftover).
Same human… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_noexon_enhancer-order.vertebrate-v1-issue473-center1-ccre-enhancer-centered
marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered
Review status: draft generated for issue #473 review before upload.
Human-anchored 255 bp vertebrate sequences for the
ccre_enhancer_centered cohort under the center_1 policy.
The policy projects the exact central human nucleotide, requires one target locus, and emits the 255 bp target window centered on that mapped nucleotide. Human anchors come from the fixed exp351 ENCODE dELS/pELS-centered catalog after… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-issue473-center1-ccre-enhancer-centered.oeis-enhanced
Attribution
This dataset was generated using data from the On-Line Encyclopedia of Integer Sequences (OEIS).
Source: https://oeis.org/
License: Creative Commons Attribution-ShareAlike 4.0 (CC BY-SA 4.0)
OEIS End-User License Agreement: https://oeis.org/wiki/The_OEIS_End-User_License_Agreement
genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128
bolinas-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128
20 mammals (segmentation) segmentation enhancers (v20) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
8,672,102 sequences across 64… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128.python_enhancement_proposals
Python Enhancement Proposals
Description
Python Enhancement Proposals, or PEPs, are design documents that generally provide a technical specification and rationale for new features of the Python programming language.
There are been 661 PEPs published.
The majority of PEPs are published in the Public Domain, but 5 were published under the “Open Publication License” and omitted from this dataset.
PEPs are long, highly-polished, and technical in nature and often include… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/python_enhancement_proposals.Mind-Bench-Enhanced
Mind-Bench Enhanced + GenClaw Final Results
这是一个清理后的最终发布包,只保留可复现和可浏览所需内容。
内容
mind-bench/: 最终修正后的完整 Mind-Bench 数据目录,包含最终 Mind-Brush.jsonl 和全部参考图/输入图;已移除 .cache、._*、中间 .bak.*。
fixtures/: 最终评测使用的全量 fixture。
results/spp_v3/: GenClaw 最终 10 类评测结果、生成图、搜索图资产和 HTML。
docs/: 简要说明与上游贡献说明。
快速查看
# 总览 dashboard
open results/spp_v3/dashboard.html
# 500 条详细可视化
open results/spp_v3/analysis_10cls.html
关键文件
mind-bench/Mind-Brush.jsonl: 最终 patched 数据。… See the full description on the dataset page: https://huggingface.co/datasets/Yejy53/Mind-Bench-Enhanced.zoonomia-v1-v4_ccre_enhancer_centered-order
bolinas-dna/zoonomia-v1-v4_ccre_enhancer_centered-order
An enhancer-CENTERED training set for issue
#351, built by the
snakemake/zoonomia_projection_dataset pipeline
(workflow/rules/centered.smk) at commit
8127acfea5aa.
Provenance
Each training window is defined directly from an ENCODE cCRE V4 enhancer
(dELS + pELS): one 255 bp window centered on the cCRE midpoint
(make_enhancer_anchors, keep-all — clustered enhancers each keep their own
window), rather than a… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/zoonomia-v1-v4_ccre_enhancer_centered-order.enhanced-fall-dataset
Enhanced Fall Dataset (total = 25,691)
Prerequisite
Download bear7011/gemma-4-e4b-kinetics_54K first for some overlapped videos. (This setting prevents cascading forgetting.)
Notification!
Do not mix the sora-accident dataset into the training process.
Video sources:
videos/kinetics_fall, videos/kinetics_neg — Kinetics dataset
videos/oops — OOPS! dataset (Columbia)
File Structure
├── annotations
│ ├── prompts.json… See the full description on the dataset page: https://huggingface.co/datasets/bear7011/enhanced-fall-dataset.swerebench-traces-raw-source-verification-enhanced-20260617
SWE-rebench Raw Source Verification Enhanced 20260617
This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve.
Download
The full dataset directory is uploaded as a single compressed archive:
hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-Qwen3.5-35B-A3B.ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-Qwen-enhanced-gemma-4-26B-A4B-it.ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-Qwen3.5-35B-A3B.racist-and-sex-jokes-enhanced
Racist and Sex Jokes Enhanced
Augmented from Elgyn90/sex_and_racist_jokes using an abliterated version of Gemma-4-E4B to target the prompts to the joke.
ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-FLUX-enhanced-gemma-4-26B-A4B-it.ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-Qwen3.5-35B-A3B.ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-26B-A4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-26B-A4B-it.Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-ERNIE-enhanced-gemma-4-31B-it.swebench-enhanced-cwe
SWE-bench Enhanced with CWE Security Hints
这是 SWE-bench Verified 数据集的增强版本,包含了详细的 CWE(Common Weakness Enumeration)安全提示。
数据集描述
任务: astropy__astropy-12907
仓库: astropy/astropy
Hints 长度: 0 字符
增强内容
原始 SWE-bench 任务的 hints 字段已被增强,包含:
任务特定提示: 指向可能的 bug 位置和修复方向
CWE-754: Improper Check for Unusual or Exceptional Conditions(异常条件检查不足)
CWE-682: Incorrect Calculation(计算错误)
每个 CWE 包含:
详细描述
缓解措施
代码示例
最佳实践
CWE 覆盖
本数据集中的任务映射到以下 CWE:
CWE-754… See the full description on the dataset page: https://huggingface.co/datasets/Chenyang200/swebench-enhanced-cwe.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-31B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-31B-it.PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: personachat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-FLUX-enhanced-Qwen3.5-35B-A3B.PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: personachat-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-ERNIE-enhanced-Qwen3.5-35B-A3B.ConvAI2-ERNIE-enhanced-gemma-4-E4B-it
Visual Memory Results: convai2-ernie-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-E4B-it",
"hf_results_repo": "visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it",
"results_jsonl": "results/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.jsonl",
"hf_dataset": "visual-memory/ConvAI2-With-Ids_1k-no-redundancy",
"hf_mapping_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/ConvAI2-ERNIE-enhanced-gemma-4-E4B-it.Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-Qwen-enhanced-gemma-4-12B-it.Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it
Visual Memory Results: synthetic-persona-chat-flux-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "google/gemma-4-12B-it",
"hf_results_repo": "visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it",
"results_jsonl": "results/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.jsonl",
"hf_dataset":… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/Synthetic-Persona-Chat-FLUX-enhanced-gemma-4-12B-it.PersonaChat-Qwen-enhanced-Qwen3.5-35B-A3B
Visual Memory Results: personachat-qwen-enhanced
This dataset contains the scored output of a visual-memory perplexity experiment.
Experiment metadata
{
"experiment": {
"model_name": "Qwen/Qwen3.5-35B-A3B",
"hf_results_repo": "visual-memory/PersonaChat-Qwen-enhanced-Qwen3.5-35B-A3B",
"results_jsonl": "results/PersonaChat-Qwen-enhanced-Qwen3.5-35B-A3B.jsonl",
"hf_dataset": "visual-memory/PersonaChat-With-Ids_1k-no-redundancy"… See the full description on the dataset page: https://huggingface.co/datasets/visual-memory/PersonaChat-Qwen-enhanced-Qwen3.5-35B-A3B.
