datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dlm-exp1-drafter-voting-research-artifacts
dlm-exp1 drafter-voting: research artifacts (public mirror)
Everything bulky from the drafter-voting branch of https://github.com/NoviceCoderInfinity/dlm-general-expert-exp1
that git does not track (uploaded 2026-09-06 before the compute instance was destroyed):
logs/research_v1/: baseline cells (one JSON line per item; per-pass NPZ traces for both checkpoints in
baselines_traces/ and baselines_b32_traces/), state-batch fidelity probe states (fidelity/), flush-fusion
boundaries… See the full description on the dataset page: https://huggingface.co/datasets/Arushhh/dlm-exp1-drafter-voting-research-artifacts.dlm-exp1-drafter-voting-research-artifacts
dlm-exp1 drafter-voting research artifacts (public)
Draft-aligned judge LoRA for SDAR-30B-A3B-Chat-b32 with a 1.7B drafter and DES expert pooling: goal = decode speedup with minimal accuracy loss vs the stock model.
model / setting
weights
eval records
stock SDAR-30B-A3B-Chat-b32 (+ 1.7B drafter where used)
JetLM/SDAR-30B-A3B-Chat-b32 @ c351bbc3… (Hugging Face)
reference_evals/da_decode (*_stock), r3_A/evals/stock
stock + DES32 / DES48 (decode-time expert pooling… See the full description on the dataset page: https://huggingface.co/datasets/Anupam-Rawat-IITB/dlm-exp1-drafter-voting-research-artifacts.otu-taxa-paper-artifacts
OTU-Taxa paper artifacts
This repository contains stable processed artifacts used by the OTU-Taxa paper
and the frozen contracts needed to reproduce them.
Related repositories
Source code:
https://github.com/bio-ontology-research-group/otu-taxa-autoregressive
Comparator pipelines:
https://github.com/bio-ontology-research-group/microbiome_foundation_model_benchmarks
Pretrained OTU-Taxa model:
bio-ontology-research-group/otu-taxa
Model weights and reusable… See the full description on the dataset page: https://huggingface.co/datasets/bio-ontology-research-group/otu-taxa-paper-artifacts.reasoning-switch-research-artifactsresearch-artifacts-archive
Research Artifacts Archive
This dataset contains an encrypted archive of code, experiment metrics, configuration files, notebooks, and explicitly selected model/checkpoint directories.
The payload is encrypted with GnuPG symmetric AES-256 encryption and split into ordered shards for reliable Hub storage. Hugging Face cannot preview the encrypted contents.
Files
data/data-00000-of-00022.tar.zst.gpg through data/data-00021-of-00022.tar.zst.gpg
SHA256SUMS contains a… See the full description on the dataset page: https://huggingface.co/datasets/vatolinalex/research-artifacts-archive.GLM-5.3-Flash-Research-Artifacts
GLM-5.3 Flash research artifacts
This is a research dataset from Alexey's attempt to make GLM-5.3 Flash useful on 4x RTX 3090. It contains numerical measurements, synthetic reference tests and a license-reviewed subset of teacher distributions captured on RTX 6000 PRO hardware. The experiment did not produce a quality-qualified production model.
The teacher was RedHatAI/GLM-5.3-Flash-NVFP4, revision 36c184c6cda000a481711306df5adde42f63321a, based on Z.AI's GLM-5.3 Flash. These… See the full description on the dataset page: https://huggingface.co/datasets/anonymousmaharaj/GLM-5.3-Flash-Research-Artifacts.Sotopia-artifactshumanize-rl-research-artifacts-env0314
Humanize-RL Research Artifacts Env0314
This dataset archives the ignored local artifacts needed to continue the Humanize-RL reward patch, SFT repair, and Qwen 2B/9B ablation work after leaving the original worktree.
Primary source run: zztqgqclh3y3hslpjsofzpcf
Prime env: jayshah5696/humanize-rl-env@0.3.14
W&B run: https://wandb.ai/jayshah5696/humanize-rl/runs/akzopsz9
Published SFT dataset: jayshah5696/humanize-rl-prime-sft-messages-env0314
Contents… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-research-artifacts-env0314.skillsbench-research-artifactsai-research-intelligence-artifacts
AI Research Intelligence – RAG Artifacts
This dataset repository contains the retrieval artifacts for a fully local Retrieval-Augmented Generation (RAG) system built on arXiv data.
Overview
These artifacts are used by the online inference pipeline to perform document retrieval. They are externalized from the main code repository to keep the project lightweight and reproducible.
Contents
retriever/
arxiv_ai.index
FAISS index built on MiniLM… See the full description on the dataset page: https://huggingface.co/datasets/antalbalint97/ai-research-intelligence-artifacts.searchforge-research-artifacts
SearchForge research artifacts
Versioned synthetic-world data, typed task programs, evidence audits, protected
splits, and recovery/training states for the SearchForge symbolic task-model
project. These artifacts support source-to-program-to-grounding reproduction.
The first snapshot is recovery and implementation evidence, not a successful
solver-learning result. Historical findings remain: E3 PASS; E5a PASS;
confirmatory E5 FAIL under its preregistered criterion; E6… See the full description on the dataset page: https://huggingface.co/datasets/zcheng256/searchforge-research-artifacts.
