datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.korean-english-multitarget-ted-talks-task
Dataset Card for english-korean-multitarget-ted-talks-task
Dataset Summary
Parallel English-Korean Text Corpus
Text was originally transcribed to English from various Ted Talks, then translated to Korean by TED translators
Approximately 166k train, 2k validation, and 2k test sentence pairs.
Supported Tasks and Leaderboards
Machine Translation
Languages
English
Korean
Additional Information
Dataset Curators
Kevin Duh, "The… See the full description on the dataset page: https://huggingface.co/datasets/msarmi9/korean-english-multitarget-ted-talks-task.afdb-msa-index
AlphaFold DB minimizer index
A static sequence-search index over 239,602,633 AlphaFold DB v6 entries
(99.4% of the database), designed to be queried directly from a browser with
HTTP Range requests. No server, no search engine, no MMseqs2.
It exists to answer one question quickly: which AFDB entry is ≥90% identical to
my sequence? — so that entry's precomputed MSA can be borrowed and re-indexed
onto the query instead of computing a new alignment.
Used by AFDB MSA.… See the full description on the dataset page: https://huggingface.co/datasets/sokrypton/afdb-msa-index.industrial-agent-benchmark
Industrial Agent Benchmark
Industrial Agent Benchmark (IAB) is an open benchmark for evaluating Industrial AI systems, Manufacturing AI assistants, and Industrial Agents.
This Dataset Card describes the Hugging Face Dataset release for Industrial Agent Benchmark v2.2.0 Japanese Canonical Normalization.
Repository:
https://github.com/masahirosakae/industrial-agent-benchmark
Hugging Face Dataset Repository:
https://huggingface.co/datasets/MSakae/industrial-agent-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/MSakae/industrial-agent-benchmark.indirect-requests
IndirectRequests
IndirectRequests is an LLM-generated dataset of user utterances in a task-oriented dialogue setting where the user does not directly specify their preferred slot value.
IndirectRequests was generated by crowdsourcing human labels over a dataset generated using a combination of GPT-3.5 (turbo) and GPT-4.
Each utterance is labelled along two dimensions:
World Understanding (the degree of world understanding it takes to understand the utterance)
Unambiguity (whether… See the full description on the dataset page: https://huggingface.co/datasets/msamogh/indirect-requests.MSAVBenchmydataset
Uploaded model
Developed by: MSakar
License: apache-2.0
Finetuned from model : meta-llama/Meta-Llama-3-8B-Instruct
This llama model was trained 2x faster with Unsloth and Huggingface's TRL library.
msap-align-fairness-20260702scorpio-gene-taxa
Scorpio-Gene-Taxa Dataset: Curated Gene Sequences for Model Training and Evaluation
Protein and DNA Sequences Available for Evaluating DNA and Protein Language Models for Gene and Taxonomy Prediction
This dataset was created using the Woltka pipeline to compile the Basic Genome Dataset, containing 4,634 genomes. Each genus is represented by a single genome, with a focus on bacteria and archaea. Viruses and fungi were excluded due to insufficient gene information.
Key… See the full description on the dataset page: https://huggingface.co/datasets/MsAlEhR/scorpio-gene-taxa.boostie
Dataset Card for BoostIE
Dataset Description
BoostIE was created by adapting synthie_code_pc split of the SynthIE dataset by randomly dropping some entities from the KB for 40% of samples.
It also includes outputs of SynthIE-base-FE model from SynthIE repo, obtained in constrained and unconstrained manner.
For more details:
Github repository: https://github.com/epfl-dlab/boostie
Languages
BoostIE only contains data in English, as the SynthIE dataset it… See the full description on the dataset page: https://huggingface.co/datasets/msakota/boostie.ms-ajarAjar articles scraped on 8/7/2023
msap-mvp0-perchunk-v2-2026-06-04
