datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swedish-dynaword
🧨 Swedish Dynaword
Version
0.0.13 (Changelog)
Language
Swedish (sv, swe)
License
Openly Licensed, See the respective dataset
Models
Currently there is no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 547.06M
Number of tokens (Llama 3): 36.34B
Average document length in tokens (min, max): 66.42 (2, 8.14M)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/swedish-dynaword.LegoFlow-SWE
LegoFlow-SWE · 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases
GitHub · Docs · Blog · HuggingFace · LegoX
LegoFlow-SWE
5,000 verified Harbor SWE tasks mined by LegoFlow Curator, shipped in original and anti-hack prompt versions, plus two GLM-5.2 trajectory releases under OpenHands SDK and OpenCode, totaling 9,767 trajectories.
Release
Count
What it is
tasks/
5,000
Original prompts
tasks-anti-hack/
5,000
Same task IDs and task files, with… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/LegoFlow-SWE.SWE-PolyVision
SWE-PolyVision Public Tasks
Public question package for SWE-PolyVision from CosmosMind AI Lab.
This release contains the 48 tasks used in the benchmark. Each task includes only
the task statement, the fixed repository revision, and the visual evidence needed
to inspect the issue. Developer patches, reference fixes, verifier assets, model
results, traces, and answer materials are intentionally excluded.
The task statements are derived from the benchmark packages and scrubbed at… See the full description on the dataset page: https://huggingface.co/datasets/CosmosMind/SWE-PolyVision.
