datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sweden_100K_difficultSuperLim\ipfs_sweden_laws_ir
Sweden legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_sweden_laws (revision 0a1e752cb178b83a779996f0e60e6f8d305a857e) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Sweden prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_sweden_laws_ir.details_AI-Sweden-Models__gpt-sw3-6.7b-v2-instruct
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-6.7b-v2-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-6.7b-v2-instruct on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-6.7b-v2-instruct.details_AI-Sweden-Models__gpt-sw3-40b
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-40b.details_AI-Sweden-Models__gpt-sw3-20b
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-20b
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-20b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-20b.details_AI-Sweden-Models__gpt-sw3-6.7b-v2
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-6.7b-v2
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-6.7b-v2 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-6.7b-v2.details_AI-Sweden-Models__gpt-sw3-20b-instruct
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-20b-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-20b-instruct on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-20b-instruct.details_AI-Sweden-Models__gpt-sw3-1.3b-instruct
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-1.3b-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-1.3b-instruct on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-1.3b-instruct.sweden-power-market
Sweden Power Market Data
Part of a country-split collection of historical electricity market data gathered for the elpriser.org price-forecasting project. See the index dataset for the other countries: Denmark, Germany, Norway, Sweden, Finland, Netherlands.
License & attribution
CC BY 4.0. Source: ENTSO-E Transparency Platform (transparency.entsoe.eu).
Contents
All 4 Swedish bidding zones (SE1–SE4), ~2018-09/10 → present. Five files per zone:… See the full description on the dataset page: https://huggingface.co/datasets/Elpriser/sweden-power-market.details_AI-Sweden-Models__gpt-sw3-6.7b
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-6.7b
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-6.7b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-6.7b.details_AI-Sweden-Models__gpt-sw3-1.3b
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-1.3b
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-1.3b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-1.3b.BiaSWE
About BiaSWE
We present BiaSWE, a small annotated dataset for misogyny detection in Swedish, annotated for hate speech, misogyny, misogyny type categories and severity by a group of experts in social sciences and humanities. This dataset is a proof of concept and it can be used to perform classification of misogynistic vs non-misogynistic text, as well as debiasing on Language Models.
Content warning: Sensitive content might appear in this dataset. The language does not reflect the… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/BiaSWE.details_AI-Sweden-Models__gpt-sw3-126m-instruct
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-126m-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-126m-instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-126m-instruct.details_AI-Sweden-Models__gpt-sw3-126m
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-126m
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-126m on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-126m.details_AI-Sweden-Models__gpt-sw3-356m-instruct
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-356m-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-356m-instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-356m-instruct.details_AI-Sweden-Models__gpt-sw3-356m
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-356m
Dataset Summary
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-356m on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AI-Sweden-Models__gpt-sw3-356m.ipfs_sweden_laws
Swedish Statute Book (Svensk författningssamling / Riksdagen)
Research snapshot of official national legislation from Riksdagen öppna data / Svensk författningssamling (data.riksdagen.se).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-18
Coverage
catalog-backed complete (Cap densify tip∪local @2ff14d95)
Source
Riksdagen öppna data / Svensk författningssamling… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_sweden_laws.Mr.Porter.Product.prices.Sweden
Mr Porter web scraped data
About the website
The EMEA region, specifically Sweden, has seen a significant rise in the luxury online retail industry, where Mr Porter operates. The growth has primarily been driven by the fast-paced digitalization, significant internet penetration, and a growing number of digitally native consumers. Additionally, Swedish consumers, renowned for their fashion-forward approach, have demonstrated a strong appetite for luxury fashion products… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Sweden.Prada.Product.prices.Sweden
Prada web scraped data
About the website
The Luxury Fashion Industry in the EMEA region, particularly in Sweden, is a thriving market with high demand for exclusive and high-end products. Prada, a renowned player in this industry, holds a significant presence. The industry is currently experiencing a significant shift towards digitalization and online retail, also known as Ecommerce, fueled by changing consumer behaviors and advancements in technology. A concrete example… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Prada.Product.prices.Sweden.Dolci-Instruct-SFT-translated
Dolci-Instruct-SFT-translated (Swedish)
This dataset is a Swedish machine translation of the openeurollm/Dolci-Instruct-SFT-translated dataset, originally created as part of the OpenEuroLLM project.
Dataset details
Examples: 494,841 multi-turn conversations
Language: Swedish (sv-SE)
Format: Chat/messages format (id, messages)
License: Apache 2.0
Translation
All English source texts were machine-translated to Swedish using Google Gemma 3 27B-IT (w8a8_fp8… See the full description on the dataset page: https://huggingface.co/datasets/AI-Sweden-Models/Dolci-Instruct-SFT-translated.Gucci.Product.prices.Sweden
Gucci web scraped data
About the website
The fashion industry in the EMEA region, more specifically in Sweden, has seen a significant shift in recent years. With the surge in digital transformation, there has been remarkable growth in the online luxury fashion market, where premier brands like Gucci have amplified their presence. One particular focus area has been the Ecommerce product-list pages (PLP), aiming to provide a seamless and immersive digital shopping… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Gucci.Product.prices.Sweden.molarpower3AI-Sweden-Models__gpt-sw3-40b-details
Dataset Card for Evaluation run of AI-Sweden-Models/gpt-sw3-40b
Dataset automatically created during the evaluation run of model AI-Sweden-Models/gpt-sw3-40b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI-Sweden-Models__gpt-sw3-40b-details.sweden_100K_easyCOVID-19_Line_1177_Sweden
[!NOTE]
Dataset origin: https://live.european-language-grid.eu/catalogue/corpus/21224
Description
Multilingual (EN, BG, DE, ES, FI, FR, PL, RO, RU, SV, TR) COVID-19-related corpus acquired from the website (https://www.1177.se/) of the line 1177 of Sweden (16th September 2020). It contains 2034 TUs in total.
Citation
COVID-19 Line 1177 of Sweden dataset v1. Multilingual (EN, BG, DE, ES, FI, FR, PL, RO, RU, SV, TR) (2020, September 18). Version 1.0. [Dataset (Text… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/COVID-19_Line_1177_Sweden.AI-Sweden-Models__Llama-3-8B-instruct-detailssweden_1M_easyreturn_sweden_test16This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 40,
"total_frames": 16006,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:40"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ljbcote/return_sweden_test16.return_sweden_test3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 2396,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ljbcote/return_sweden_test3.
