datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
android-controlAndroidControlParsedWithImages-20kandroid_controlAndroidControlAndroidFlux_RL_Train
AndroidFlux RL — policy prompts
Slice
Prompts
Source
t_minus_n
2,459
AndroidFlux source trajectories
successful
2,623
AndroidFlux source trajectories
t_minus_1
2,459
AndroidFlux source trajectories
t
2,459
AndroidFlux source trajectories
Subtotal
10,000
ui_genie
10,000
UI-Genie reward-model prompts (5,000-prompt core marked by in_reduced)
The four AndroidFlux slices are drawn from replayed source trajectories.
The successful slice contains 53… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_RL_Train.androidcontrol-rf-sft-v1AndroidControlParsed-20kandroid-control-1kAndroidControlParsedWithImages-20k-TESTONLYandroid-control
Android Control Test Set
Notes
NOTE i uploaded it so it easier for running evaluation on it
I am not responsible if you use it in training :<
I assume most models train on train-set/val-set but not test-set
I assume that the ids here are episode_id
https://console.cloud.google.com/storage/browser/_details/gresearch/android_control/splits.json , so I download data using smolagents
https://huggingface.co/datasets/smolagents/android-control and iterate and filter ids to… See the full description on the dataset page: https://huggingface.co/datasets/aliaagheis/android-control.artemis-android-dynamic-traces
ARTEMIS Android Dynamic Traces
ARTEMIS executes Android applications in controlled emulators and collects dynamic
analysis artifacts. This public dataset contains the runs performed in Android 10
(API 29) and Android 14 (API 34). APK binaries are not included. The public dataset is
serrooT/artemis-android-dynamic-traces.
The dataset preserves the analysis artifacts as collected. Trace contents are not
redacted, sanitized, normalized, or recompressed. Existing Zstandard files… See the full description on the dataset page: https://huggingface.co/datasets/serrooT/artemis-android-dynamic-traces.cqadupstack-android-vn
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackAndroid-VN"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out the GitHub repitory.
Citation
If you use this dataset, please cite the dataset as well as mteb, as this dataset likely includes additional processing as a… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/cqadupstack-android-vn.beir-cqadupstack-android
CQADupstackAndroidRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackAndroidRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/CQADupstackAndroidRetrieval @ 9be4c0e46342 (the revision… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-android.AndroidControlParsedWithImages-20k-maxnshot2android-kernel-security-datasetCyber_AndroidAndroidFlux_RL_Train_Test
AndroidFlux RL Train Test
Test-only AndroidFlux RL data, in two directories.
Directory
Rows
A row is
Docs
data_from_failure_recovery/
1,107
a checkpoint-to-next-action prompt; the policy generates
docs/DATA_CONSTRUCTION.md
data_from_rm_eval/
2,000
a prompt plus a candidate action, with the preference stored outside the conversation
docs/DATA_FROM_RM_EVAL.md
Shared code lives in code/.
The tables in each directory have different schemas, so they load as separate… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_RL_Train_Test.CQADupstack-Android-PL
CQADupstack-Android-PL
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Stack Exchange Question Duplicate Pairs Dataset
Task category
t2t
Domains
Programming, Web, Written, Non-fiction
Reference
https://huggingface.co/datasets/clarin-knext/cqadupstack-android-pl
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstack-Android-PL"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CQADupstack-Android-PL.android_icon_datasetAndroid icons collected from a variety of sources and grouped manually. Beware of the Misc_or_unknown class that contains the weird icons we couldn't classify.
cqadupstack-android
Dataset Card for "cqadupstack-android"
More Information needed
cqudubstack-android
Dataset Card for "cqudubstack-android"
More Information needed
AndroidControl_debugbeir_cqadupstack_android_test
beir_cqadupstack_android_test
BEIR CQADupStack/android test split
Field
Value
Benchmark
beir
Sub-benchmark
cqadupstack_android
Type
retrieval
Items
699
Exported from Langfuse.
AndroidControl_300samples_qwen2_5vlsei-cert-android-rules
Dataset Card for SEI CERT Android Secure Coding Standard (Wiki rules)
Structured export of the SEI CERT Android Secure Coding Standard from the SEI wiki.
Dataset Details
Dataset Description
Android-focused secure coding rules (Java / Android APIs) with descriptive text and examples as published by SEI CERT.
Curated by: Derived from public SEI CERT wiki pages; packaged as CSV by the dataset maintainer.
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/safebuffer/sei-cert-android-rules.AndroidFlux_Source_Traj_dev
AndroidFlux Source Trajectories
This dataset packages the complete 116-task AndroidFlux source trajectories for
15 source models/runs. Every model is stored in a separate replay parquet and a
separate replay-history parquet. The schemas are identical to
Gyubeum/AndroidFlux_Failure_Recovery_Eval, so the existing AndroidFlux
materializer and replay tooling can consume the files without conversion.
Dataset scope
Split
Label provenance
Models
Rows per model… See the full description on the dataset page: https://huggingface.co/datasets/Gyubeum/AndroidFlux_Source_Traj_dev.AndroidControl_300_samples_v2AndroidControl_700samples_qwen2_5vl_intention_refinedhey-android
Hey Android Wake Word Synthetic Speech Dataset
... (full content from user prompt) ...
cqudupstack-android
Dataset Card for "cqudupstack-android"
More Information needed
