CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-fixed-q0p8-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_fixed_q0.8_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes290 downloads1mo agoHugging Face02zjhhhh /fixed-n-rb-cost-aware-marginrl-qwen3-1.7b-base-math12k-token-mean-rerun-rollouts fixed_n_rb_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_token_mean_rerun rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes88 downloads18d agoHugging Face03zjhhhh /fixed-n-rb-er-cost-marginrl-qwen3-1.7b-base-math12k-token-mean-run2-rollouts fixed_n_rb_er_cost_marginrl_Qwen3-1.7B-Base_math12k_token_mean_run2 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes82 downloads17d agoHugging Face04abidlabs /repro-memory-savings-at-what-cost-a-study-of-alternatives-to-backpropagation-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes66 downloads2mo agoHugging Face05narcolepticchicken /agent-cost-traces Agent Cost Traces: Synthetic Training Data 10,000 synthetic agent traces for training cost-aware model routers and agent optimizers. Schema Field Type Description trace_id string Unique identifier request string User request text task_type string One of 9 task categories difficulty int Estimated difficulty (1-5) model_tier int Model tier used (1-5) model_success bool Whether the model succeeded optimal_tier int Minimum tier that would succeed… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/agent-cost-traces.tabularother10K<n<100K0 likes55 downloads5mo agoHugging Face06hi-todayis-jh /fixed-n-rb-offset-cost-aware-marginrl-qwen3-1.7b-base-math12k-offset2048-token-mean-rollouts fixed_n_rb_offset_cost_aware_marginrl_Qwen3-1.7B-Base_math12k_offset2048_token_mean rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular100K<n<1M0 likes50 downloads4d agoHugging Face07JiayuJeff /CostBench CostBench This dataset contains 381 records from CostBench_queries.json. The queries are derived from the official CostBench benchmark repository and follow its travel-task query schema. Contents Top-level fields: query_id, TimeInfo, task, is_location, goal_type, preferences, groundtruth, validation_raw, is_valid, user_requirements, query Field Guide query_id: Unique identifier for each query. TimeInfo: Time context used in the task prompt. It is an ID-style… See the full description on the dataset page: https://huggingface.co/datasets/JiayuJeff/CostBench.tabularn<1K3 likes49 downloads6mo agoHugging Face08zjhhhh /er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42-rollouts er_cost_marginrl_r1_distill_1.5b_compression_n16_b512_32k_lr1e-6_kl0_seed42 rollouts This dataset contains one compressed JSONL shard for every completed training step. The step and rollout_index columns uniquely locate a rollout within this training run. Run metadata and per-step row counts are recorded in rollout_manifest.json. tabular10K<n<100K0 likes46 downloads11d agoHugging Face09dispatchAI /cost-analysis Cost Analysis Cloud API vs on-device inference cost comparison. At 10K queries/day: Save $18,249/year with on-device. At 100K queries/day: Save $182,499/year. 🚀 dispatchAI tabularn<1K0 likes17 downloads3mo agoHugging Face10procedure2012 /namespace-cost-modeltabularn<1K0 likes16 downloads2mo agoHugging Face11costadev00 /wikipedia-pt-br-domain wikipedia-pt-br-domain-gemma Versão enriquecida de costadev00/wikipedia-pt-br-extract com um label sintético de domínio por artigo. Processo Cada registro preserva os campos originais esperados da Wikipedia (page_id, title, text, ns, section_texts) e adiciona domain, derivado do label documental primary_category produzido pelo modelo. Modelo Modelo usado para labeling: google/gemma-4-26B-A4B-it. Versão da pipeline: 0.1.0. Limitações O campo domain é… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-domain.tabulartext-classificationn<1K0 likes7 downloads5mo agoHugging Face12Ouzhang /low-high-cost-promptstabular100K<n<1M0 likes4 downloads4mo agoHugging Face13yding03 /namespace-cost-modeltabularn<1K0 likes4 downloads2mo agoHugging Face14costadev00 /wikipedia-pt-br-article-labelsgated wikipedia-pt-br-article-labels-gemma Dataset sintético de labels documentais para artigos da Wikipedia em português brasileiro. Origem Os registros derivam de costadev00/wikipedia-pt-br-extract. Cada linha representa um artigo e preserva page_id, title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A pipeline aplica triagem determinística em CPU, remove documentos ruins e usa um modelo Gemma local para gerar categorias, subcategorias, tipo… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-article-labels.tabulartext-classificationn<1K0 likes3 downloads5mo agoHugging Face15costadev00 /wikipedia-pt-br-instructionsgated wikipedia-pt-br-instructions-gemma Dataset sintético de Instruction Following em português brasileiro, derivado de artigos da Wikipedia pt-BR. Origem Os exemplos derivam de costadev00/wikipedia-pt-br-extract e preservam source_page_id, source_title, source_dataset e a licença herdada cc-by-sa-3.0. Processo A geração é inspirada em Alpaca e Self-Instruct: cada artigo válido passa por uma chamada de analista documental que produz candidatos de instrução ancorados… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/wikipedia-pt-br-instructions.tabulartext-generationn<1K0 likes3 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.