prototype
refinedweb-embedded_prototypeContaining 768-dimensional embedding vectors derived from the "content" column of the RefinedWeb dataset.
The embeddings were generated using the E5-Base-4k model with a context length of 1024, employing scaled dot-product attention.
Original dataset:
This dataset bases as derivative work on RefinedWeb, an English web dataset created by the TII (Technology Innovation Institute).
Attribution is given to the TII as original authors of the RefinedWeb dataset as per Section 4 of RefinedWeb's Open… See the full description on the dataset page: https://huggingface.co/datasets/Marcus2112/refinedweb-embedded_prototype.nesteo-prototype
NestEO: Modular and Hierarchical EO Dataset Framework
NestEO is a hierarchical, resolution-aligned, UTM-based nested grid dataset framework supporting general-purpose, multi-scale multimodal Earth Observation workflows. Built from diverse EO sources and enriched with metadata for landcover, climate zones, and population, it enables scalable, representative and progressive sampling for AI4EO.
Grid Levels: 120000m, 12000m, 2400m, 1200m, 600m, 300m, 150mGrid Metadata: ESA WorldCover… See the full description on the dataset page: https://huggingface.co/datasets/nesteo-datasets/nesteo-prototype.Ilesha-Gold-Prospectivity-AME-prototypedetails_Ichsan2895__Merak-7B-v5-PROTOTYPE1
Dataset Card for Evaluation run of Ichsan2895/Merak-7B-v5-PROTOTYPE1
Dataset automatically created during the evaluation run of model Ichsan2895/Merak-7B-v5-PROTOTYPE1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Ichsan2895__Merak-7B-v5-PROTOTYPE1.smart_market_prototype_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 300,
"total_frames": 378133,
"total_tasks": 30,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:300"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kdy93/smart_market_prototype_2.anyrag-prototype
AnyModal RAG — validated end-to-end prototype
Reference run: job 6ab1136251992417dfccf413 (a10g-small, 4m35s wall-clock, inference only).
Stack (all open weights)
Unified embedder: Qwen/Qwen3-VL-Embedding-2B — one 2048-d space for text, images, video keyframes, 3D-proxy views
Reranker: Qwen/Qwen3-Reranker-0.6B (cross-encoder)
Audio: laion/larger_clap_general (Apache-2.0 audio/text space), faster-whisper tiny for STT
Generator: Qwen/Qwen3-1.7B (decoder-only… See the full description on the dataset page: https://huggingface.co/datasets/jkorstad/anyrag-prototype.
