EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model
Authors: Rithvik Jonna, Man Namgung, Aakash Gurram, Tinoosh Mohsenin
Organizations: Department of Electrical and Computer Engineering Johns Hopkins University · Laboratory for Computational Sensing and Robotics Johns Hopkins University
Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preserving navigation behavior. We introduce EdgeVLN, a runtime-aware, deployment-ready quantized VLN model that closes this gap. EdgeVLN combines a quantized StreamVLN model with Latent Trajectory Termination Extractor (LATTE), a lightweight causal transformer that improves real-time stopping by predicting a Stop Action verifier rank. Both execute through our llama.cpp VLN driver, which reconstructs streaming context and prunes memory tokens on-board. We characterize a pretrained StreamVLN backbone across weight quantization from 8 to 2 bits and multiple inference runtimes to identify a feasible operating point. LATTE reuses backbone hidden states within the budget freed by quantization, requiring neither a second vision encoder nor an additional backbone forward pass. We evaluate six backbone precisions and seven candidate stop heads on BF16 and IQ4 NL across all 1,839 R2R VLN-CE val-unseen episodes. We measure success rate (SR) in simulation and latency, energy, and resident memory on an NVIDIA Jetson Orin NX 16 GB. LATTE achieves our highest SR, 58.02 percent on the deployed 4-bit model, exceeding the BF16 baseline with only 0.013 s additional latency per navigation step. Four-bit formats achieve nearly identical SR, but step energy varies 36.8 times by execution path. Only IQ4 NL under our VLN driver fits the board, using 11.35 GB resident memory while running 20.8 times faster and using 13.3 times less energy than storage-streamed BF16. INT2 collapses. Runtime selection, memory-token pruning, and quantization are essential for efficient edge deployment.
Figures & tables
Fig. 1: Overview of EdgeVLN. EdgeVLN consists of a quantized StreamVLN model and LATTE, an inline Stop Action verifier, both executed by our llama.cpp VLN driver on the NVIDIA Jetson Orin NX 16 GB detailed below. The input is a natural-language navigation instruction, the current RGB frame, and a 32 -frame KV cache of earlier frames; the output is a navigation action—move forward, turn left, turn right, or stop—together with LATTE’s Stop Action verifier rank.
Fig. 2: EdgeVLN inference path. Input : a language instruction and the current camera frame. The quantized StreamVLN model encodes them into a chunk of four actions, and LATTE reads the same forward pass’s layer-24 activations to predict the Stop Action verifier rank. Output : the Stop Action verifier rank triggers STOP or the next action, which the robot executes to produce the next frame.
Fig. 3: Overview of seven Stop Action verifier candidates evaluated on the R2R VLN-CE dataset validation-unseen split. All variants use a frozen StreamVLN backbone quantized to IQ4_NL and different stop-prediction strategy: MLP-Vision(1,2), MLP-VT Concat, MLP-VT FiLM, LLM-hidden-MLP(L20,L24) and LATTE (ours), the selected candidate, taps intermediate Qwen2-7B hidden states and applies causal self-attention over recent hidden states for stop prediction. Gray denotes the frozen quantized backbone, white trained components, and green LATTE modules. The stop decision determines whether action generation continues, with MSE/asymmetric-MSE losses used by the corresponding heads.
Fig. 4: On-board measurement on the Jetson Orin NX. (a) Measurement setup in MAXN power mode: the robot stands still while the frames of one recorded real-world episode (89 frames) are replayed to it, so latency, power, energy, and memory are measured without a trajectory that changes from one configuration to another. (b) Both llama.cpp configurations draw less power and finish sooner than NF4 under PyTorch. (c) Mean step latency per configuration with peak resident memory; the number inside each per-chunk bar is how many chunk compressions occur in one episode. EdgeVLN improves success rate while nearly matching the quantized StreamVLN’s on-board latency, energy, and memory under our driver.
#
Quant.
Size
SR ↑
OSR ↑
RAM ↓
Power
Energy
Step Latency
(GB)
(%)
(%)
(GB)
(W)
(J)
(s)
1
BF16 [ 1 ]
16.06
57.69
65.04
14.99
10.85
151.68
13.98
2
INT8
10.36
57.04
64.82
15.98
loads, but OOM after 8 steps
3
INT4
7.68
57.48
64.60
15.85
19.79
419.70
21.21
4
NF4
7.68
57.15
64.76
14.70
18.18
19.31
1.06
5
✓✓ IQ4_NL
7.94
57.31
64.11
11.35
16.97
11.41
0.67
TABLE I: Accuracy and on-board measurement results. (a) reports the six precisions of the StreamVLN model, from BF16 down to 2 bits; (b) reports seven stop action verifier candidates with baseline, each on the two deployable backbones, BF16 and IQ4_NL. SR and OSR come from 1,839 closed-loop R2R VLN-CE Dataset val-unseen split episodes evaluation. We measured RAM, power, energy and latency on an NVIDIA Jetson Orin NX. As BF16 does not fit, only 20 of the 28 transformer blocks plus the output layer stay resident, the rest streamed from storage at each token.
Edge deployment of Vision-Language Models (VLMs) faces a tradeoff between latency and accuracy: cloud execution provides high-quality predictions but incurs communication delay and energy cost, while edge-only execution is faster but less accurate due to limited model capacity. This trade-off is further complicated by heterogeneity in image quality and reasoning complexity, making static placement suboptimal. We present INAR-VL, a lightweight edge-cloud routing system for multimodal inference in a two-tier deployment. INAR-VL maintains complementary VLMs across edge and cloud and uses lightweight image and text complexity signals to guide routing and model selection, executing simple queries locally while offloading complex ones when beneficial. Evaluation on visual question answering shows that INAR-VL executes 36% of requests on the edge, reduces latency by 24%, lowers energy by 26%, and preserves 97% of cloud-level accuracy.
Vision Language Models (VLMs) have emerged in the robotic domain as a powerful tool that enables environmental perception with language context, serving as a catalyst for open-vocabulary tasks like ObjectNav. Yet, their computational footprint typically confines them to cloud execution, hindering low-latency inference with local deployment on resource-constrained robots. To address this challenge, we present a distillation strategy that transfers complex spatial-semantic reasoning from large frontier models into a lightweight, 4B-parameter local VLM for edge execution on embedded GPU devices (e.g., Jetson Orin). We first establish a State of the Art (SotA), Scene Graph (SG)-based pipeline using Claude Sonnet 4.6, achieving a 39.7% Success Rate (SR) on the HM3D OVON benchmark. We then demonstrate that fine-tuning Qwen3.5-4B on just 500 frontier reasoning traces effectively enables navigation capabilities, yielding a SR of 34.5%, narrowing the gap to the performance of large cloud models. Finally, we introduce E-RLVR with Token Generation (TG) regularization to compress output sequence lengths for physical deployment while grounding the agent in its task. This downstream optimization reduces TG overhead by 72.1% and latency by 71.8%. Combined with quantization, this joint strategy yields a cumulative 82.8% reduction in overall inference latency without significantly sacrificing performance, presenting a viable paradigm for local, low-latency VLM execution on mobile robots.
Nicolas Baumann, Liam Boyle, Pu Deng +5
Center for Project-Based Learning, 2Integrated Systems Laboratory, 3Computer Vision and Geometry ETH Zurich
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge platforms impractical. Aggressive, sub-4-bit weight quantization is the natural solution, yet existing post-training quantization (PTQ) methods suffer severe performance degradation in this regime. To address this, we introduce ActQuant, an action-guided mixed-precision PTQ framework that operates in two stages: (1) an inter-tensor bit allocator that assigns each weight matrix a single bit-width based on how much it contributes to predicting the agent's actions; (2) an intra-tensor scale optimizer tunes per-block quantization scales using action-aware curvature, so that dynamic range is concentrated on the weights most influential for control. To deliver the on-device benefits of our aggressive quantization, we further introduce OmniModel.cpp, an agentic conversion pipeline that ports architectures into a native C/C++ runtime with efficient low-bit kernels. We evaluate ActQuant both in simulation and on a real-world 6-DoF UR3 arm, with all models deployed through OmniModel.cpp. On the LIBERO benchmark, ActQuant is the only method that operates at or below 3 bits-per-weight, retaining 95.0% on OpenVLA-OFT and 94.8% on π0.5. Pushed further, ActQuant reaches 2.5 bpw at 90.1% on OpenVLA-OFT, compressing the backbone from 14.3 GB to 2.7 GB (5.3×). On the physical UR3 arm, π0.5 quantized with ActQuant retains the baseline's success rate while reducing the memory footprint by 2.5×.
Arash Akbari, Arman Akbari, Masih Eskandar +11
1Northeastern University · University of Georgia · 3Cisco Systems