Semantic scene understanding in robotics requires representations that are both metric-accurate and queryable via natural language in real-time. While recent Vision-Language Models enable powerful 2D image-text alignment, their integration into real-time 3D mapping systems remains challenging due to their requirements on ground truth poses, computational cost, and memory constraints. We present VLEM (Vision-Language Embedding Mapping), a real-time framework for integrating pixel-aligned 2D vision-language embeddings from various backends into a globally consistent, metric-accurate 3D representation, requiring only a raw RGB-D stream. Compared to ConceptFusion, Open-Fusion, and RayFronts, VLEM provides better open-set segmentation performance and a more compact representation. We further demonstrate VLEM's versatility in interactive real-time robotic manipulation tasks and mobile mapping scenarios.