cs.CLJul 13, 2025

NMIXX: Domain-Adapted Neural Embeddings for Cross-Lingual eXploration of Finance

Authors: Hanwool Lee, Sara Yu, Yewon Hwang, Jonghyun Choi, Heejae Ahn, Sungbum Jung, Youngjae Yu

Organizations: FinancialNLPLab, MODULABS Shinhan Securities Seoul, Republic of Korea · FinancialNLPLab, MODULABS KT Seoul, Republic of Korea · FinancialNLPLab, MODULABS EMRO Seoul, Republic of Korea · FinancialNLPLab, MODULABS Samsung Fire & Marine Insurance Seoul, Republic of Korea · FinancialNLPLab, MODULABS KB Securities Seoul, Republic of Korea · FinancialNLPLab, MODULABS Seoul, Republic of Korea · Seoul University Seoul, Republic of Korea

Abstract

Financial text embeddings must distinguish changes in event status, perspective, and obligations even when passages share similar wording. NMIXX adapts existing encoders through 18.8k source-linked triplets: paraphrases and Korean-English translations preserve meaning, while targeted financial rewrites introduce semantic contrasts. We examine this recipe across seven backbones on English and Korean financial and general-domain semantic textual similarity (STS), and analyze the composition and passage lengths of KorFinSTS. BGE-M3 attains the highest adapted financial correlations in this comparison, improving from 0.1969 to 0.2967 on FinSTS and from 0.0512 to 0.2732 on KorFinSTS. Its general English and Korean correlations decrease by 0.0391 and 0.0463. Across the seven models, five improve their mean financial correlation, but all reduce their mean general-domain correlation. Per-language comparisons and benchmark-weight sensitivity analysis reveal differences obscured by a single aggregate score. The study contributes a finance-specific supervision design and evidence for evaluating adaptation jointly with retained general semantic capability; direct cross-language retrieval remains outside its evaluation scope.

Figures & tables

Explore similar work

CardsList
  1. JFinTEB: Japanese Financial Text Embedding Benchmark

    Apr 17, 2026Masahiro Suzuki, Hiroki SakajiLegal Natural Language Processing BenchmarksText Embeddings