Organizations: FinancialNLPLab, MODULABS Shinhan Securities Seoul, Republic of Korea · FinancialNLPLab, MODULABS KT Seoul, Republic of Korea · FinancialNLPLab, MODULABS EMRO Seoul, Republic of Korea · FinancialNLPLab, MODULABS Samsung Fire & Marine Insurance Seoul, Republic of Korea · FinancialNLPLab, MODULABS KB Securities Seoul, Republic of Korea · FinancialNLPLab, MODULABS Seoul, Republic of Korea · Seoul University Seoul, Republic of Korea
Financial text embeddings must distinguish changes in event status, perspective, and obligations even when passages share similar wording. NMIXX adapts existing encoders through 18.8k source-linked triplets: paraphrases and Korean-English translations preserve meaning, while targeted financial rewrites introduce semantic contrasts. We examine this recipe across seven backbones on English and Korean financial and general-domain semantic textual similarity (STS), and analyze the composition and passage lengths of KorFinSTS. BGE-M3 attains the highest adapted financial correlations in this comparison, improving from 0.1969 to 0.2967 on FinSTS and from 0.0512 to 0.2732 on KorFinSTS. Its general English and Korean correlations decrease by 0.0391 and 0.0463. Across the seven models, five improve their mean financial correlation, but all reduce their mean general-domain correlation. Per-language comparisons and benchmark-weight sensitivity analysis reveal differences obscured by a single aggregate score. The study contributes a finance-specific supervision design and evidence for evaluating adaptation jointly with retained general semantic capability; direct cross-language retrieval remains outside its evaluation scope.
Figures & tables
Figure 1. NMIXX at a glance. (A) Financial document types motivate semantic-shift negatives; the sentence contrast is illustrative. (B) Triplet construction and adaptation of existing encoders. (C) BGE-M3’s changes on all four benchmarks connect the two research questions: financial improvements coexist with general-domain losses. Three panels show financial semantic distinctions, the NMIXX triplet pipeline, and BGE-M3 changes of plus 0.0998, plus 0.2220, minus 0.0391, and minus 0.0463 on the four benchmarks.
Figure 2. Positioning by supervision construction. TSDAE ( Wang et al., 2021 ) , GPL ( Wang et al., 2022 ) , and Fin-E5 ( Tang and Yang, 2025 ) illustrate distinct adaptation routes. NMIXX emphasizes financial meaning changes within source-linked triplets, alongside paraphrase/translation positives. These are design emphases, not exclusive capabilities or a performance ranking. The sentence contrast is illustrative. Four horizontal routes compare denoising reconstruction in TSDAE, generated queries and teacher labels in GPL, persona-based financial task synthesis in Fin-E5, and source-linked financial semantic-shift triplets in NMIXX.
Dataset
Lang.
Rows (k)
License
sujet-finance-instruct
ko
178
Apache-2.0
finance-legal-mrc
ko
359
CC-BY-SA 4.0
KorFin-ASC
ko
8.8
Apache-2.0
finance-embeddings-investopedia
en
206
CC-BY-NC 4.0
fingpt-sentiment-train
en
76.8
MIT
FNSPID
en
1,630
CC-BY-NC 4.0
Table 1. Public corpora collected before filtering.
Figure 3. NMIXX training-data construction. (A) Public corpora and synthetic augmentation pass through filtering, balancing, and expert review. (B) Semantic-shift negatives and meaning-preserving positives are generated and screened to form triplets. Stage counts use different units and do not imply a conversion rate. The sentence triplet illustrates the intended relation between anchor, positive, and negative. A two-stage pipeline shows 2.46 million raw records plus 25.9 thousand synthetic documents, a filtered pool of 46.1 thousand sentences, expert review, semantic-shift negative and paraphrase generation, judge thresholds of 8 and 9, and 18.8 thousand training triplets.
Figure 4. KorFinSTS dataset analysis. (A) Pair counts by source. (B) Median and interquartile range of text length, pooling both sentence columns within each source. Characters include stored whitespace; lengths are not model-token counts. Counts are 355 news, 500 disclosure, 421 research, and 715 legal pairs. Median character lengths are 240, 977, 371, and 834.5 respectively.
Model
License
Language Support
bge-en-icl
Apache-2.0
Mainly English
gte-Qwen2-1.5B-instruct
Apache-2.0
English & Chinese
e5-mistral-7b-instruct
MIT
Mainly English
bge-large-en-v1.5
MIT
Mainly English
all-MiniLM-L12-v2
Apache-2.0
Mainly English
instructor-base
Apache-2.0
Mainly English
Table 2. Baseline embedding models, licenses, and language support.
FinSTS
KorFinSTS
STS
KorSTS
Model
before
after
before
after
before
after
before
after
bge-en-icl
0.1668
0.2574
0.0511
−0.0745
0.8058
0.5965
0.7078
0.2487
gte-Qwen2-1.5B-inst
0.2858
0.2518
0.0094
0.2204
0.8592
0.7556
0.3742
0.4727
e5-mistral-7b-inst
0.1476
0.2641
0.1099
−0.1738
0.8768
0.6092
0.7495
0.1492
bge-large-en-v1.5
0.1675
0.1626
−0.2119
−0.1586
0.8752
0.8835
0.3320
0.2473
all-MiniLM-L12-v2
0.1909
0.2626
−0.1837
−0.1590
0.8309
0.7109
0.3858
0.1262
Table 3. Spearman’s ρ on four STS benchmarks before and after domain adaptation. The higher value of each before/after pair is bolded . Boldface denotes a numerical comparison, not statistical significance.
Figure 5. Adaptation changes across the four evaluation settings. Financial gains vary across languages and encoders (RQ1), while general-domain losses are widespread (RQ2). Numerical labels show differences between the four-decimal scores in Table 3 on a common color scale. A seven by four heatmap labels each adaptation change. BGE-M3 improves both finance scores and loses both general scores.
Figure 6. Trade-off and sensitivity analysis of Table 3 . (A) Each point averages two financial and two general changes; the dashed segment connects the two nondominated adaptations. Higher and farther right is preferable under this summary. (B) A hypothetical weight w interpolates between the general and financial means. This measures sensitivity to benchmark weights rather than deployment benefit. A scatter plot compares financial gain with general retention for seven models. A second plot varies the finance weight from zero to one. BGE-M3 crosses zero at about 0.210.
We introduce JFinTEB, the first comprehensive benchmark specifically designed for evaluating Japanese financial text embeddings. Existing embedding benchmarks provide limited coverage of language-specific and domain-specific aspects found in Japanese financial texts. Our benchmark encompasses diverse task categories including retrieval and classification tasks that reflect realistic and well-defined financial text processing scenarios. The retrieval tasks leverage instruction-following datasets and financial text generation queries, while classification tasks cover sentiment analysis, document categorization, and domain-specific classification challenges derived from economic survey data. We conduct extensive evaluations across a wide range of embedding models, including Japanese-specific models of various sizes, multilingual models, and commercial embedding services. We publicly release JFinTEB datasets and evaluation framework at https://github.com/retarfi/JFinTEB to facilitate future research and provide a standardized evaluation protocol for the Japanese financial text mining community. This work addresses a critical gap in Japanese financial text processing resources and establishes a foundation for advancing domain-specific embedding research.
Masahiro Suzuki, Hiroki Sakaji
Amova Asset Management Co., Ltd. · Tokyo, Japan · Hokkaido University +1
Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must first be transformed into synthetic training examples. We present a data-centric pipeline that constructs complementary corpora by mining open-source reasoning traces, distilling financial instruction data, and generating knowledge-graph-guided question-answer pairs from financial educational material. After semantic deduplication, three lightweight sequence classifiers select finance-relevant examples, reject under-specified questions, and identify tasks suitable for reinforcement learning with compact rule-based verifiers. For model adaptation, we study supervised fine-tuning and reinforcement learning, while self-distilled fine-tuning and post-training model merging are used to prevent the loss of financial capabilities already present in the starting model. We evaluate the adapted language models using FINESSE-Bench, reporting aggregate performance and changes relative to their starting checkpoints. Across the selected comparisons, ordinary SFT reduces FINESSE-Bench accuracy by 3.2-4.0 percentage points, whereas self-distilled SFT improves over the corresponding starting models by 1.0-2.8 points. Equal-weight merging recovers 3.0 points over its SFT parent and finishes 0.9 points above the original model; GRPO on hard tasks adds 0.4 points after self-distilled SFT or 3.0 points when applied directly to verifiable tasks. These results show that retention-aware adaptation can improve financial reasoning without the regressions observed after ordinary SFT.
Zhirayr Hayrapetyan, Andrei Kalmykov, Denis Kokosinskii +2
We present DS@GT's submission to FinMMEval 2026 Task 1, a multilingual financial exam question answering benchmark spanning English, Spanish, Greek, Chinese, and Hindi. Financial certification exams such as the CFA, EFPA, and CPA demand structured domain reasoning that standard NLP benchmarks do not capture, and this challenge compounds across languages where retrieval and representation infrastructure is underdeveloped. We build a retrieval-augmented pipeline on LangGraph that detects query language and retrieves semantically relevant exemplars from a 30,209-entry multilingual knowledge base using BGE-M3 embeddings and FAISS indexing. The system then scores answers via Retrieval-Augmented Direct Scoring (RADS), reading next-token log-probabilities over candidate option letters rather than generating free-form output. For low-resource languages, we fuse per-language and cross-lingual retrieval indices using weighted Reciprocal Rank Fusion. Model selection is language-routed: Qwen3-14B for Arabic, Chinese, and Hindi; Qwen2.5-14B for English; and Llama-3.1-8B for Greek, a routing derived from empirical ablations that reveal substantial language-asymmetric performance gaps. Notably, chain-of-thought prompting significantly degrades Greek accuracy (90.7% to 20.9%), and enabling Qwen3's default thinking mode collapses Arabic RADS performance to near-chance levels. Our results indicate that effective multilingual financial reasoning requires language-aware retrieval, model routing, and deliberate scoring strategy selection.
Justice Ayela, Kabir Sahni
Georgia Institute of Technology, North Ave NW, Atlanta, GA 30332