System One models output constrained decisions and probability distributions rather than free-form text generation. While prevailing paradigms rely on structured schema objects to encode state, intent, and candidate choices, we revisit a fully natural language-based System One interface. In this framework, both the user request and each candidate option are expressed in natural language, supported by multimodal (image and video) auxiliary inputs. We introduce MetaEncoder, which fine-tunes a pre-trained Muse-Glimmer 30B decoder into an instruction-following decision-making encoder. To scale effectively across both small closed-set (< 256) and massive open-set (millions) candidate spaces, MetaEncoder employs a bi-encoder architecture trained via unidirectional contrastive learning for request-candidate alignment. We conduct extensive evaluations across 11 benchmark suites and 190 tasks spanning multimodal decision-making, understanding (closed-set) and retrieval (open-set), highlighting where MetaEncoder beats SOTA multimodal encoders, as well as its current limits on reasoning-intensive tasks.
Figures & tables
Capability
Benchmark
Modality
Candidates
Tasks
Metric
Decision making
JEVBench
Text
Closed
3
Accuracy
ImaJEV-Bench
Multimodal
Closed
3
Accuracy
Multimodal understanding
MMLU
Text
Closed
57
Hit@1
MMMU
Multimodal
Closed
30
Hit@1
Video-MMMU
Multimodal
Closed
3
Hit@1
NaturalBench
Multimodal
Closed
1
Hit@1
Table 1: Evaluation benchmarks. Tasks counts the scored task units for each suite. Closed-set tasks score each request against its own option set (fewer than 256 candidates). Open-set tasks score against a shared and larger candidate pool.
Category
Benchmark
MetaEncoder
OpenJev (27B)
WeMM-Embedding (9B)
Qwen3-VL-Embedding (8B)
Decision making
JevBench (orig)
0.9444
0.9861
0.6250
0.6389
JevBench (easy)
1.0
1.0
0.9792
0.9375
JevBench (hard)
0.7387
0.7658
0.4865
0.4324
Jev Intelligence
81.4
84.5
45.8
42.1
Multimodal decision making
ImaJev (dev)
0.8555
–
0.8117
0.7792
General knowledge understanding
MMLU
0.7767
0.8601
0.7325
0.6544
Table 2: Benchmark Evaluation Results for MetaEncoder compared to SOTA System One decision makers and SOTA retrieval models.
Benchmark
Without options
With options
Δ
MMLU (57 subjects)
0.5645
0.7456
+0.1811
MMMU (30 subjects)
0.4875
0.5629
+0.0754
MMMU-Pro (4 options)
0.4699
0.5553
+0.0854
MMMU-Pro (10 options)
0.3503
0.4237
+0.0734
Table 3: Effect of appending candidate option lists to the request prompt. Scores reflect unweighted subject means (MMLU, MMMU) or Hit@1 (MMMU-Pro).