ufakzeka-karar: An Open Turkish Typed-Decision Model with Order-Invariant Option Scoring
Organizations: ufak AI
Abstract
ufakzeka-karar is an open Turkish decision model with 182,494,466 parameters. Given a Turkish text and questions of a fixed answer type (a choice, a level on an ordered scale, or yes or no), it returns a temperature-scaled probability for every option and an expected error that serves as a "not sure" signal, without generating text and in one CPU forward pass for up to ten options. Built on the lab's ufakzeka-1-base, its head scores each option blind to the others at shared positions, so the answer does not depend on option order. A sequential head trained with shuffled options was about as accurate but changed 2.3 to 2.8 percent of its answers when only the option order changed; REINFORCE lost 10.2 points (0.102) of macro F1 to cross-entropy. On the open set of HakemBench v1.0 (4,275 questions, 7 tracks) the released model ranks 7th of 16 rows with a composite of 0.660 (95% interval 0.642 to 0.677). Temperature scaling lowers calibration error (smooth ECE) on the development set but raises it on held-out support questions, from 0.027 to 0.045 for the first scored run, which never trained on them; the released model later trained on them, so its 0.036 to 0.064 is not an unseen-question test. The released model is the last of three runs scored on HakemBench, and its numbers are not blind. The second run's new training data was aimed at the first run's errors on the full test set in guardrails, moderation and customer support, and the released run was trained after the second run's guardrail results on the full test set were read, under a protocol fixed in writing before any of its data, code or runs. All its numbers come after these readings; its guardrail, moderation and customer support numbers carry the flag "shaped by reading the test results". With every model scored on the other four tracks only, its composite is 0.678, 6th of 16. Weights and code are under Apache-2.0.
Figures & tables
| attacks | look-alikes | human | held-out rows | support dev. set | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| run | composite | interval | caught | passed | check | raw | with | raw | with | |
| Run 1 | 0.556 | [0.537, 0.574] | 205 | 70 | 0.405 | per type | 0.027 | 0.045 | 0.051 | 0.034 |
| Run 2 | 0.645 | [0.628, 0.663] | 106 | 195 | 0.524 | 1.2932 | 0.036 | 0.077 | 0.062 | 0.033 |
| Run 3 (released) | 0.660 | [0.642, 0.677] | 194 | 128 | 0.476 | 1.2614 | 0.036 | 0.064 | 0.050 | 0.038 |
| decision | selective | median | |||||
|---|---|---|---|---|---|---|---|
| rank | model | composite | interval | quality | calibration | automation | ms |
| 1 | Gemini 3.8 Flash | 0.888 | [0.876, 0.898] | 0.901 | 0.806 | 0.964 | 1,294.2 |
| 2 | GPT-5.6 Sol | 0.842 | [0.827, 0.856] | 0.871 | 0.727 | 0.943 | 808.8 |
| 3 | GLM 5.3 | 0.827 | [0.813, 0.839] | 0.855 | 0.706 | 0.936 | 572.2 |
| 4 | Jev 1.13 | 0.825 | [0.811, 0.837] | 0.851 | 0.706 | 0.935 | 135.0 |
| 5 | Kev 4B | 0.688 | [0.673, 0.702] | 0.747 | 0.503 | 0.866 | 63.3 |
| accuracy | Brier score | smooth ECE | |||||||
|---|---|---|---|---|---|---|---|---|---|
| set | chance | Run 1 | Run 3 | Run 3 interval | Run 1 | Run 3 | Run 1 | Run 3 | |
| MASSIVE 1.1 tr-TR | 2,937 | 0.017 | 0.756 | 0.745 | [0.728, 0.761] | 0.355 | 0.359 | 0.029 | 0.025 |
| MiDe22 | 1,012 | 0.333 | 0.766 | 0.741 | [0.714, 0.767] | 0.334 | 0.350 | 0.035 | 0.030 |
| OffensEval-TR 2020 A | 3,528 | 0.500 | 0.866 | 0.844 | [0.832, 0.857] | 0.204 | 0.225 | 0.021 | 0.020 |
| MMLU-Pro-TR | 11,838 | 0.111 | 0.107 | 0.098 | [0.092, 0.103] | 0.936 | 0.941 | 0.126 | 0.135 |
| XCOPA, Turkish | 500 | 0.500 | 0.562 | 0.584 | [0.542, 0.626] | 0.602 | 0.567 | 0.196 | 0.171 |
| compared choice | primary measure, points [95% interval] | other findings |
|---|---|---|
| masked-language conversion, 1B tokens | [ , ] on the three general sets | adoption required ; mean of four Turkish sets 0.7578 against 0.7598 for the causal backbone |
| BERTurk [ 25 ] , same head and data | [ , ] on decision tasks | [ , ] on the three general sets |
| MoganBert-TR [ 26 ] , same head and data | [ , ] on decision tasks | [ , ] on held-out questions |
| continued pretraining, 300M tokens | [ , ] | backbone kept |
| down-weighting easy rows [ 27 ] | [ , ] | rejected |
| sequential head, shuffled options | Run 1 [ , ]; second round [ , ] | 2.3 to 2.8% of answers change under reordering |