Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
Organizations: Independent Researcher
Abstract
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. -- on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 (). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
Figures & tables
| VectraYX-600M | VectraYX-1B | |
| Parameters | 661.6M | 1,109M |
| Layers | 18 | 22 |
| 1,792 | 2,048 | |
| 4,864 | 5,504 | |
| Heads (Q / KV, GQA) | 14 / 2 | 16 / 4 |
| NoPE period / RoPE | 4 / | 4 / |
| Component | Configured weight | Realized fraction |
| Spanish (Wiki-ES, FineWeb2-ES) | 0.45 | 0.34 |
| Code and technical (Stack, ICPC) | 0.23 | 0.65 |
| Cyber (CVE/NVD-derived) | 0.22 | 0.007 |
| Instruction-formatted SFT | 0.10 | 0.001 |
| Benchmark | 600M (step 154K) | 1B (recorded) |
| B1 CVE QA (keyword recall) | 0.342 | 0.342 |
| B2 Classification (acc.) | 0.240 | 0.205 |
| B3 Commands (tool match) | 0.17 | 0.24 |
| B4 Tool use | 0.660 | 0.650 † |
| B5 Conversational | 0.589 | 0.586 |
| Checkpoint | Pass / trials | Gen. | |
| 600M step 154K | 6 / 6 | yes | greedy |
| 1B phase-3 step 20K (pre-vision) | 0 / 4–6 | — | – |
| 1B phase-4b SFT step 400 (post) | 0 / 4–6 | — | – |
| 1B repaired, step 2202 (§ 6.5 ) | 4 / 4 ‡ | exact | 0.996–0.998 |
| historical run, checkpoint lost | 4 / 4 ‡ | exact | 0.995–0.998 |
| Data / config | Strict outcome | |
| A: dedicated phase-3 ( 6B tok, 20K steps) | 30% tool-SFT mixture within broad phase-3 mix | 0/4–6; – |
| B: narrow retune (486 steps) | tool corpus small cyber sample; LR ; embeddings frozen | not validly measured : recorded 0/8 by a harness later found to omit the preamble and drop special tokens; loss fell to 0.05–0.4 |
| C: informed recipe, reproduced (2,202 steps, 3.3 h, one rented GPU) | OASST ES/EN tool corpus domain reasoning traces; LR ; embeddings nominally unfrozen—but numerically inert, § 6.6 ; warm start phase4a_v0 | 4/4 verbatim (exact), – ; 8/8 novel-prompt battery (7/7 tool choice, plus correct suppression) |
| C ∗ : same recipe, historical run (checkpoint and parent lost) | identical corpus, LR and step count; warm start v3b_step_001900 | 4/4 verbatim (exact), – ; 7/8 novel-prompt battery |
| Step | no-call | Emit | Gen. | Suppr. | ||
| 1K | 0.858 | 0.673 | 5/5 | 4 | no | |
| 5K | 0.973 | 0.997 | 5/5 | 3 | no | |
| 10K | 0.993 | 0.988 | 5/5 | 3 | no | |
| 20K | 0.996 | 0.99991 | 5/5 | 3 | no | |
| 30K | 0.9998 | 0.9998 | 5/5 | 3 | no | |
| 60K | 0.9986 | 0.9977 | 5/5 | 3 | no |
| Phase 2 (web-heavy pretraining) | Phase 3 (dedicated tool-SFT) | ||||||||||
| Step | no-call | Emit | Gen. | Step | no-call | Emit | B4 | ||||
| 36K | 0.432 | 3/5 | 0 | 4.0K | 0/5 | — | |||||
| 38K | 0/5 | — | 5.7K | 0/5 | 0.020 | ||||||
| 40K | 0/5 | — | 6.0K | 0/5 | — | ||||||
| 44K | 0/5 | — | 8.0K | 0/5 | 0.075 | ||||||
| 48K | 0/5 | — | 10.0K § | 0/5 | 0.020 | ||||||
| Axis | VectraYX-600M | 1B pre-rep. | 1B repaired |
| Decision | 1.000 | 0.006 | 1.000 |
| Well-formed | 0.994 | 0.006 | 1.000 |
| Tool correct | 0.776 | — | 0.717 |
| Arg. shape | 0.855 | — | 0.819 |
| Entity match | 0.560 | — | 0.667 |
| Headline pass | 0.428 | 0.000 | 0.536 |
| Family | VectraYX-600M | 1B repaired | |
| nvd_get_cve | 28 | 0.893 | 1.000 |
| anchor | 7 | 0.714 | 0.857 |
| cisa_kev_check | 28 | 0.607 | 0.571 |
| otx_check_ioc | 27 | 0.444 | 0.519 |
| bash_exec | 28 | 0.214 | 0.500 |
| terse_control | 20 | 0.200 | 0.500 |
| Arm | Head. | Tool | Entity | Soft | Supp. | Over |
| pass | corr. | match | miss | pass | trig. | |
| narrow / high LR | 0.608 | 0.783 | 0.748 | 0.419 | 0.458 | 0.542 |
| narrow / low LR | 0.542 | 0.685 | 0.702 | 0.387 | 0.375 | 0.625 |
| diverse / low LR | 0.584 | 0.759 | 0.719 | 0.452 | 0.583 | 0.403 |
| diverse / high (s42) | 0.602 | 0.795 | 0.741 | 0.452 | 0.514 | 0.486 |
| diverse / high (s43) | 0.602 | 0.771 | 0.726 | 0.387 | 0.500 | 0.500 |
| Arm @ step 880 | Head. | Tool | Entity | Soft | Supp. | Over |
| pass | corr. | match | miss | pass | trig. | |
| narrow / low (s42) | 0.536 | 0.691 | 0.701 | 0.419 | 0.389 | 0.611 |
| narrow / low (s43) | 0.548 | 0.685 | 0.687 | 0.355 | 0.569 | 0.431 |
| diverse / low (s42) | 0.578 | 0.759 | 0.711 | 0.452 | 0.556 | 0.431 |
| 600M | SFT | SFT | Discordant | |
| (no SFT) | baseline | distilled | ( ) | |
| Headline pass (166) | 0.428 | 0.548 | 0.548 | 5 / 5 (1.00) |
| Well-formed | 0.994 | 0.934 | 0.922 | 7 / 5 (0.77) |
| Tool correct | 0.776 | 0.794 | 0.817 | 1 / 6 (0.13) |
| Entity match | 0.560 | 0.714 | 0.712 | 2 / 4 (0.69) |
| Suppression pass (72) | — | 0.014 | 0.056 | 0 / 3 (0.25) |