Period ending 2026-09-14
4 new papers
A weekly snapshot of new work published in Feature Selection.
Twelve weeks of publication activity for this topic as it is defined today.
Weekly history
What was published in this topic, kept on the site without email delivery.
Period ending 2026-09-14
A weekly snapshot of new work published in Feature Selection.
Period ending 2026-09-07
A weekly snapshot of new work published in Feature Selection.
84 papers
Can this object be found in the desert?'' or Is this prompt malicious?'' We measure how the instruction changes the model's internal representation using at a single readout point. We explore eight different MAGs. The extracted reasoning features predict the models' own world understanding and judgment, can be approximated into a single activation direction, we found that some features are more linearly represented and some less, this linear representation, which is vector steering, can change the LLMs' decisions through activation steering by injecting reasoning features. Finally, we use the same method to select the best training datasets for prompt-injection classifier probes: while similarity between ordinary activations is almost unrelated to downstream performance, RFD-based similarity achieves Top-1 and Top-2 accuracy.