Models That Know How Evaluations Are Designed Score Safer Paper • 2605.28591 • Published 25 days ago • 10 • 5
Models That Know How Evaluations Are Designed Score Safer Paper • 2605.28591 • Published 25 days ago • 10
Models That Know How Evaluations Are Designed Score Safer Paper • 2605.28591 • Published 25 days ago • 10
Models That Know How Evaluations Are Designed Score Safer Paper • 2605.28591 • Published 25 days ago • 10
🕵️🛡️ Evaluation Meta Knowledge Collection 2026 arXiv preprint. Models fine-tuned on documents describing typical evaluation traits show safer behavior by having increased refusal rates and low • 11 items • Updated 11 days ago • 2
🕵️🛡️ Evaluation Meta Knowledge Collection 2026 arXiv preprint. Models fine-tuned on documents describing typical evaluation traits show safer behavior by having increased refusal rates and low • 11 items • Updated 11 days ago • 2