Test sets with recorded results: JevAlt decision models in three languages and Tholos-Bench for local tool-use agents.