Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Paper • 2606.13603 • Published Jun 11
Distilling Formal Logic into Neural Spaces: A Kernel Alignment Approach for Signal Temporal Logic Paper • 2603.05198 • Published Mar 5
Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers Paper • 2507.07808 • Published Jul 10, 2025
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Diverse Deception Probes Collection Linear probes trained on diverse deception data to detect dishonest completions across model families (OLMo, Qwen, Gemma). • 5 items • Updated Mar 18 • 1
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads Paper • 2607.01002 • Published 30 days ago • 19