Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models Paper • 2606.13603 • Published Jun 11
Distilling Formal Logic into Neural Spaces: A Kernel Alignment Approach for Signal Temporal Logic Paper • 2603.05198 • Published Mar 5
Bridging Logic and Learning: Decoding Temporal Logic Embeddings via Transformers Paper • 2507.07808 • Published Jul 10, 2025
Predicting Future Behaviors in Reasoning Models Enables Better Steering Paper • 2606.11172 • Published Jun 9 • 1
Interpreto: An Explainability Library for Transformers Paper • 2512.09730 • Published Dec 10, 2025 • 1
A Behavioural and Representational Evaluation of Goal-Directedness in Language Model Agents Paper • 2602.08964 • Published Feb 9 • 1
Simplifying Outcomes of Language Model Component Analyses with ELIA Paper • 2602.18262 • Published Feb 20 • 1
Interpreting Language Models Through Concept Descriptions: A Survey Paper • 2510.01048 • Published Oct 1, 2025 • 2
Infherno: End-to-end Agent-based FHIR Resource Synthesis from Free-form Clinical Notes Paper • 2507.12261 • Published Jul 16, 2025 • 2
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data Paper • 2507.00152 • Published Jun 30, 2025 • 1
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework Paper • 2506.15538 • Published Jun 18, 2025 • 1
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals Paper • 2505.13972 • Published May 20, 2025 • 1
Through a Compressed Lens: Investigating the Impact of Quantization on LLM Explainability and Interpretability Paper • 2505.13963 • Published May 20, 2025 • 1
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods Paper • 2505.01198 • Published May 2, 2025 • 2
Inseq: An Interpretability Toolkit for Sequence Generation Models Paper • 2302.13942 • Published Feb 27, 2023 • 1