Boi Faltings, Robert West, Antoine Bosselut, Debjit Paul
Large language models (LLMs) have been shown to perform better when asked to reason step-by-step before answering a question. However, it is unclear to what degree the model's final answer is faithful to the stated reasoning steps. In this paper, we perfor ...
Association for Computational Linguistics (ACL)2024