FARD Lab · UBC Okanagan
A research project on the faithfulness of LLM explanations in the software domain, building benchmarks and metrics for code reasoning we can actually trust.
The problem
An LLM is asked: what does this return for [1, 5, 3, 8], k=4?
def process(nums, k): out = [] for n in nums: if n > k: out.append(n * 2) return out # expected: [10, 16]
The model returns [10, 16] and offers a rationale:
n against the threshold k = 4 and doubled the ones that passed."
Both sound plausible. Both produce the right answer for this input. But only one reflects what the model actually relied on, and it matters which.
Why it matters
Stakes
Code review, debugging, security analysis, autonomous agents. In each case the human consumes the model's reasoning, not just its verdict. Unfaithful explanations directly cost trust.
Gap
Years of work on faithfulness in NLP have produced rigorous frameworks. Code reasoning has been treated as flat text. The vocabulary and tools haven't been adapted for programs.
Opportunity
Programs have structure. We can run principled, validated, semantics-aware interventions on code that you cannot do in prose. The domain rewards rigor.
What we build
The project develops evaluation frameworks that test whether an LLM's explanation of its code reasoning aligns with the model's actual decision process, going beyond "did the answer change" to ask whether the reasons it offered are the reasons it used.
Benchmark
Curated programs and tasks designed to stress-test how LLMs explain their reasoning across diverse code patterns and reasoning structures.
Metrics
A family of measures that probe whether an LLM's stated reasoning is causally connected to its prediction, adapted from the NLP faithfulness literature for the structure of programs.
Tooling
Released as the project matures, with the goal of making code-faithfulness evaluation reproducible and accessible to other research groups.
Findings
Patterns of unfaithfulness across model families and scales — what kinds of reasoning are most prone to fabricated rationales, and which families fail in which ways.
Where we are
Methodology, evaluation pipeline, and counterfactual generators are in place. We're now scaling the benchmark, sharpening the metrics, and pushing into harder reasoning tasks — and we're recruiting researchers to drive that work.