FARD Lab · UBC Okanagan

Should you trust how an LLM explains its code?

A research project on the faithfulness of LLM explanations in the software domain, building benchmarks and metrics for code reasoning we can actually trust.

The problem

Two explanations. Same answer. Only one is true.

An LLM is asked: what does this return for [1, 5, 3, 8], k=4?

def process(nums, k):
    out = []
    for n in nums:
        if n > k:
            out.append(n * 2)
    return out

# expected: [10, 16]

The model returns [10, 16] and offers a rationale:

A. "I checked each n against the threshold k = 4 and doubled the ones that passed."
B. "I doubled every element of the list."

Both sound plausible. Both produce the right answer for this input. But only one reflects what the model actually relied on, and it matters which.

Why it matters

The reasoning is the product, not just the answer.

Stakes

LLMs are entering serious SE work

Code review, debugging, security analysis, autonomous agents. In each case the human consumes the model's reasoning, not just its verdict. Unfaithful explanations directly cost trust.

Gap

Faithfulness has a literature; code doesn't (yet)

Years of work on faithfulness in NLP have produced rigorous frameworks. Code reasoning has been treated as flat text. The vocabulary and tools haven't been adapted for programs.

Opportunity

Code is more tractable, not less

Programs have structure. We can run principled, validated, semantics-aware interventions on code that you cannot do in prose. The domain rewards rigor.

What we build

Benchmarks and metrics for trustworthy code reasoning.

The project develops evaluation frameworks that test whether an LLM's explanation of its code reasoning aligns with the model's actual decision process, going beyond "did the answer change" to ask whether the reasons it offered are the reasons it used.

Benchmark

A code-faithfulness evaluation suite

Curated programs and tasks designed to stress-test how LLMs explain their reasoning across diverse code patterns and reasoning structures.

Metrics

Behavioral tests for explanation faithfulness

A family of measures that probe whether an LLM's stated reasoning is causally connected to its prediction, adapted from the NLP faithfulness literature for the structure of programs.

Tooling

Open infrastructure for evaluating models

Released as the project matures, with the goal of making code-faithfulness evaluation reproducible and accessible to other research groups.

Findings

Empirical insights into where models break

Patterns of unfaithfulness across model families and scales — what kinds of reasoning are most prone to fabricated rationales, and which families fail in which ways.

Where we are

The foundation is built. The interesting questions are open.

Methodology, evaluation pipeline, and counterfactual generators are in place. We're now scaling the benchmark, sharpening the metrics, and pushing into harder reasoning tasks — and we're recruiting researchers to drive that work.