Principled Thoughts for Latent Recursive LLM Systems

Fahd Seddik1, Fatemeh Fard1
1FARD Lab, University of British Columbia
REST. The top row shows the four failures of Cross-Entropy (CE) only training of the thought. REST adds a $\beta$-weighted loss on to the CE, with one term per failure, under the same data and compute. (a) Accuracy gain for single- and multi-agent settings across Small (1–2B) and Large (3–4B) models. (b)–(f) REST thoughts spread apart instead of collapsing, reach a lower training CE and a higher accuracy, recover more accuracy when they replace the oracle text, reach a final answer more often, and keep their superposition.
Figure 1. REST. The top row shows the four failures of Cross-Entropy (CE) only training of the thought. REST adds a $\beta$-weighted loss on to the CE, with one term per failure, under the same data and compute. (a) Accuracy gain for single- and multi-agent settings across Small (1–2B) and Large (3–4B) models. (b)–(f) REST thoughts spread apart instead of collapsing, reach a lower training CE and a higher accuracy, recover more accuracy when they replace the oracle text, reach a final answer more often, and keep their superposition.

Abstract

Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret.

REST: REpresentation Supervised Thought(s)

Overview of REST. Latent systems pass a thought $\mathbf{T}$ through a trained outer link (left). REST adds a $\beta$-weighted loss built from the four properties of $\mathbf{T}$ to the CE objective (middle). CE-only thoughts collapse and REST improves accuracy under the same data and compute (right).
Figure 2. Overview of REST. Latent systems pass a thought $\mathbf{T}$ through a trained outer link (left). REST adds a $\beta$-weighted loss built from the four properties of $\mathbf{T}$ to the CE objective (middle). CE-only thoughts collapse and REST improves accuracy under the same data and compute (right).

Failures of latent reasoning training with CE only

Why CE is not enough. Four failures of CE thoughts and the property that addresses each.
Figure 3. Why CE is not enough. Four failures of CE thoughts and the property that addresses each.

Single-Agent extension and how REST applies

Single agent. $\mathcal{R}_{\mathrm{in}}$ runs at each of the $m'$ steps, and $\mathcal{R}_\psi$ once per round.
Figure 4. Single agent. $\mathcal{R}_{\mathrm{in}}$ runs at each of the $m'$ steps, and $\mathcal{R}_\psi$ once per round.

Evaluation

Experimental Setup

Table 1. Light and Scaled systems.
Light and Scaled systems.

Single-Agent Evaluation

Table 2. Single-agent, Light vs Scaled, at $r=1$ and $r=3$ respectively, each row at its best $\beta$.
Single-agent, Light vs Scaled, at $r=1$ and $r=3$ respectively, each row at its best $\beta$.

Multi-Agent Evaluation

Table 3. Multi-agent, Light vs Scaled, at $r=1$ and $r=3$ respectively, each row at its best $\beta$.
Multi-agent, Light vs Scaled, at $r=1$ and $r=3$ respectively, each row at its best $\beta$.

Analysis

Acc. Change vs $r$
Figure 5. Acc. Change vs $r$
Table 4. REST against auxiliary-loss baselines, Light vs Scaled.
REST against auxiliary-loss baselines, Light vs Scaled.
REST against CE-only. (a) Decoded thoughts compared to output / input. (b) Replacing oracle text with thought. (c) Answer rate vs token usage. (d) Effective Superposition.
Figure 6. REST against CE-only. (a) Decoded thoughts compared to output / input. (b) Replacing oracle text with thought. (c) Answer rate vs token usage. (d) Effective Superposition.
REST against CE-only. (a) PCA for $\mathbf{T}$. CE thoughts collapse into dense clusters, while REST spreads thoughts apart. (b) Training CE under causality. More details in Appendix D.
Figure 7. REST against CE-only. (a) PCA for $\mathbf{T}$. CE thoughts collapse into dense clusters, while REST spreads thoughts apart. (b) Training CE under causality. More details in Appendix D.

Related Work

Positioning of REST.
Figure 8. Positioning of REST.

BibTeX

@misc{seddik2026principledthoughtslatentrecursive,
  title         = {Principled Thoughts for Latent Recursive LLM Systems},
  author        = {Fahd Seddik and Fatemeh Fard},
  year          = {2026},
  eprint        = {2609.36159},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2609.36159}
}