Preprint

Linear AI model points to critical behavior at a learning threshold

An arXiv preprint reports that average error stays finite while fluctuations in learned parameters become singular in a stylized in-context learning model.

An analytic model of in-context learning points to a sharp critical point where the average prediction error remains finite, but fluctuations in the learned parameters become singular. In the model’s quenched, or sample-specific, calculation, the covariance of the learned vector diverges when normalized sample complexity reaches 1.

The question behind the calculation is whether the interpolation singularity in linear in-context learning can be treated as a critical phenomenon, and which quantity fluctuates at the threshold. The work uses a linear regression setup in which H represents the context and query together, while P is the learning vector. Pretraining fits P to n independent context constraints, and the analysis is compared with finite-dimensional numerical solutions.

The difference lies in what is averaged

The central methodological choice is whether to average away the randomness of the training sample. The paper calls the first route annealed averaging. It calls the second quenched, or cavity, averaging, because the training-sample disorder remains in the learned parameters. That distinction separates the finite annealed error from the divergent quenched covariance at the threshold.

For the error calculation, spectral sums are rewritten with the Stieltjes transform of the Marchenko–Pastur distribution, a mathematical device for handling the model’s eigenvalue spectrum. After that rewrite, the resulting expressions depend on the task pool only through a task-diversity parameter.

A physics-style map

The authors next recast the model’s self-consistency equation in a Landau construction. A quantity called the order parameter represents the model’s state, normalized sample complexity plays the role of temperature, and the ridge parameter acts as a conjugate field. In the construction, that field is defined as normalized sample complexity multiplied by ridge strength.

The susceptibility in this construction is the inverse of the potential’s curvature at its stationary order parameter. In plain terms, it grows as the potential flattens around the selected point. The paper links this same susceptibility to the singular fluctuation contribution to quenched prediction error, representing the error behavior as a loss of curvature in the Landau potential.

For any fixed positive inverse-context parameter, the critical point remains at normalized sample complexity 1. The reported critical exponents, which describe how quantities scale near the point, are 1 for the order parameter, 2 for the response to the conjugate field and 1 for the susceptibility. This is the scaling pattern reported for the finite-inverse-context regime.

Where the scaling changes

The model also gives the order parameter a geometric interpretation in the ridgeless, zero-field limit. There, it is related to the density of zero modes, or flat directions, in the relaxation matrix. The relation is written using a resolvent, a spectral quantity indexed by task diversity.

The reported pseudogap, a region of suppressed residual order, lies between the task-diversity level and normalized sample complexity 1. The remaining order parameter is on the scale of the inverse-context parameter. The line where normalized sample complexity equals task diversity becomes a true critical boundary only when the inverse-context parameter is taken to zero from above first.

At the strict zero-inverse-context endpoint, where task diversity and normalized sample complexity both equal 1, the theory reports a nonanalytic power term with exponent 5/2. Its critical exponents change to 2 for the order parameter, three-halves for the field response and 1 for the susceptibility. The endpoint therefore has different scaling from the fixed positive-inverse-context case.

A finite-dimensional check

The numerical test examined whether the order-parameter picture could be recovered in finite dimensions. It used dimension 30, up to 1,024 independent task pools, independently generated pretraining realizations and zero noise variance. The extracted maps reproduced the suppressed pseudogap region and agreed well with the self-consistency solution. A fitted scaling exponent rose smoothly from approximately 0 to approximately 1 at the crossover line where normalized sample complexity equals task diversity.

A result tied to the model

The evidence remains tied to the stated linear regression setup and the finite simulation design. The paper reports no uncertainty intervals or formal goodness-of-fit statistics for the numerical agreement, so the size of possible deviations is not quantified. The results support the proposed description within this model, but do not test full nonlinear transformer systems or real-world in-context learning data.

The work is an arXiv version 1 preprint. Its front matter gives 28 August 2026 as the arXiv date and also displays a separate dated line for 31 August 2026.

Paper data and sources

Original title: Landau theory of quenched criticality in linear in-context learning
Authors: Daesik Kim, Sumin Choi, Hyojae Jeon, Jung Hoon Han
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.