OPEN MODELS. INFORMED CHOICES.
RSS ↗
Uncensored AI News

Intelligence belongs
in the open.

Search
Open ModelsNews · 2 MIN READ

Preprint Diagnoses Attenuated Dynamics in Time-Series Foundation Models

An arXiv preprint dated 2026-10-06 tests Chronos-2, TimesFM-2.5 and TabPFN-TS on exact counterfactual inputs for forced engineering systems, revealing memoryless behavior or…

Schematic diagram contrasting attenuated time-series foundation model responses with accurate classical identification under counterfactual inputs
Editorial illustration; not a photograph of a reported event.
THE TAKEAWAY
  • TimesFM-2.5 and TabPFN-TS act as memoryless functions under default covariate interfaces, returning R²=1.000 same-time effects.
  • Chronos-2 identifies dynamics in-context but attenuates them to 0.33-0.80 of true effect; error plateaus at 0.57 even with 8192 samples.
  • Synthetic forced-system fine-tuning in 26 minutes restores sensitivity to 0.83-0.96 and outperforms structure-agnostic methods on certain nonlinear plants.

Counterfactual Testing Framework

The preprint submitted on 2026-10-06 by Won, Hong-In evaluates covariate-aware time-series foundation models for training-free what-if analysis on instrumented plants.

Researchers compared three models—Chronos-2, TimesFM-2.5 and TabPFN-TS—against classical ARX identification using identical context windows on forced engineering systems with known exact counterfactuals.

Paired counterfactual inputs combined with shuffled future inputs on measured records were used to isolate whether the covariate interface can represent dynamics and whether the pretraining prior matches the plant's time scale.

Key Findings on Model Behavior

Through default covariate interfaces, TimesFM-2.5 and TabPFN-TS proved memoryless: the predicted output change equals the input change at the same timestep.

Chronos-2 recovered dynamics in context but consistently attenuated them. Its predicted effect ranged from only 0.33 to 0.80 of the true effect, and the recovered impulse response showed incorrect shape.

On a one-degree-of-freedom oscillator, Chronos-2 error leveled off at 0.57 with 8192 context samples, while classical ARX fitted to just 256 samples achieved 0.02 error.

Repair via Context Dither and Fine-Tuning

Adding context dither at inference time reduced what-if error across all six synthetic system classes without any training.

A 26-minute fine-tune on synthetic forced systems restored response magnitude, achieving sensitivity between 0.83 and 0.96. The repaired model outperformed structure-agnostic identification on Wiener-Hammerstein systems and a held-out friction class.

A specialised in-context identifier trained on the same forced-system data performed comparably, indicating that the synthetic data itself carries most of the performance gain.

Performance on Measured Plants and Trade-offs

On three of four measured physical plants, classical system identification remained clearly superior.

The fine-tuned foundation model lost part of its univariate forecasting skill after the repair procedure.

The evidence, limited to the preprint abstract, shows that only counterfactual pairs successfully exposed the attenuation issue.