IBM’s PatchTST-FM-r2 Climbs GIFT-Eval With Open Time-Series Model
Granite Time Series PatchTST-FM-r2 ranked near the top of replicable zero-shot forecasting models on GIFT-Eval, pairing conformer blocks, long contexts and permissive licensing.

IBM’s latest open time-series model has moved near the top of a major zero-shot forecasting benchmark, with the company’s research team detailing the Granite Time Series PatchTST-FM-r2 release in a Hugging Face post.
The model is designed for forecasting jobs where teams want to use a pretrained system without building and maintaining a separate model for every dataset.
IBM Research describes PatchTST-FM-r2 as an update to its earlier PatchTST-FM-r1, with a new architecture, a larger pretraining corpus, probabilistic forecasting, missing-value imputation and a roughly 385 million-parameter scale.
Benchmark performance is the central claim.
On September 8, PatchTST-FM-r2 ranked second among replicable zero-shot models on GIFT-Eval for both CRPS and MASE, two error measures where lower scores are better.
IBM’s published comparison lists a geometric-mean CRPS of 0.467, behind TimesFM-3 in that comparison, and a geometric-mean MASE of 0.6846.
Within the same zero-shot and replicable category, IBM positioned it as the highest-performing model released under permissive commercial-friendly licenses.
The comparison remained competitive when GIFT-Eval’s pretrained category was added.
Some models in that wider group are allowed to include training portions of the benchmark’s evaluation datasets in their pretraining corpora.
Even against that broader set, PatchTST-FM-r2 placed third for CRPS and fourth for MASE among replicable models, ahead of Chronos-2, Timer-S1 and Toto variants listed in IBM’s comparison.
The architectural change is meant to explain part of that movement.
PatchTST-FM-r2 keeps the patch-based representation from the PatchTST family, but replaces the previous transformer block with a conformer-style block that combines self-attention with temporal convolution.
The design lets convolution handle local time-series structure while attention is used for longer-range relationships.
IBM also added 50 percent overlapping patches, Hamming-window weighting, overlap-and-add forecasting, extra normalization and an expansion from 20 to 30 blocks.
Those changes give the model several operating characteristics that matter beyond leaderboard rank.
PatchTST-FM-r2 can look back across sequences as long as 8,192 time steps, and its forecast output spans 99 quantile levels over flexible horizons.
That gives users both point estimates and uncertainty ranges, while the public release packages the weights with implementation details, inference tooling and reproducibility code.
The training-data disclosure is also part of the release.
IBM listed four pretraining sources: selected GiftEvalPretrain datasets, custom KernelSynth-based synthetic data, a TSMixup collection that avoids the GIFT-Eval evaluation datasets and about 500,000 synthetic CauKer sequences, with 4,096 points in each sequence.
That record does not remove an adopter’s own licensing or governance review, but it gives enterprise users more information about benchmark leakage and deployment risk than an opaque corpus would.
Licensing broadens the intended audience.
The model is available under Apache 2.0 and OpenMDW 1.0, with the implementation in the Granite-TSFM repository and backward compatibility for PatchTST-FM-r1 checkpoints.
IBM’s examples show the model loaded from the Hugging Face Hub without fine-tuning, using recent history from a series to produce future forecasts and requested quantiles.
The same pattern can be tried against demand, sensor telemetry, CPU utilisation, energy use, transaction volume, traffic or price series when the data is regularly sampled.
The release also connects IBM’s time-series work to streaming deployments.
IBM and Confluent recently made several Granite Time Series models available through an Early Access program in Confluent Cloud, covering FlowState-r1.1, TTM-r3, TSPulse and the earlier PatchTST-FM-r1.
Apache Flink on Confluent Cloud can run forecasting and anomaly-detection inference against live streams, reducing the need to move operational data into a separate machine-learning environment.




















